You can find our previous writings on why evaluating scheming risks require insight into the training pipeline (and not just the final checkpoint) here:
apolloresearch.ai/science/we-nee…
We strongly feel it’s crucial that evaluators have the right to publish important findings without editorial control or undue redaction by the AI company. Ideally, regulators would provide strong protections for such transparency requirements. 2/3
We’re excited to see Dario and Sam commit to “embedded evaluators who have employee-like access to verify safety practices and report incidents.” [...] “whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes.”
We’ve long advocated for third-party evaluators to have deep access. Risks from internal deployment and scheming cannot be appropriately assessed without direct insight into the training pipelines and internal deployment practices.
Apollo is looking forward to working on Embedded Evaluations with all Frontier AI developers. 1/3
As always, Apollo is hiring! New this time around: Founding Staff for our DC office.
It's becoming increasingly urgent and impactful to accurately communicate cutting edge research and its implications to the USG, orient their actions, and advance meaningful interventions. Help us with this impactful work and apply here:
jobs.lever.co/apolloresearch…
In the rest of the announcement post, we describe our design and methodology of Watcher Live in detail.
1. Threat modelling: deciding which model failures and attacks Watcher Live should address
2. Rubric design: mapping different threats and failures to decision criteria that ultimately result in an approval or block decision.
3. Evaluation: measuring the effectiveness, cost and latency of different Watcher configurations. We also describe how we created the evaluation datasets.
We’re releasing a new and improved version of Watcher Live!
Watcher Live is a real-time coding agent monitor that blocks dangerous actions to prevent data leaks, repo deletions, or scope overreach.
There is a free and an Enterprise version.
As should be very clear from the last months, security is becoming even more important in the AI age. We’re looking to grow our security team to defend against external and internal, human and non-human adversaries.
AI Security Researcher: jobs.lever.co/apolloresearch…
Security Engineer: jobs.lever.co/apolloresearch…
Every role is open in both offices, London and San Francisco. We provide visa sponsorship. A formal AI safety background is not required.
All open roles: jobs.lever.co/apolloresearch
We’re doubling down on our control and monitoring research. We have recently successfully red-teamed Anthropic’s auto-mode and will run more such campaigns in the future with multiple labs.
We will also continue to make better monitors ourselves and improve the research frontier for the blue team.
There are many low hanging fruits to pick for monitoring and not a lot of time. Our main blocker is having more hands on deck.
Come join us!
RS (Control): jobs.lever.co/apolloresearch…
AI Red Team Engineer: jobs.lever.co/apolloresearch…
AI Security & Control Researcher: jobs.lever.co/apolloresearch…
Every role is open in both offices, London and San Francisco. We provide visa sponsorship.
Our monitoring agenda: apolloresearch.ai/monitoring/a-s…
Auto-mode red-teaming: apolloresearch.ai/monitoring/pil…
All open roles: jobs.lever.co/apolloresearch
We're growing the Watcher team
Watcher is now deployed across many organizations and we’re seeing a lot of traction and inbound interest.
Many people are starting to feel how misaligned agents start affecting their organizations and Watcher is one of the few AI security products that is trying to tackle immediate and future misalignment failures like deception, overclaiming, oversight subversion or scope overreach in addition to security concerns.
If you’re interested in the intersection of AI safety and product, come join us!
- Engineering Manager (Product): jobs.lever.co/apolloresearch…
- Full-Stack Engineer (Product): jobs.lever.co/apolloresearch…
- Forward Deployed Engineer: jobs.lever.co/apolloresearch…
- Product Security Engineer: jobs.lever.co/apolloresearch…
Every role is open in both offices, London and San Francisco. We provide visa sponsorship. A formal AI safety background is not required.
Watcher: watcher.apolloresearch.ai
All open roles: jobs.lever.co/apolloresearch
We hosted a webinar with @Tailscale on safely scaling AI coding agents.
We cover how Aperture + Watcher provide visibility, runtime controls and behavioural monitoring, with a demo of Watcher blocking data exfiltration after a prompt injection.
Video: youtu.be/T3W6XPqJwe0?si…
We are hiring!
If you’re interested in this work, please apply. We have two roles open for exactly this kind of work:
AI security & control engineer: jobs.lever.co/apolloresearch…
AI control research scientist: jobs.lever.co/apolloresearch…
We’re committed to keep pushing the bar in monitoring capabilities across the industry in ensuring autonomous agents are safe.
Our own coding agent security product, Watcher, was instrumental in building the tools to red-team auto-mode.
watcher.apolloresearch.ai
Apollo Research ran its first external red-teaming campaign for Anthropic’s auto mode.
Now auto mode is the default permissions mode in Claude Code. After hardening based on Apollo’s findings, the classifier's miss rate fell from 12% to 7%. 🧵
127 Followers 158 FollowingAI/ML Engineer | LLMs + Classical ML 🤖
Building production-grade AI with proper SE
Freelancing | Studying | Building in public
59K Followers 18K Following#DigitalTransformation obsessed #CXO and #GrowthHacker. Find me at the intersection of #Tech and #Humanity. #AI #DataAnalytics #FinTech #CX #IoT #CyberSecurity