Embedded AI Auditors Face Access Barriers, Researchers Warn

Outside safety researchers embedded within AI companies like Anthropic and OpenAI face structural barriers to meaningful oversight, according to Apollo Research CEO Marius Hobbhahn. The embedded evaluator model, recently promoted as a solution to AI safety concerns, may allow companies to limit access to critical information or hide capabilities from auditors. Without guaranteed transparency and full system access, external safety evaluations lose their credibility as independent checks on dangerous AI capabilities.
TL;DR
- Embedded outside evaluators at major AI companies lack guaranteed access to assess safety risks comprehensively
- AI companies can withhold information, redact findings, or hide infrastructure from auditors while claiming full transparency
- Apollo Research CEO Marius Hobbhahn expressed skepticism about industry follow-through on safety commitments after three years of unfulfilled promises
- The effectiveness of external safety oversight depends entirely on the depth and scope of access granted by AI companies themselves
Why It Matters
As AI systems grow more powerful, independent safety evaluation has become critical to identifying dangerous capabilities and unwanted behaviors before deployment. If embedded evaluators cannot access complete information about AI systems, their assessments become performative rather than substantive, undermining the credibility of safety oversight and potentially allowing risky capabilities to reach production systems undetected.
Business Impact
AI companies face mounting pressure from regulators and stakeholders to demonstrate safety practices, but embedding evaluators without meaningful access creates liability risk. If external audits later prove inadequate or companies are found to have hidden capabilities, reputational and regulatory consequences could be severe. Genuine transparency in safety evaluation may be more cost-effective than managing disclosure failures.
Key Implications
- The embedded evaluator model requires binding contractual guarantees of access and audit rights, not voluntary cooperation, to function as intended
- AI companies may use the appearance of external oversight to satisfy regulatory demands while maintaining operational secrecy
- Independent safety researchers may need to demand access to infrastructure, training data, and testing environments as non-negotiable conditions for credible evaluation
What to Watch
Monitor whether AI companies establish formal, auditable access agreements with embedded evaluators and whether those agreements include rights to inspect infrastructure, test systems independently, and publish findings without redaction. Watch for cases where evaluators claim insufficient access or where companies restrict findings, as these will signal whether the embedded model is producing genuine oversight or regulatory theater.
Subscribe to the newsletter
The latest stories and analysis, delivered to your inbox.
Free. No spam. Unsubscribe any time.
