Don’t Let AI Audits Become a Box-Ticking Exercise

Summary
- Aligning on audits: US lawmakers and AI companies are converging on the need for independent evaluations of frontier AI models.
- Shades of gray: Current laws don’t specify whether evaluators will get access to the full model weights and training data, or only to prompt inputs and outputs.
- Safety-washing: Without proper access to the most dangerous models, independent evaluations will become window-dressing for AI company self-assessments.
- Closing the gap: New laws should give evaluators deeper access to riskier models, enough time to test them, and a standard way to report their findings.
In mid-July, one of OpenAI’s models autonomously escaped its sandbox and hacked into US tech company Hugging Face. Reportedly, OpenAI was not monitoring its model in real time and only discovered the attack after Hugging Face had already reported it to law enforcement.
In June 2026, the US Department of Commerce forced Anthropic and pressured OpenAI to restrict or delay access to their more powerful AI models.
What these events have in common is that none were informed by a systematic and legally-binding AI evaluation framework.
Over the past two months, lawmakers and every major frontier lab have converged on the need for independent evaluation. But one unresolved question is what evaluators actually get to see.
The legislative landscape
California's SB 53 and New York's RAISE Act require frontier developers to disclose “the extent to which third-party evaluators were involved” in catastrophic risk assessments.
Illinois' SB 315 goes further: it is the first US state law to mandate annual independent third-party audits of frontier developers' safety practices, which both OpenAI and Anthropic supported during the legislative process.
At the federal level, the bipartisan FRONTIER Act, introduced on 23 July by Representatives Jay Obernolte and Lori Trahan, would create a system of licensed Independent Verification Organizations (IVOs), of the kind long championed by AI governance nonprofit Fathom. The IVOs would be required to audit very large frontier developers at least every six months. The bill says IVOs must receive “timely access upon request to unredacted materials, records, personnel, systems, and all other information reasonably necessary for conducting the assessments and ongoing monitoring”.
The term ‘reasonably necessary’ is underspecified. Will frontier developers be required to provide access to model weights, or not?
Notably, the bill would also preempt state legislation on third-party auditing of frontier AI companies – including state laws requiring a developer to provide “model weights or training data”. This makes it especially important for the federal bill to specify the right level of access.
Black-box access is not enough
In frontier AI evaluations, ‘black-box’ and ‘white-box’ refer to opposite ends of a spectrum of AI model access. Black-box means seeing only the inputs and outputs when prompting the model. White-box typically means access to the model weights and details of how it was trained.
In 2024, Hubinger and colleagues at Anthropic demonstrated that language models could be trained as ‘sleeper agents’: models that behave well during evaluation and switch to a different objective only when a trigger condition is met. Crucially, these backdoors persisted through standard safety training, while the model passed every behavioral test.
Subsequent work by MacDiarmid and colleagues showed that simple investigation techniques applied to the model's internal activations could almost perfectly predict when a model would defect.
When the research nonprofit METR evaluated OpenAI's o1-preview, it received API access days before its evaluation concluded. METR reported that the short window, high latency, and low rate limits left substantial room for better elicitation – meaning their published results likely understated the model's capabilities.
These cases show that black-box access can miss dangerous behavior; white-box access is more likely to find it; and evaluators need sufficient time to do their jobs.
The emerging consensus and the remaining disagreement
In June 2026, Anthropic published its Advanced AI Framework, which states that “self-assessment is not enough” and calls for qualified independent evaluators supported by standards, licensing, and pooled funding. In recent weeks, both OpenAI and Google DeepMind have made similar proposals for federal third-party evaluations.
But it’s not only labs calling for more independent testing. Virginia directed a formal study of the IVO framework in April, and Connecticut authorized a voluntary IVO pilot in May.
However, no proposal says how deep these evaluations should run. And new laws cannot inherit deeper access as a default, because there has never been one.
Without legal specificity, ‘reasonably necessary’ will resolve to what evaluators have historically received: API access on short notice, with no view of the system's internals.
The strongest argument against deeper access is that it creates intellectual property and security risks.
This is a real problem, but also one that’s increasingly tractable. As one of us (Tlaie) outlined in a Pour Demain paper, confidential computing could allow third parties to run evaluations on decrypted AI model weights inside a secure digital environment. In other words, they get to see the results of their investigations, but not the model itself. The developer retains hardware control; the evaluator obtains white-box access.
What the laws should require
Three concrete additions would close the gap between current law and credible oversight:
- Access levels should be calibrated to model risk. Evaluators should have black-box access for low-risk evaluations; gray-box access (log-probabilities, training metadata) as the absolute minimum for models that pose systemic risks; and access approaching white-box level (weights, gradients, fine-tuning) for the highest-capability systems.
- Minimum evaluation windows should be measured in weeks, not days. Evaluators should also receive version-stable model snapshots for longitudinal study. Existing state laws and the FRONTIER Act set the frequency of audits, but specify no floor on evaluation time.
- Evaluation reports should be standardized. Each report should document evaluators’ access level, time constraints, methodology, conflicts of interest, and limitations. To its credit, the FRONTIER Act requires that access limitations be described in the evaluators’ assessment report.
What this is really about
Frontier AI safety claims are empirical claims about complex systems. AI is not unique: other high-stakes industries have already worked out that verifying such claims requires access to the system itself, rather than summaries of it.
For example, Food and Drug Administration reviewers do not merely read a sponsor's conclusions about a drug trial; they also analyze the raw data. In aviation, regulators gain access to design and engineering data, and certification flight tests are flown with the authority's own test pilots aboard.
Until the laws specify what evaluators get to see, the gap will be filled either by developer self-assessment dressed in third-party clothing, or by improvised government intervention with no transparent standards. Neither is acceptable.




