All views are my own and do not represent my employer.
In the wake of the recent wave of misalignment incidents, both OpenAI and Anthropic have reported slowing down RL training to improve safety. These incidents, combined with an apparent acceleration in the already-blistering pace of AI progress,1 have led a number of researchers and leaders in the industry to believe that the risk that humanity loses control of AI is now urgent enough to warrant slowing down the pace of AI development soon.
This has led to a lot of discussion about the role of third party evaluators in verifying “pacing commitments”, evaluating safety cases, or auditing compliance with safety policies. I think these are valuable roles for third party groups to aim to fulfill, but I also worry we’re putting the cart before the horse in all this talk of “verifying” and “auditing” things.
The science on loss-of-control risk is, to put it generously, nascent. Companies are not in the business of making structured, standardized claims about risk and safety that can be cleanly verified or falsified. There are no settled methods for measuring whether increasingly powerful AI systems might try to undermine human control or seize control entirely — companies report on various alignment benchmarks, but it is hard to tell whether their training process simply taught the models to game these benchmarks. It is hard to confidently bound risk even over a horizon of months because there is vast and hard-to-reduce uncertainty about whether recursive self-improvement will suddenly lead to AI systems capable of toppling states or overpowering humanity.
To the extent that AI companies choose to proceed with AI development, they can and should do their best to manage their own safety practices. They should continue to unilaterally slow as needed to do this, and there is room to do more in that vein. But to durably reduce loss-of-control risk to an acceptable level, we will probably need to develop shared technical standards for how to manage this risk adequately and enforce those standards uniformly across the industry, including internationally.2
We currently lack the basic prerequisites needed to have a conversation about safety standards: a shared understanding of the current state of loss-of-control risk, what companies are currently doing or not doing to try to manage it, and how well these efforts do or don’t work. The claims that AI companies make about these topics are currently far too high-level and imprecise to cleanly “verify” or “disprove.”3 Moreover, no company is even making a set of public safety claims that adequately cover all the important questions about loss-of-control risk.4
In this situation, I think what we most need is not verification of the safety claims companies are making, but production of a far greater quantity and quality of concrete evidence about loss-of-control risk and how companies are managing it.
AI companies themselves can and should be generating and publishing much more of this evidence, but third party investigations of major questions relevant to risk can also help. These investigations would look less like trying to prove or disprove claims in a structured “safety case”, and more like generating and operationalizing hypotheses relevant to risk and coming up with a wide range of creative ways to generate evidence about whether these hypotheses are true: in other words, like doing science.
And like all good science, I believe third party investigations should aim for evidence transparency: that is, publicly sharing as much of the empirical evidence underlying their conclusions as possible. External scientists will generally have to trust that third party evaluators like METR are not outright lying, but the goal is that insofar as possible they should not have to trust that we are making the right judgment calls on tough scientific questions — they should be able to read our reports and form their own views. Evidence transparency has three big benefits:
Evidence transparency creates common knowledge between companies: To start working toward serious standards, AI companies need to understand what other companies are doing in enough detail that they can actually replicate it on their own stack and learn whether it works for them. Because it is very tough socially and legally to privately share detailed information with competitors, the public channel is actually the main way competitors learn about each others’ work. I often talk to researchers at AI companies who pay attention to papers, system cards, and risk reports from other companies to learn how they approach practical challenges like alignment training or monitoring — but these public artifacts generally don’t get into the nitty gritty enough to be as helpful as they could be.
Evidence transparency empowers the outside world: In theory, we could set up a private cross-lab evidence sharing program to achieve 1 (though that’s been tried and it’s very difficult in practice). But public transparency also brings in scientists outside of any AI company, who currently struggle to participate effectively in conversations about standards because of the massive information asymmetry. External scientists will be from a different milieu and can bring a fresh perspective to many questions. Importantly, they also have very different incentives — they stand to be harmed by catastrophic loss of control, but don’t stand to gain nearly as much as company employees do from going faster. With outside voices in the room, the industry would likely arrive at standards that make very different tradeoffs between safety and speed.
Evidence transparency allows for evaluation of the evaluators: If evaluators just went into a company and came out with a high-level judgment about overall risk or the adequacy of practices, it would be very hard for anyone to tell if they conducted a good investigation and came to a reasonable conclusion. Standard scientific transparency norms force third party evaluators to prove themselves to the external scientific community — if the norm is that evaluators should carefully explain the evidence they relied on and the methods they used to analyze that evidence, it gets much easier to tell which evaluators are competent and hard-hitting, and much harder for AI companies to get away with engaging overly-friendly evaluators.
Given capacity constraints and the need to navigate redactions for IP, it will not be possible to achieve perfect evidence transparency for third party investigations. For example, in our Hugging Face report, we wrote a Methodology appendix but did not share our prompts (and certainly not our code). But evidence transparency was our north star, and I think we got pretty far — if you read the report carefully, your interpretation of what it means is about as good as mine.5
We need to admit that understanding and managing loss-of-control risk is an open scientific problem that no AI company or external research group has settled. Standard scientific norms of evidence transparency evolved to handle this situation, and we should lean into them. If AI companies and third party researchers make it a priority over the next several months, I believe we can dramatically improve the state of public scientific evidence about loss-of-control risk and mitigations. I think this would put industry leaders and policymakers in a much better position to implement a functional governance regime for loss-of-control risk.
Anthropic and OpenAI have reported that measures like the amount of code written by AI and number of agent-workdays per human workday have sped up a lot recently. No company systematically reports verified measures of outputs like how much algorithmic progress has sped up, and from public evidence it is not clear how the input measures companies report translate into the output measures that matter. Most researchers I know seem to have the sense that progress in AI capabilities has sped up recently, but they are unsure whether this is the start of a continuous acceleration.
The extent to which a set of standards reduces risk depends on how stringent they are, which is in turn ultimately a political question. For example, some researchers believe that alignment practices should not be considered "adequate" unless they are provably robust; trying to meet this bar would likely necessitate pausing development for many years or decades. The political process could settle on working standards that are far short of this while still raising the bar substantially on top of what any developer is currently achieving.
For example, it would take a lot of judgment and thought to operationalize a claim like “We don’t believe we’re overfitting to our methods of detecting misalignment,” and testing it could require poring over training data or running novel evaluations.
For one thing, I believe no company is making a crisp claim that they are confident they will not be able to build uncontrollable superintelligence within six months.
As stated in the redaction summary statement at the top of the report, METR and OpenAI were able to agree on how to describe redactions for all the important evidence underlying our views. This means that if an outside researcher carefully read the report and came to a different interpretation about how concerning the agents’ behavior was or what it means for the future, I wouldn’t feel like they should trust me because I had seen so much more private evidence that informed my view.

