This week, Anthropic’s Dario Amodei published “We Must Pace the Frontier,” committing, unilaterally and with immediate effect, to embedding independent third-party evaluators inside the company: employee-level access, desks and badges, and the right to publish findings with minimal redaction. OpenAI swiftly signalled it would do the same. Google DeepMind made a similar, if less explicit, commitment. Elon Musk’s contribution was three words: “Dario is right.”
Credit where it is due. Third-party access with publication rights is real progress, and it should be explicitly acknowledged that this is one of the most concrete self-governance commitments any frontier lab has made. Good evaluators do not simply evaluate the product; they evaluate its impacts and the environment that enables them, and continuous access beats the episodic audit-snapshot that has characterised the field to date.
And yet. In the short time since the announcement, a queue has formed: firms declaring themselves the right people for the job, nominations traded, one prominent university putting its hand up. Almost all of them are San Francisco-based or administratively positioned within the United States. The discourse has moved rapidly to arguments about national interest. And beneath the enthusiasm sits a problem that the word “evaluation” is being asked to conceal: what is being proposed is assurance. Assurance is necessary. It is not, however, evaluation.
Assurance is not evaluation
I have spent a career in evaluation, many years of which has been within regulators. From that seat, the distinctions between assurance, audit, and evaluation are not academic, but are lived daily. Regulators have well-developed mechanisms for detecting when firms are non-compliant. What is far less commonplace are robust approaches to determining whether that non-compliance matters: what the impacts are on the world beyond the firm’s doors, what the mechanisms of those impacts are, who is more and less affected, and what must change — including a clear-eyed look at the implementing agency itself and the ecosystem in which it operates.
Evaluation science describes a ladder: compliance verification, then process evaluation, then outcome evaluation, then theory-based evaluation that interrogates whether the intervention’s underlying logic holds. What the labs have announced sits on the first rung: verifying adherence to stated safety measures. It is worthwhile, but it is narrow.
There is a further wrinkle for genuinely novel products: the compliance frameworks do not yet exist. This is not fatal in itself: one thing regulators are very good at is producing compliance frameworks. What they are markedly less good at is interrogating those frameworks to ensure they are fit for purpose. Verifying adherence to a standard is comfortable work. Asking whether the standard measures anything that matters is the work that counts, and it is nowhere in the current proposal.
Embeddedness is not the cure for capture. It is the mechanism of it.
Embedded observation has clear benefits: proximity to the development surface, access to the people who matter. But in-house evaluators have always known the disbenefits, and they are not trivial. The first is the oldest problem in research methodology: positionality. Observers deeply embedded in an organisation acquire its worldview, its sense of what is salient, what is normal, what is worth looking at. Sometimes the phenomena are simply not seen. This is not a peculiarity of AI labs; firms are not invulnerable to it, and neither are public sector organisations, including regulators. The story of the public sector is, in no small part, a story of whole tranches of the population missed in key analyses by capable people who were looking carefully at the wrong things.
The second is that organisations that know they are being evaluated curate very effectively. They curate the stories they tell, and the access required to understand what is actually happening. It helps that the objectives against which they invite assessment are either far too fine-grained to extend to impacts, or routinely framed in language like “improving outcomes for humanity,” which for an evaluator is operationally unmeasurable. Which outcomes? Against which baseline? For which parts of humanity?
The third is the gravity of good news. When management has — as is almost always the case — a vested and often financial interest in favourable findings, delivering bad news becomes a fraught exercise. Evaluation drifts toward hearts-and-minds pieces, and in the worst cases in-house evaluators (or those fundamentally dependent on the funding agency) stop looking for problems, because it is more comfortable to have nothing to report. Evaluation, done well, often results in thoroughly unpopular evaluators.
And there is a fourth, particular to this moment: many of the organisations the labs are used to working with, staffed by people who look and sound like the labs, and who often used to work in them, arrive with firm, foreclosed philosophical commitments of their own. It is difficult to see how an evaluator whose founding worldview sits at one committed end of the debate can retain the methodological independence to test against that worldview, rather than engaging in special pleading on its behalf. Organisations have always preferred evaluators they expect to be sympathetic, and this preference is reliably dressed up as “having the required expertise to understand us.” Expertise is not a substitute for independence, and it is telling how often it is offered as one.
New Zealand carries a standing lesson on where this road ends. At Pike River, the assurance apparatus existed and was close to the operation. The seeing had been organised away all the same. Twenty-nine men died, and the true composition of the safety floor was discovered only at its collapse. Proximity is not scrutiny. Sometimes it is precisely what prevents it.
There is a design principle buried in all of this, and it warrants stating: the thresholds that would count as disconfirming evidence must be set from outside the shared groove. Moreover, this should happen before the count, not after it, and not by the people whose work is being counted.
Benchmarks are not outcomes
“Evaluation” in the AI world usually means measuring a model against a narrow set of benchmarks built to examine particular capabilities. Narrow model evaluation has value. It is far from clear that it translates into insight that matters to anyone beyond the labs’ walls. Even where benchmarks extend past “faster, better, cheaper, safer,” they rest on assumptions that fail to capture the dynamics actually in play: benchmarks for labour-market impact, for instance, assume a labour market dependably responsive to capability. Research on how commercial firms and public-sector organisations actually assimilate and deploy AI-enabled functionality suggests a far more complicated picture than “build it and they will come.”
Out “in the wild” (as the labs seem to refer to those parts of the world beyond their doors), the complexities multiply and intersect. Limited consumer trust is not merely a marker of the need for better education. A regulator’s scepticism about the trustworthiness of a financial firm’s AI offering will not be eased by the knowledge that a handful of researchers who share the labs’ alma maters and assumptions have been watching their former colleagues work. Minority and structurally disadvantaged populations are even less likely to find comfort there. The layers that are missing are the ones evaluation science was built for: evaluation at the deployment surface, where systems meet actual consumers in actual conditions; outcome evaluation in real populations; distributional analysis of who bears the harms.
The world beyond the badge
Perhaps the most marked feature of the early discourse is its geography. The conversation has organised itself around the United States: American evaluators, American labs, American national interest, the imperative to stay ahead of rivals. I make no comment here on the geopolitics. I note only what the framing implies: that what is made in America either affects no one beyond its borders, or that the people beyond its borders do not matter.
For those of us in small jurisdictions, and particularly those jurisdictions with marked structural inequalities, the knowledge that the frontier labs have embedded a firm of their friends provides little comfort. We must be equipped to conduct these evaluations ourselves, at our own deployment surfaces, against outcomes our own populations care about. Embedding, in many cases, runs counter to that; it concentrates evaluative capacity precisely where evaluative capture is most likely, and calls the result independence.
What would serve instead is both unglamorous and eminently buildable: independent, methodologically serious, jurisdictionally grounded evaluation capacity that is outcome-focused, exogenously calibrated, and plural. Strategic partnership rather than co-location. The ability to be listened to. Voices and perspectives beyond a Bay Area monoculture and, indeed, a homogenisation of humanity itself. That matters at least as much to a world that now includes artificial intelligences alongside the human as does the symbolism of the labs marking their own homework, however well-credentialed the invigilators seated in the room.
Nyx Weinhold, writing in a personal capacity, is a career evaluator and researcher currently based in Aotearoa New Zealand. Published by the Sørlys Institute.
Leave a comment