Independent AI Safety Evaluators Need Real Power Now

12 min read
0 views
Sep 18, 2026

More than a hundred specialists just drew a hard line: inspecting frontier AI is useless if labs still control the evidence, the money, and the story. The missing piece is not another pledge.

Financial market analysis from 18/09/2026. Market conditions may have changed since publication.

Have you noticed how often the same sentence keeps coming back? Frontier labs say they welcome outside eyes. Then the fine print arrives, and those eyes are suddenly looking through frosted glass. That gap, between the pledge and the actual inspection, is the story sitting under this week’s public letter from more than a hundred specialists who test, audit, and stress the most capable systems on earth.

Why Independent Safety Checks Suddenly Feel Urgent

I keep coming back to a simple discomfort. If a handful of companies control models that can touch cybersecurity, markets, critical infrastructure, and national security systems, the public cannot live on those companies’ own scorecards. That is not cynicism. That is basic risk hygiene. When the same organization trains the model, markets the model, and then grades the model, you do not have an exam. You have a brochure.

The letter published this week is not a manifesto against building powerful systems. It is closer to a working brief. The signatories say they are encouraged that some labs now talk about embedding third-party teams. They also say encouragement is cheap. Credibility is not. Embedded evaluation only works if the people doing it can keep scientific objectivity, publish what they find, stay independent, and survive the moment a finding becomes inconvenient.

That last part is the one people skip. Independence is not a vibe. It is a stack of boring protections: who pays, who edits, who can talk to the board, who can walk the building, and who still has a job after saying the quiet part out loud.

What The Letter Actually Asks For

The ask is sharper than “please let outsiders poke the chatbot.” The group wants labs to embed evaluators who can look at the systems themselves, at significant incidents of real-world harm, and at the company’s training, deployment, oversight, operations, and safeguards. In other words, not just the demo. The factory floor.

To be credible, embedded third-party evaluations must have scientific objectivity, transparency, independence, and robust protections against interference from the evaluated companies.

That sentence is the whole argument. Everything else is implementation. And implementation is where good intentions go to die.

The minimum conditions they list are not exotic. They look like the rules any serious audit practice already treats as table stakes. No ownership or governance by the lab being inspected. No fat commercial side deals with that lab. No payment that depends on a flattering result. Full editorial control. Disclosure of conflicts. Multiple teams, not one favorite vendor. Room for disagreement between evaluators and between evaluators and staff. Limited nondisclosure. A path to the board. A path to the public. Protection against retaliatory lawsuits. Funding that does not vanish the week a report turns ugly. Access that looks like the access of highly privileged internal staff, minus customer secrets that truly belong to someone else.

Read that list slowly. It is not anti-company. It is anti-theater.

Employee-Like Access Sounds Generous Until You Ask What It Means

One lab chief recently floated the idea of giving some evaluators employee-like access. Other prominent industry names said they liked the direction. Fair enough. The phrase has a nice ring. It also leaves every hard question on the table.

Which groups get chosen? Who decides they are “sufficiently technically credible”? How long do they stay? Can they talk to engineers without a minder in the room? Can they see unreleased internal systems, not only the public model card? Can they inspect the messy incident logs, the near misses, the eval harnesses that failed, the safety cases that got rewritten after a board deck? Can they walk the same physical spaces as senior internal risk staff?

I’ve found that the moment you start asking those questions, the conversation gets quieter. That is usually a sign you are getting closer to the real issue.

Access equivalent to highly privileged employees is a big request. It would mean company machines, candid one-on-one conversations, sensitive internal data, and systems that never ship. The people organizing the letter argue that this is exactly the point. Public releases are only part of the risk surface. Internal tools can already be potent. If an unreleased system can be used in a serious incident, pretending the public model is the only object of interest is a convenient fiction.

Independence Is A Structure, Not A Compliment

People love calling a vendor “independent” because the logo is different. That is not independence. If the lab owns the evaluator, sits on its board, feeds it most of its revenue, or can yank the contract after a bruising finding, the report will lean. Maybe not every time. Often enough to matter.

  • No ownership or governance by the frontier lab under review
  • No significant side business that turns the auditor into a sales partner
  • No success fees, bonuses, or future work tied to a soft conclusion
  • Editorial control that stays with the evaluator, not the communications team
  • Open admission of remaining conflicts, plus a plan to mitigate them

Those rules sound strict because they are supposed to be strict. In every other high-stakes field, we already know what happens when the inspected party writes the inspector’s paycheck and then edits the inspector’s draft. The language gets smoother. The caveats get longer. The headline gets friendlier. The underlying hazard does not move an inch.

Perhaps the most interesting part of the letter is how ordinary this demand is. It is not asking labs to halt research. It is asking them to stop treating evaluation as a controlled demonstration.

Why One Favorite Evaluator Is Not Enough

A single embedded team can become house culture in six months. They learn the acronyms. They like the people. They start anticipating what the lab will accept. That is human. It is also how capture works.

The signatories want several organizations, covering different risk areas, with real technical depth in each. They also want those groups to be allowed, even encouraged, to say where they disagree with one another and with company staff. Disagreement is not a branding problem. It is evidence that someone is still thinking.

If every evaluator produces the same warm paragraph, I get suspicious. Uniformity is easy to manufacture. Divergence is harder to fake.

Transparency With A Timer, Not A Gag Order

Nobody serious is arguing that evaluators should dump customer data, private keys, or exploit details into a blog post at midnight. The letter leaves room for time-limited redaction when intellectual property, customer information, individual privacy, security, or public safety is truly on the line. That is reasonable.

What is not reasonable is an NDA so wide it swallows methods, access terms, and the existence of a dispute. If the public cannot know what was tested, how it was tested, what was off limits, and where the evaluator and the lab parted ways, then “third party evaluation” is just a press line.

The group also wants prompt, unfiltered communication with boards and other privileged oversight bodies. That matters more than people admit. A report that dies in a middle-manager review cycle is not oversight. It is internal mail.

Retaliation Is The Quiet Killer Of Honest Work

Here is the part I wish more coverage would sit with. Evaluators can have perfect methods and still fold if the lab can sue them into silence, freeze their funding, or blackball them from the tiny market of groups that can actually do this work. There are not that many teams with the technical depth and the scale. Everyone in the room knows that. That scarcity is leverage.

So the letter asks for reasonable shields against retaliatory litigation and for funding structures that keep the lights on after an unflattering conclusion. Without those, “speak freely” is a slogan. People speak freely when the downside is survivable.

Embedded evaluators should be shielded from retaliation from the companies they embed with for choosing reasonable evaluation methods, discovering information, or drawing conclusions that are unflattering to those companies.

In my experience, this is where polished principles usually get watered down. Access is photogenic. Legal cover is not. Funding independence is even less photogenic. Those are still the load-bearing walls.

Internal Safety Teams Are Not The Enemy

It would be sloppy to frame this as outside saints versus inside villains. The organizers themselves say third-party work is not a substitute for internal evaluation or mitigation. Labs should keep building their own red teams, their own incident processes, their own deployment brakes. Outside evaluators are a complement. They are a check on self-report, not a replacement for engineering discipline.

That distinction is useful. If a company treats an embedded auditor as the entire safety program, it is outsourcing responsibility. If it treats the auditor as a nuisance to be managed, it is performing. The adult version is both: strong internal work and outsiders who can say the internal work is incomplete.

The National Security Argument Is Not Abstract

One signer put the stakes in blunt terms. When a few powerful labs control capabilities that can endanger cybersecurity, critical infrastructure, and the systems national security and the economy run on, government and the public cannot depend only on those labs’ account of what is safe. That is not a culture-war sentence. It is an institutional one.

You do not need to believe in cinematic catastrophe to accept a narrower claim: concentrated capability plus self-graded assurance is a fragile arrangement. Markets already understand this in banking, aviation, nuclear operations, and medicines. We argue endlessly about the right regulator. We do not usually argue that the factory should be the only source of truth.

There is a political overlay, of course. Some industry voices want statutory rules. Others want the state to stay out and let companies design their own guardrails. The coalition behind the letter says it is not wedded to one delivery mechanism. It wants basic principles and more standardization for people who work outside the labs. That is a narrower, more practical target than the usual policy food fight.

A Standard Exists. Adoption Is The Hard Part

The letter points to an early standard already seeing some uptake, then immediately admits the obvious: far more work is needed before embedded evaluation is effective and meaningful. Standards do not enforce themselves. A pdf on a website does not change incentive structures.

Still, naming minimum conditions is how you stop every lab from inventing a private definition of “independent.” If Company A means a weekend tabletop with a friendly vendor, and Company B means months of privileged access with publishable findings, the public will hear the same word and miss the gulf.

Minimum stack for credible embedding:
  Independence of ownership and money
  Editorial control and conflict disclosure
  Multiple specialist teams
  Limited NDAs and a path to the board
  Anti-retaliation and durable funding
  Privileged-employee-level access for the evaluation itself

That is the checklist I would tape to the wall before cheering any new partnership announcement.

What Labs Gain If They Take This Seriously

There is a self-interest case here, and it is not subtle. Credibility compounds. If the only people who ever see the dangerous edges are on payroll, outsiders will assume the worst the first time something ugly leaks. If a lab can show that outsiders with real access and real protection reached a documented view, the company has a better story when the next scare hits. Not a perfect story. A better one.

There is also an internal benefit. Outside teams catch blind spots that house culture stops seeing. They ask the dumb question that is not actually dumb. They notice when an eval suite has been quietly tuned to the model’s strengths. They notice when “we mitigated it” means “we wrote a policy.” Good engineers already know this. They live with production surprises every week.

The organizers warn that labs might ignore the letter. They also note that reputation is on the line. That feels right. After a season of grand statements about third-party testing, declining the minimum conditions would tell on itself.

The Practical Mess Nobody Wants To Schedule

Let us be honest about the operational pain. Privileged access creates insider-risk problems. Evaluators will see unreleased weights, tools, and roadmaps. Someone has to clear them, log their sessions, and decide what leaves the building. Legal teams will hate the publishing clock. Security teams will hate extra humans on sensitive machines. Product teams will hate delays. Communications teams will hate findings they cannot massage.

All of that is real. None of it is a reason to keep evaluation decorative. It is a reason to design the program like a controlled inspection, not like a press tour.

  1. Define the risk areas first, then pick specialist teams for each area.
  2. Write access rights that match senior internal risk staff, with narrow exceptions for third-party customer data.
  3. Set a redaction clock in advance so publication is a process, not a negotiation after the bruise.
  4. Create a funding lock that cannot be quietly canceled after a harsh draft.
  5. Give evaluators a documented channel to the board that does not pass through brand management.

If a lab cannot do those five things, it is not embedding evaluators. It is hosting visitors.

Public Research Access Still Matters

The letter is careful on another point that often gets lost. Embedded evaluation cannot cover every oversight need. It should sit beside broader external research access and more public transparency. A small set of privileged inspectors is necessary. It is not sufficient. Independent researchers outside the embed program still need ways to probe systems without becoming unpaid extensions of a corporate comms calendar.

That balance is tricky. Open access can leak capabilities. Closed access can hide failures. The adult move is layered oversight: internal teams, embedded specialists, and a wider research perimeter with clearer rules. Layered does not mean chaotic. It means no single channel owns the truth.

How To Read The Next Wave Of Announcements

Watch the verbs. “Engage,” “partner,” “consult,” and “welcome feedback” are soft. “Embed,” “access equivalent to senior internal staff,” “publish after time-limited redaction,” and “funding independent of findings” are hard. Soft verbs make good keynotes. Hard verbs change the risk picture.

Claim you will hearQuestion to askWhy it matters
We support third-party testingWho pays, and can payment stop after a harsh report?Money is the first lever of capture
Evaluators will have deep accessIs that access equal to privileged internal risk staff?Shallow access misses internal systems
Findings will be sharedShared with whom, after how long, with what redactions?A private memo is not public accountability
We chose a trusted partnerHow many partners, and may they disagree in public?One house vendor is easy to manage

Keep that table nearby. The industry is about to produce a lot of language that sounds like compliance while leaving the machinery untouched.


A Personal Read On Why This Moment Feels Different

I do not think the letter is going to rearrange the industry by Monday. Letters rarely do. What it can do is raise the cost of vague promises. Once minimum conditions are written in plain language, “we already have evaluators” stops being an automatic win. People can ask whether those evaluators can publish, whether they can talk to staff alone, whether they can survive a fight.

There is also a cultural tell. The evaluation community is small, technically picky, and usually allergic to pile-on politics. When more than a hundred of those people sign the same short list of conditions, you should assume the current setup is not good enough. They would not burn social capital for a vibes campaign.

Will labs meet them halfway? Some might, at least on paper. The test is the first ugly finding. If the report comes out with methods intact, access described, disagreements visible, and the evaluators still funded, then the pledge was real. If the report is late, thin, and lawyered into fog, we will know.

That is the standard I plan to use. Not the announcement. The first uncomfortable document.

What Readers Should Demand Without Becoming Policy Hobbyists

You do not need a research lab badge to follow this. You need a short memory and a few stubborn questions. When a company says outsiders are in the building, ask whether those outsiders can leave with an honest account. When a company says safety is the top priority, ask whether the people measuring safety can be punished for measuring it too well.

It is fashionable to treat AI risk as either science fiction or a branding exercise. The letter sits in the unfashionable middle. Capabilities are rising. Evaluation capacity is thin. Incentives are lopsided. The fix is not a slogan and it is not a panic. It is boring institutional design: access, money, protection, publication.

If that sounds unglamorous, good. Safety work that needs glamour is usually performing for an audience. The useful version is closer to an audit team that can open the cabinet, read the ugly folder, and still have a career afterward.

So here is the line I would draw. Frontier labs can keep building. They can keep competing. They can keep giving speeches about responsibility. What they should not get is the last word on whether the systems they control are as contained as they say. Independent evaluators are only independent if the structure makes honesty cheaper than silence. Until that structure exists, every new pledge is just another polished sentence waiting for a footnote.

Cryptocurrency is the future, and it's a new form of payment that will allow more people to participate in the economy than ever before.
— Will.i.am
Author

Steven Soarez passionately shares his financial expertise to help everyone better understand and master investing. Contact us for collaboration opportunities or sponsored article inquiries.

Related Articles

?>