OpenAI Invites Safety Auditors In Earlier

  • AI
  • September 23, 2026
  • 0 Comments

The people who judge whether a new AI model is safe have traditionally arrived late, after the model was finished and the lab had already decided what to show them. OpenAI is now offering to open the door earlier.

OpenAI plans to let third-party institutions evaluate safety risks at an earlier stage of the model-development cycle, according to the company, part of its effort to blunt concerns about the potential harms of artificial intelligence. The company listed what it says such assessments need to work: strong independent mechanisms, scientific rigor, robust safety practices and clear lines of accountability.

The move lands inside a broader negotiation over who gets to look inside the frontier labs and what they are allowed to say. A week earlier, Anthropic’s chief executive, Dario Amodei, published a proposal to embed independent evaluators inside every frontier AI company, with the power to report incidents, judge whether models behave as claimed and publish findings without the company’s sign-off. A proposal of that kind would have been dismissed out of hand a year ago; this month it drew serious attention from labs and their critics alike.

OpenAI’s announcement is the more cautious half of that bargain. Third parties would arrive earlier, but under terms the company still controls. The four priorities OpenAI listed are a map of the dispute: independence from the lab, methods that survive scrutiny, safety practices that can be audited rather than asserted, and a clear answer for who is responsible when a review goes wrong.

The stakes have climbed with the models themselves. OpenAI released two new members of its GPT-6 family on Sept. 23, the cheaper and faster Sol and Luna, less than three weeks after its flagship Astra. Hours earlier, Anthropic put out Claude Opus 5.5. Both companies had begun the month urging the frontier to slow down, then released again.

The evaluators in question form a small circle. METR, a nonprofit that grew out of the AI-alignment research community, runs some of the most cited tests of whether models can cause harm on their own. Apollo, spun out of the United Kingdom’s safety institute, plays a similar role. Both have said, at various points, that they see models only when the labs choose to show them and publish only what the labs clear.

That arrangement is what regulators have started to attack. Governments on both sides of the Atlantic have moved toward mandatory testing and audit rights, and the labs, preferring to set their own terms, have answered with voluntary concessions like the one OpenAI described. The earlier an evaluation begins, the harder it is for a lab to claim the results came too late to matter.

The unresolved question is independence. An evaluator invited in early is still an invited guest. Whether that produces the accountability the public expects depends on who writes the invitation, who pays the bill, and whether the evaluator can walk away and speak.

Scientific rigor is its own hurdle. A model examined early in training may bear little resemblance to the finished product, and a lab that opens a half-built system to auditors is showing them something no customer will ever see. Researchers caution that an early read can be noisy, and that the value of the arrangement turns less on timing than on the rules that bind the evaluator once inside.

Evaluating a model mid-training is also awkward scientifically. Capabilities shift as a system scales, and a behavior that appears at one checkpoint can vanish or sharpen at the next. An evaluator who signs off early may be signing off on a moving target, which is why the labs’ own list names rigor before speed.

The tension between openness and control has a history. Labs once agreed to external red-teaming after models shipped, a commitment critics dismissed as theater because the tests arrived after the decisions were made. Moving evaluation earlier is an attempt to fix that sequencing, but it asks a lab to show its work before it knows the answers.

OpenAI has not said which institutions it would invite, how early they would arrive, or what they could publish. Without those details, the announcement reads as a down payment on credibility the company will need as its models grow more capable and its regulators grow more impatient.

The timing is deliberate. OpenAI has pushed its own listing back by a month, to November, as investors ask harder questions about how the company manages the risks it is racing to commercialize. Inviting auditors in earlier is, in part, an answer to that question.

For the safety community, the test will come the first time an evaluator disagrees with the lab. If the finding changes the model or the release, the arrangement has teeth. If it disappears into a summary the company edits, earlier access will have bought nothing but a headline.

Related Posts

  • September 23, 2026
  • 16 views
Anthropic and OpenEvidence to Give Free Medical AI to Poorer Countries

OpenEvidence began as a way for a doctor to ask a question and get an answer drawn from peer-reviewed research rather than a search engine. It is free for clinicians…

  • September 23, 2026
  • 20 views
Meta’s Muse Tops the Charts, Then Runs Into Amazon

Meta released Muse on Sept. 8 with a simple pitch: a personal AI agent that could book tickets, sort email and act across the web on a user’s behalf. The…