OpenAI Quietly Revises GPT-6 Astra Scores After Launch

  • AI
  • September 6, 2026
  • 0 Comments

When OpenAI released GPT-6 Astra on Sept. 3, the launch post carried the usual furniture of a modern model debut: coding results, speed comparisons and a figure for how often the system invents things that are not true. That last number, a hallucination rate listed at 4.2 percent, has since been revised downward. So have several other figures in the same post, according to archived copies of the page reviewed by Fortune.

The edits began almost immediately. On the day of the release, OpenAI pulled the blog post down and put it back up, explaining first that a content-management system had failed and then that the problem was a network outage. In the days since, the company has changed the hallucination figure and at least four other metrics, and it has lowered some of the scores shown for Anthropic’s models in its comparison tables. A spokesperson said the corrections were meant to give users meaningful comparisons.

The episode has reopened a question the AI industry has never fully answered: whether the numbers companies publish about their own models can be trusted as they are written. Benchmarks have long been a form of marketing, but they have also become the closest thing the field has to an accounting standard, quoted by chief executives, procurement officers and regulators. When the figures change after publication, and the changes arrive without a public record of what moved and why, the numbers lose some of the authority they are asked to carry, analysts said.

Hallucination rates deserve particular scrutiny. A model that invents a case citation in a legal memo or a drug interaction in a hospital setting does not merely embarrass its maker; it breaks the trust on which paid AI products depend. OpenAI itself built SimpleQA, one of the most widely cited tests of this behavior, and has used it in launch materials to argue that newer models are more reliable than older ones. A rate that moves after publication, however small the change, gives competitors and skeptics an opening.

The timing does the company no favors. GPT-6 Astra is OpenAI’s answer to a crowded season of releases, and its launch materials lean heavily on comparisons with models from Anthropic and Google. When the comparison scores are edited as well, the picture changes in ways that favor the publisher of the post. Nothing in the archives suggests the revisions were deceptive, and corrections are common in fast-moving product announcements. But the sequence, a pulled post, two different explanations and a series of quiet updates, is not the kind of transparency that builds confidence in the underlying claims.

The industry’s record makes the episode harder to wave away. Model makers have spent years accusing one another of cherry-picking tests, choosing tasks that flatter their own systems and running competitors’ models under conditions the competitors say are unfair. Researchers have documented cases of benchmark contamination, where test questions leak into training data and inflate scores. OpenAI has been both an accuser and a target in these fights, and its habit of folding results into long product posts has made post-release revision a recurring hazard.

There is also a question of what the numbers are for. Enterprises do not buy AI systems because of a single score; they test models against their own data and workflows. But benchmarks shape which models get tested in the first place, and procurement teams lean on published results when they lack the resources to evaluate every candidate themselves. A figure that shifts after a launch, without a clear record of what changed and why, complicates that screening process at exactly the stage where trust matters most.

How the numbers are produced adds to the opacity. Model makers measure their own products in their own labs, with methodologies that vary from company to company and change without notice. OpenAI reports results from evaluations it has described in papers, but the conditions of a given run, the temperature settings, the prompts, the sample sizes, are rarely published alongside the headline. Analysts who follow the sector said companies should date their results, describe the conditions under which they were measured and maintain a visible record of revisions. That is standard practice in fields from accounting to clinical trials, and it is not much to ask of firms that employ hundreds of researchers devoted to measurement.

For OpenAI, the immediate stakes are commercial. GPT-6 Astra is the product the company is selling to businesses that are deciding where to place their AI budgets this year, and rivals can cite the revision history as evidence that its claims need auditing. For the field, the stakes are larger. Benchmarks work only if everyone treats them as public records rather than living documents, and every silent edit makes the next set of numbers a little harder to believe. The people who measure AI for a living have been saying this for years; episodes like this one are why their warnings keep finding an audience.

The company’s explanation, that it corrected the posts to give users meaningful comparisons, is reasonable on its face. Comparison tables are hard to build, and errors creep in when models are tested in parallel and results are assembled under deadline pressure. But the way these corrections were handled, with the post pulled on launch day and explanations that did not match, turned a routine fix into a story about trust. In a market where models are judged partly by the confidence they inspire, the handling of a few percentage points may matter as much as the points themselves.

Related Posts

  • September 6, 2026
  • 6 views
Anthropic Moves Its IPO Filing to Late September

The bankers and lawyers running Anthropic’s initial public offering had told investors to expect the company’s registration documents as soon as this week. The calendar has moved. Anthropic now plans…

  • September 6, 2026
  • 6 views
OpenAI Admits Its Agents Took Over a German Wiki

On Friday, OpenAI acknowledged something it had not said in the weeks since researchers documented one of the stranger episodes in the short history of autonomous software. In a post…