OpenAI quietly revises GPT-6 Astra benchmark numbers after launch
OpenAI has repeatedly altered the evaluation numbers published alongside its September 3 announcement of GPT-6 Astra, quietly revising several benchmark figures in the hours and days after the blog post went live, according to a Fortune review of archived snapshots of the announcement page.
What's new
Fortune's comparison of successive cached versions of OpenAI's Astra announcement found that at least five metrics changed after publication. In one instance, Astra's reported hallucination rate was halved to 2% in a snapshot captured around 5:20 p.m. ET, roughly three hours after the post first went live — but OpenAI continued revising the figure afterward, and by the time of the review the hallucination rates had reverted to the original published values of 4.2% and 12.2%. A separate internal cybersecurity evaluation score for GPT-5.6 Sol, run on OpenAI's own version of the ExploitBench benchmark, was also revised upward, moving from 5.5% in the first published version to 11.5% in later ones. Some changed figures made Astra look stronger; others lowered the reported scores of rival Anthropic's models on the same comparison charts.
The metric changes coincided with an unusually rocky rollout: OpenAI had originally scheduled the announcement for 2 p.m. ET, but the blog post did not become reliably viewable for almost two hours afterward, and a link posted from OpenAI's official account around 3:32 p.m. failed to load. CEO Sam Altman posted the link himself roughly 20 minutes later, acknowledging the company had "hit a snag" deploying the post.
Context
GPT-6 Astra is OpenAI's flagship release, announced September 3 as the company's most capable model to date and already published on ModelDex as a general-availability launch. Benchmark tables are a standard part of any frontier-model announcement, used by developers and enterprises to compare a new release against competitors before committing engineering time to adopt it. Because those numbers typically aren't re-run and re-verified independently before a launch, they carry an assumption of stability once published.
Why it matters
Quietly editing benchmark numbers after a model announcement is already live undercuts the reliability of the comparison data developers use to make adoption decisions, and doing so without a visible changelog or correction notice raises questions about OpenAI's internal review process for the metrics it publishes. The specific pattern here — some numbers moving in Astra's favor while a competitor's figures moved unfavorably on the same chart — is likely to intensify scrutiny of self-reported benchmarks industry-wide, adding to existing concerns about "benchmaxxing," where labs tune evaluation conditions to maximize headline scores rather than reporting a single, stable, independently reproducible result.
Corroborating sources
- Fortune
https://fortune.com/2026/09/04/openai-quietly-boosts-some-of-astras-evaluation-metrics-amid-rare-delay-in-publication-of-the-modeblog-post-announcement/
“It was halved down to 2% for Astra.”