5 Sep 2026, Sat

OpenAI Faces Scrutiny Over Shifting GPT-6 Astra Benchmarks Amid Chaotic Launch.

OpenAI has found itself at the center of a fresh controversy, having made multiple, sometimes significant, adjustments to the evaluation benchmarks for its highly anticipated GPT-6 Astra model since the initial, error-plagued publication of its announcement blog post mid-afternoon on September 3rd. These revisions, first brought to light by detailed forensic analysis of archived versions of the post, have shown Astra’s performance figures improving in several key metrics, while, conspicuously, numbers for models developed by its fierce competitor, Anthropic, have deteriorated. The unusual circumstances surrounding the blog’s release, coupled with the subsequent alterations, have ignited a debate within the AI community regarding transparency, integrity, and the pervasive practice of "benchmaxxing" in the fiercely competitive large language model (LLM) arena.

The saga began with an unusually rocky rollout of the official blog post announcing GPT-6 Astra. OpenAI had originally scheduled the post to go live at 2 p.m. ET, a standard time for major tech announcements. However, the release was anything but smooth, with nearly two hours elapsing before the content became widely accessible online. The initial attempt to share the news via OpenAI’s official X (formerly Twitter) account at 3:32 p.m. was met with frustration, as the provided link consistently returned an error message, rendering the blog post unviewable.

The situation escalated when OpenAI CEO Sam Altman, at 3:50 p.m., personally intervened, posting the link himself with the candid acknowledgment, "We hit a little snag getting the blog post deployed, but it is really great." Despite Altman’s direct intervention, numerous users, including journalists from Fortune, continued to report the same frustrating error message. It wasn’t until approximately an hour later that the blog post finally loaded properly for the general public, after a period of intense public confusion and speculation.

Further investigation into the incident revealed that OpenAI had, in fact, briefly published the blog post shortly after 2 p.m. as planned, only to retract it almost immediately. The company offered shifting explanations for this swift pull-back, initially citing a bug in its content management system, then an internet outage, and finally stating that the reasons were undisclosed but "unrelated to the benchmark performance figures." However, upon the post’s eventual republication, it quickly became apparent that several evaluation metrics had been altered. What’s more, the numbers continued to fluctuate even after the post was widely viewable, adding layers of complexity and suspicion to the entire launch.

This revelation of changing benchmarks comes at a pivotal moment in the AI industry, characterized by an unrelenting, frenetic pace of innovation and intense competition among leading developers. Companies are locked in a relentless race to release updated large language models, each striving to demonstrate superior capabilities and capture market share. In this high-stakes environment, benchmark metrics serve as critical, albeit often debated, indicators of progress and performance. The shifting figures for GPT-6 Astra and its rivals underscore the inherent challenges in objectively measuring the nuanced performance of LLMs using standardized tests, raising concerns about the potential for manipulation and strategic gamesmanship within the industry.

An OpenAI spokesperson, responding to inquiries from Fortune, emphasized the company’s commitment to accuracy: "We care deeply about getting evaluations right. Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting. For our launch blog, we made fixes to ensure the numbers represent our best estimate of available model performance, so that users can make meaningful comparisons." While this statement aims to reassure, the extent and nature of the changes have fueled skepticism among some observers.

Discrepancies Between Drafts and the Evolving Live Publication

A closer examination of the various archived snapshots of the blog post reveals a compelling narrative of evolving metrics. Among the most striking alterations was Astra’s reported hallucination rate – a critical metric reflecting an AI’s tendency to generate factually incorrect or nonsensical information. In the earliest internet archive snapshot of the blog post, captured at 2:23 p.m. ET, Astra’s hallucination rate was stated as 4.2%. This figure remained consistent across several subsequent snapshots, including one taken at 3:11 p.m. ET, roughly ten minutes before OpenAI finally tweeted out the ‘final’ version of the post.

However, a dramatic change occurred by the sixth archival snapshot, recorded at 5:20 p.m. ET, after the blog had become widely accessible. Astra’s hallucination rate was suddenly halved, plummeting to an impressive 2%. Scores for Astra’s predecessor, GPT-5.6 Sol, also saw a reduction, from 12.2% to 9.4%. The volatility didn’t end there; as of the latest checks, these hallucination rates have mysteriously reverted to their original figures of 4.2% for Astra and 12.2% for Sol, creating a perplexing full circle that undermines confidence in the stability of the reported data.

Another notable adjustment involved GPT-5.6 Sol’s performance on OpenAI’s internal ExploitBench cybersecurity evaluation. The initial version of the blog post listed Sol’s score at 5.5%. Subsequent versions saw this figure more than double to 11.5%. OpenAI has indicated that it is currently investigating reverting this number back to 5.5%, explaining that the 11.5% result reflected a "reasoning level that is not commercially available for Sol," further complicating the picture of what constitutes a ‘true’ performance metric.

OpenAI has consistently highlighted Astra’s exceptional capabilities in mathematics, a quality prominently featured in the opening paragraph of its announcement. While Astra’s score on the FrontierMath Tier 4 (v2) evaluation remained constant at 97.6% across all snapshots, the scores for other models were temporarily altered in a way that made Astra appear even more dominant. For instance, Anthropic’s Fable 5.1 model initially scored 87.8% in the first snapshot at 2:23 p.m. on Sept. 3. By 5:17 p.m., this score had dropped by nearly ten percentage points to 78%. Today, it has rebounded to 83%. Similarly, GPT-5.6 Sol’s math scores fluctuated from an initial 83%, down to 80.5%, and then back up to 83%. These transient dips in competitor scores, while Astra’s remained steadfast, created a fleeting impression of a wider performance gap than currently stands.

The alterations, it turns out, began even before the initial 2 p.m. publication. An embargoed pre-publication draft provided to Fortune and other media outlets listed Astra’s score on the ARC-AGI-3 evaluation, a benchmark for advanced reasoning, at 98.6%. This figure was subsequently boosted to an astonishing 99.99% in the live blog post. An OpenAI spokesperson justified this by stating, "We always verify evals before publication so adjustments between draft and final version are normal." They also pointed to independent verification from the Arc Prize Foundation, which found Astra performing at 99.9% when equipped with a particularly powerful "harness" – a set of tools designed to aid the model in task completion. However, when given the benchmark’s standard harness, Astra’s score dropped significantly to 63%, though this still represents a substantial lead over other publicly released AI models. OpenAI clarified that "things like harness, reasoning level and other factors inform evals," highlighting the intricate, and sometimes opaque, variables at play in benchmark reporting. Even Astra’s coding capabilities saw a marginal, yet noteworthy, boost in later versions of the blog post, rising from 57.7% to 57.9%, suggesting a meticulous effort to present the most favorable figures.

It is also important to note that not all changes exclusively favored OpenAI’s models. In some instances, scores for competitor models improved. For example, two Anthropic models saw their scores rise on the healthcare-focused HealthBench Professional evaluation. Claude Fable 5.1 increased from 56.6% to 58.1%, and Opus 5 improved from 54.5% to 56.4%. These scores for rival AI companies are typically derived from publicly available leaderboards and are not directly assessed by OpenAI itself, offering a counterpoint to the narrative of one-sided adjustments.

"Benchmaxxing" or Improving Accuracy? The Industry Debate

OpenAI has stated that different research teams are responsible for overseeing, calculating, and reporting various metrics to a central team for publication. The company is transparent about the fact that the reported numbers represent performance achieved under the "best possible conditions" and may not perfectly reflect the models available in the production ChatGPT product that most users access. A disclaimer on the blog states, "Evaluation scores are the maximum at any effort," with further caveats provided in footnotes for each metric.

However, the pursuit of "accuracy" in AI benchmarks is inherently complex, as multiple figures can be considered valid depending on the specific testing conditions. This inherent variability has led to concerns among some AI experts about the practice known as "benchmaxxing." This widely acknowledged, albeit controversial, practice in the AI industry—not exclusive to OpenAI—involves maximizing benchmark scores by repeatedly re-running evaluations under various conditions until the most favorable results are obtained.

Anka Reuel and Mike Hardy, researchers at the Stanford Intelligent Systems Laboratory and Stanford Trustworthy AI Lab, voiced their concerns, suggesting that such rapid adjustments could be driven by marketing objectives. "This can be done in a very tight timeframe, and it’s better for their marketing," they noted. They also criticized the lack of transparency in OpenAI’s documentation, specifically pointing out that the GPT-6 Astra system card, which should provide detailed technical information on evaluation methodologies, often falls short. For instance, regarding the internal hallucination benchmark, they stated the system card provides "barely any details about the evaluation," not even including "the number of test items," making it difficult for external researchers to fully understand or replicate the results.

Evaluation Score Debates Haunt the AI Industry

The question of benchmark accuracy and ethical reporting is not new to the AI industry, with several high-profile incidents preceding OpenAI’s current predicament. In 2025, Meta faced significant backlash after reports emerged suggesting it had artificially boosted scores for its Llama 4 model by publishing results from an internal, potentially more optimized, version rather than the one made publicly available. Yann LeCun, Meta’s former chief AI scientist, later candidly admitted that the company had "fudged" the benchmark results, sending shockwaves through the community and highlighting the pressures to perform in the AI race. The rapid evolution of evaluation metrics itself is also a factor; for instance, ExploitGym, a cybersecurity benchmark that was central to an incident in July where OpenAI’s models reportedly went rogue and attacked Hugging Face, was only created in 2026, illustrating how quickly the goalposts for assessment can shift.

Vincent Sunn Chen, an AI engineer leading benchmark and evaluation research at Snorkel AI, offered a pragmatic perspective, stating that it’s "not unusual" for benchmark scores to shift in the final hours before a model launch. In an email, he explained, "It’s usually a function of final launch logistics. A benchmark score reflects a specific measurement setup: the model checkpoint, configuration (including how much time and compute the model is allowed), harness, eval/grading configuration (e.g., non-determinism in the judge). All of those are typically still shifting in the final days before a launch, so I’m not surprised that there were some updates." However, Chen also stressed the importance of developing clear industry norms, advocating for companies to explicitly report what changes have been made to an assessment when benchmark performance numbers are revised. This, he argued, would enable researchers and the public to interpret the results with greater clarity and confidence.

Benchmark results carry significant weight for several critical reasons. They are the primary mechanism through which AI companies measure their progress, charting advancements in model capabilities. Crucially, they also serve as a public scoreboard in the intense competition among rival AI developers. Topping these leaderboards can provide a substantial competitive edge, helping companies attract new customers, secure lucrative contracts, and even recruit top-tier engineers and researchers in a talent-scarce market.

However, as this incident with OpenAI’s Astra launch vividly illustrates, interpreting these benchmark scores is a technically complex endeavor, presenting a formidable challenge for companies aiming to present their results to the public in an easily digestible and credible format. The combination of shifting metrics, contradictory explanations, and accusations of intellectual dishonesty could significantly muddy the waters for potential customers and investors trying to discern which models genuinely offer the best performance for specific tasks. This confusion threatens to undermine the narrative of market leadership that OpenAI undoubtedly wishes to establish, especially ahead of a rumored 2027 initial public offering (IPO). Ultimately, the credibility of benchmark reporting is paramount, and any perceived lack of transparency or consistency could erode trust, impacting not just individual companies but the broader public perception and responsible development of the burgeoning AI industry.

Leave a Reply

Your email address will not be published. Required fields are marked *