TrueSeeker AI · Verified claim report Case ddedfa1ec1 · 2026-09-05

§ Claim under review · Benchmark

"OpenAI launched GPT-6 Astra, a new AI model achieving 98.6% on ARC-AGI-3, 100% on ExploitBench, 97.6% on FrontierMath Tier 4 (v2), and 74.1% on DeepSWE v1.1, outperforming GPT-5.6 Sol and Claude Fable 5.1"

Circulating claim, as submitted.

Verdict

Source exists but framing is misleading

Confidence

High
§

Summary

OpenAI really did launch GPT-6 Astra on September 3, 2026, and all four numbers in this post are real rather than invented. The problem is what was left out. The 98.6% on ARC-AGI-3 came from a special OpenAI-supplied test setup; the benchmark's own operator, ARC Prize, ran the identical model through its neutral setup and got 62.7%, and reporting indicates the setup behind the high score is not part of the product people can buy. The 100% on ExploitBench came from a test OpenAI itself flagged as possibly contaminated, and on the clean version it built to correct for that, the model scored around 39%. The claim that Astra outperforms Claude Fable 5.1 is contradicted by OpenAI's own comparison table, where Fable 5.1 wins on Humanity's Last Exam, and by independent benchmarking firm Artificial Analysis, which rated Astra level with its own predecessor and behind Fable 5.1 while costing 2.5 times more. On the coding benchmark cited, Astra's 74.1% is statistically tied with two rival models and behind Meta's Muse Spark 1.3 at 75.4%. Astra does appear genuinely strong at computer use, cybersecurity and hard mathematics, but the sweeping framing goes well beyond what the evidence supports.

§

The readings

key figures from the evidence
98.6 %

ARC-AGI-3 via OpenAI's Provider Adapter harness

62.7 %

ARC-AGI-3 via ARC Prize's neutral Standard harness

39.0 %

ExploitBench on contamination-controlled internal port, vs claimed 100%

§

Why this verdict

As of 2026-09-04, the launch is real, and all four cited numbers exist in real, retrievable artifacts, so "False" and "Unverified" are both wrong, and refuting correct measurements would be an error rather than rigor. But the accurate family is closed by the contradiction test: the claim's operative proposition is that these scores show Astra outperforming GPT-5.6 Sol and Claude Fable 5.1, and OpenAI's own comparison table shows Fable 5.1 ahead on Humanity's Last Exam while the independent Artificial Analysis index puts Astra level with Sol and behind Fable 5.1 overall. That is a source I cite denying the operative proposition, not a simplification, so "Mostly accurate" is unavailable. I chose "Source exists but framing is misleading" over "Partially accurate but misleading" because the numbers are transcribed correctly rather than garbled: the entire distortion lives in stripped conditions, the harness for ARC-AGI-3, the contamination caveat for ExploitBench, the comparator choice and noise band for DeepSWE, and in an unscoped superiority verb. Confidence is High because the deciding artifacts were retrieved from the benchmark operators themselves, ARC Prize and Datacurve and Epoch AI, and from OpenAI's own documents, rather than resting on vendor claims alone. ---
§

Evidence

The release is real and fully documented on official channels. OpenAI published a launch page, a system card, a preparedness write-up and an API model page for GPT-6 Astra on September 3, 2026, and Microsoft published a Foundry availability post. This part of the claim needed no adjudication.

All four numbers are real and traceable. None is fabricated. What the sources show is that each one carries a condition that the post removed.

ARC-AGI-3. ARC Prize, the benchmark's operator, ran the model itself and published verified results. With ARC Prize's own provider-neutral Standard harness, Astra scored 62.7% at max reasoning, at a run cost near $26,000. With OpenAI's Provider Adapter harness, which ARC Prize describes as preserving opaque reasoning state between requests and using compaction, the model reached 99.9% at high reasoning. The 98.6% figure in the post is the same Provider Adapter condition at max reasoning. So 98.6% and 62.7% are the same model weights, same benchmark, two interfaces. ARC Prize said it will label the two conditions separately going forward. The New Stack reported that the harness producing the headline number is not itself part of the shipped product. Note that OpenAI's own launch page headlines 99.9%, not 98.6%, because its tables report the maximum score at any reasoning effort.

ExploitBench. The 100% is confirmed in OpenAI's own preparedness document. That same document states OpenAI had contamination concerns about the benchmark, and therefore built a contamination-controlled internal port using high-severity V8 vulnerabilities disclosed between June and August 2026. On that controlled version, reporting of OpenAI's figures puts Astra at roughly 39.0% arbitrary code execution. OpenAI's own footnote further warns that some included vulnerabilities may not permit arbitrary code execution under the evaluation's constraints, so a 100% success rate may not be achievable in principle.

FrontierMath Tier 4 (v2). The 97.6% is OpenAI's reported figure. Epoch AI's own benchmark page carries a standing disclosure that FrontierMath was developed with OpenAI funding and that OpenAI has exclusive access to a subset of the benchmark. Epoch also notes a June 2026 update that corrected errors in 42% of problems. No independent re-run of Astra on this benchmark was found.

DeepSWE v1.1. This is the best-corroborated number in the claim. The independent Datacurve leaderboard lists Astra at roughly 74% with a stated uncertainty of ±3 in the shared mini-swe-agent setup, consistent with OpenAI's 74.1%. But the benchmark is 113 tasks, so one task is about 0.9 points, and the surrounding field is bunched: Claude Opus 5 at 73.7% and Gemini 3.8 Flash at 73.8%. Astra's margin over Opus 5 is 0.4 points, less than half a single task and far inside the published uncertainty band. Meta's Muse Spark 1.3 sits at 75.4% and was the leaderboard's number one on the claim's own publication date.

On the comparative claim. OpenAI's own comparison table contains rows Astra loses. On Humanity's Last Exam with tools, Astra scores 57.2% against Claude Fable 5.1's 65.0%. On FrontierCode 1.1 splits, Fable 5 and Opus 5 edge Astra. Independently, Artificial Analysis measured Astra at about 61 on its Intelligence Index v4.1.1, level with GPT-5.6 Sol at 61 and behind Claude Fable 5.1 at roughly 66 and Opus 5 at roughly 63, while priced at 2.5 times Sol's rates. Artificial Analysis did find genuine gains for Astra in coding-agent performance and token efficiency.

One contrary source deserves explicit handling. CryptoBriefing published a piece arguing the 98.6% claim was unverified and disconnected from a leaderboard topping out near 30%. The benchmark operator subsequently ran and published verified results confirming the figure under the stated harness. On this point the operator's own artifact governs, and that commentary is superseded.


§

Findings

✓ What's accurate 7

  • OpenAI did launch GPT-6 Astra on September 3, 2026. This is confirmed on OpenAI's launch page, system card, API documentation and partner channels.
  • The staged rollout described in the caption matches OpenAI's own wording: limited organizations first, expanding over coming days to Plus, Pro, Business and Enterprise users, plus the API and AWS.
  • 98.6% on ARC-AGI-3 is a real, operator-verified result under a specified harness. It is not invented.
  • 100% on ExploitBench is stated by OpenAI in its own preparedness document.
  • 97.6% on FrontierMath Tier 4 v2 and 74.1% on DeepSWE v1.1 both appear in OpenAI's published table, and the DeepSWE figure is independently corroborated by Datacurve within rounding.
  • Astra does clearly lead on several axes: computer use, cybersecurity capability, hard mathematics, token efficiency, and long-horizon agentic work.
  • Greg Brockman's "Welcome to the AGI era" remark is confirmed by multiple outlets that attended the pre-launch briefing.

≈ What's misleading 8

  • **Omitted qualifier:** The claim presents 98.6% on ARC-AGI-3 as a property of the model. The operator's own verified results show the same model scoring 62.7% on the same benchmark through ARC Prize's provider-neutral harness. The number is a property of the model plus a specific stateful scaffold, and reporting says that scaffold is not what customers buy. A reader takes away "the model solves ARC-AGI-3," which is not what was measured.
  • **Cost compute omission:** The headline runs cost roughly $19,000 to $26,000 each and used maximum or high reasoning effort. OpenAI states its tables report the maximum score at any effort. None of this survives into the claim, which implies a routine capability.
  • **Eval contamination:** The 100% on ExploitBench is presented as a clean sweep. OpenAI itself flagged contamination concerns on that benchmark and built a controlled replacement, on which Astra scores roughly 39%. OpenAI also notes 100% may not be achievable in principle under the evaluation's constraints. The claim propagates the compromised number and drops the corrective one, which is the single largest gap between claim and evidence here.
  • **Benchmark cherry picking:** On DeepSWE the claim uses Fable 5.1's 67.4% as the reference point. Opus 5 (73.7%), Gemini 3.8 Flash (73.8%) and Muse Spark 1.3 (75.4%) are all at or above Astra, and Muse Spark led the public leaderboard on the day the post ran. The lead is manufactured by comparator selection.
  • **Margin presented beyond noise:** A 0.4 point gap on a 113-task benchmark with a published ±3 band is not a ranking. It is a tie.
  • **Exaggeration by overgeneralization:** "Outperforming GPT-5.6 Sol and Claude Fable 5.1" is stated without scope. OpenAI's own table shows Fable 5.1 ahead on Humanity's Last Exam with tools, 65.0% to 57.2%, and Fable 5 and Opus 5 ahead on FrontierCode splits. Independently, Artificial Analysis places Astra level with Sol and behind Fable 5.1 overall, at 2.5 times the price. The sweeping form of the comparison is contradicted by the very table it was drawn from.
  • **Marketing as evidence:** Three of the four numbers are vendor-run and were reproduced as neutral fact. For quality and superiority claims, the vendor channel establishes what OpenAI reports, not what is independently true. The FrontierMath figure additionally rests on a benchmark whose operator publicly discloses OpenAI funding and exclusive access to part of the problem set.
  • **Capability extrapolation:** The caption's "Welcome to the AGI era" and "new highs for the industry across... coding" convert a mixed benchmark table into a general intelligence verdict. The coding claim in particular is contradicted by the DeepSWE and FrontierCode evidence.

? What's uncertain 5

  • The original ExploitBench's total item count and composition are not established in what I retrieved. The 20-vulnerability figure belongs to OpenAI's contamination-controlled internal port, not necessarily the benchmark that produced the 100%.
  • No independent re-run of Astra on FrontierMath Tier 4 v2 was found. The 97.6% rests entirely on OpenAI's report, on a benchmark OpenAI funded.
  • The exact discrepancy between the post's 98.6% and OpenAI's headline 99.9% resolves to different reasoning-effort settings under the same harness, but I did not retrieve a statement from OpenAI explaining why the lower figure circulated first.
  • Whether ARC Prize's harness-labelling change alters how Astra's result is displayed going forward is announced but not yet observable.
  • Artificial Analysis index values are reported slightly differently across write-ups (61 versus 61.2, 65.7 versus 66). The ordering is consistent; the decimals are not load-bearing here.
Distortion flags omitted qualifier harness mismatch cost compute omission eval contamination benchmark cherry picking exaggeration marketing as evidence capability extrapolation
§

Sources

10 of 11 linked to records
[1]

ARC Prize, "GPT-6 Astra on ARC-AGI-3" blog and verified results page

primary benchmark operator of record, independent evaluator
https://arcprize.org/blog/astra ↗
[2]

OpenAI, "GPT-6 Astra: A new generation of intelligence" launch page including evaluation footnotes

primary vendor channel of record (duality rule applies)
https://openai.com/index/gpt-6-astra/ ↗
[3]

OpenAI, "Path to Astra: critical capabilities and frontier safeguards"

primary vendor safety documentation
https://openai.com/index/path-to-astra/ ↗
[4]

OpenAI, GPT-6 Astra System Card, Deployment Safety Hub

primary vendor system card
https://deploymentsafety.openai.com/gpt-6-astra ↗
[5]

Epoch AI, FrontierMath Tier 4 (v2) benchmark page with conflict-of-interest statement

primary benchmark operator
https://epoch.ai/benchmarks/frontiermath-tier-4-v2 ↗
[6]

Datacurve DeepSWE leaderboard and v1.1 methodology page

primary independent benchmark operator
https://deepswe.datacurve.ai/blog/deepswe-v1-1 ↗
[7]

Artificial Analysis, "Benchmarking GPT-6 Astra," Intelligence Index v4.1.1

primary independent evaluator with published methodology
https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra ↗
[8]

OpenAI API docs, GPT-6 Astra model page

primary vendor documentation
https://developers.openai.com/api/docs/models/gpt-6-astra ↗
[9]

The New Stack, "OpenAI will sell you Astra, but not the system that scored 98.6% on ARC-AGI-3"

secondary named-outlet trade journalism
https://thenewstack.io/openai-astra-harness-arc-agi-3/ ↗
[10]

VentureBeat, Fortune, DataCamp, Vellum, officechai benchmark write-ups

secondary named-outlet and vendor-adjacent analysis
This citation could not be independently verified.
[11]

CryptoBriefing, "Claims of GPT-6 Astra scoring 98.6% on ARC-AGI-3 don't hold up to scrutiny"

secondary commentary, superseded by operator verification (see below) ---
https://cryptobriefing.com/gpt-6-astra-arc-agi-3-claims-unverified/ ↗
How links are chosen. A source is linked only when the address comes from the investigation's own retrieval or from a registry lookup (PubMed, Crossref) that matches the citation's title and year. Author lists shown as registry-verified come from the registry record, not from the report text. Citations that cannot be matched are labeled, never guessed.
This is one case on the record See the full case, browse the archive, and search every checked claim on TrueSeeker AI Open on ai.trueseeker.com →