TrueSeeker AI · Verified claim report Case dec2aaa484 · 2026-09-09

§ Claim under review · Release

"OpenAI has released GPT-6 Astra, its most powerful AI yet, capable of using computers and browsers, handling long multi-step tasks, creating documents and presentations, and building software autonomously" Secondary claims in the post (also investigated): (a) "Astra scored 97.6% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 100% on ExploitBench" (b) "It is rolling out to ChatGPT Plus, Pro, Business, and Enterprise users, with API access also coming through OpenAI, Microsoft Azure, and AWS Bedrock" (c) Image card: "'Welcome to the AGI era.' - Greg Brockman, President of OpenAI", paired with a definition of AGI as matching or surpassing human cognition "across virtually all domains and tasks"

Circulating claim, as submitted.

Verdict

Source exists but framing is misleading

Confidence

High
§

Summary

The release is real. OpenAI did launch GPT-6 Astra on September 3, 2026, and its own announcement does describe computer use, browser use, multi-step workflows, and creating documents and presentations. The three benchmark numbers in the post are also OpenAI's real published figures, not inventions. The problem is what got left out. The headline 99.9% score on ARC-AGI-3 came from OpenAI's own testing setup, and the organization that built that benchmark ran the same model through its standard setup the same day and got 62.7%. The post also omits results where Astra did worse, including a reasoning test where it scored 57.2% against a competitor's 65%, and one independent evaluator initially rated it no better than OpenAI's previous model. The word "autonomously" was added by the post; OpenAI said "best model for software engineering," which is not the same thing. Finally, when the post went up the model was only available to a small set of vetted organizations, not to the subscriber tiers listed, and OpenAI's CEO apologized for a messy rollout the next day. The "Welcome to the AGI era" quote is genuine, but it was OpenAI's president speaking to his personal view at a press briefing, and the benchmark's own authors have said they are not claiming AGI.

§

The readings

key figures from the evidence
99.9 %

ARC-AGI-3 score using OpenAI's Provider Adapter harness

62.7 %

ARC-AGI-3 independent standard-harness score, same model

61 pts

Artificial Analysis intelligence rating, tied with predecessor

§

Why this verdict

The underlying release is real and primary-confirmed: OpenAI's announcement, system card, safety overview, and API model page all exist and were published September 3, 2026, and the capability list in the claim is close paraphrase of OpenAI's own copy. What the post does is convert a vendor launch announcement into reported fact, strip the caveats that the primary sources themselves supply, and add an autonomy claim OpenAI did not make. The decisive gap is ARC-AGI-3: the benchmark's own operator published 62.7% under its standard harness on the same day as the 99.9% Provider Adapter figure, and the post carries only the number that supports the AGI framing. I considered and rejected "Mostly accurate," because the omitted harness condition and the omitted regressions materially change what a reader concludes rather than merely simplifying it. I rejected "False" and "Unverified," because the model, the numbers, and the quote all check out against primary sources. I rejected "Partially accurate but misleading" as the closer call: the framing problem here is exaggeration of a real and correctly cited source rather than misuse of a fact, which is exactly the "framing" category. As-of date for the availability and comparative elements: 2026-09-09. Confidence is High because the primary artifacts were retrieved on both sides, the vendor announcement and the independent benchmark operator's contradicting run, though the specific "most powerful" superlative rests on vendor framing and is contested by independent aggregators.
§

Evidence

The release is real and confirmed on OpenAI's own channels. OpenAI's news index lists "GPT-6 Astra: A new generation of intelligence" dated Sep 3, 2026, alongside a "Safety overview: GPT-6 Astra" and a "GPT-6 Astra System Card" on the same date.

The API documentation lists the model as built for "complex reasoning, coding, computer use, research, and document creation," with reasoning.effort supporting low, medium, high, xhigh, and max.

The advertised capabilities in the claim track OpenAI's own wording closely. OpenAI states that Astra "combines the intelligence required for complex problems with the ability to carry out multistep workflows and produce polished documents, spreadsheets, and presentations," and that it is "our best model for adhering to existing templates and producing slides."

The announcement says Astra "is state-of-the-art on computer use, browsing, software engineering, cybersecurity, science, and professional work," that it "saturates FrontierMath Tier 4 with a 98% score," and that it "also saturates ARC-AGI-3 with a 99.9% score and ExploitBench with a 100% score."

The three headline numbers in the caption are OpenAI's own published figures, but the ARC-AGI-3 number is harness-dependent, and the benchmark's own operator published a very different result the same day. ARC Prize reports that "GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness... and 99.9% for $19K with a Provider Adapter harness," where the Provider Adapter "preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work."

ARC Prize's own summary describes Astra as scoring "63% on ARC-AGI-3, 99% via a new provider adapter harness." The widely shared comparison was also not like-for-like: "The figure that spread was 99.9% against GPT-5.6 Sol's 7.8%. Those are not the same test. Astra's 99.9% came from the Provider Adapter; Sol's 7.8% came from the standard harness."

Independent aggregate evaluators disagree with the "most powerful" framing on general intelligence. The Decoder reports that Epoch AI places Astra clearly first with 169 points while "Artificial Analysis rates it no better than its predecessor," at 61 points, level with GPT-5.6 Sol and behind Claude Fable 5.1 at 66.

Artificial Analysis' own writeup reports "a 6 point gain in Humanity's Last Exam" offset by "a drop of ~80 Elo points in GDPval-AA v2... measuring economically valuable tasks across 44 occupations," plus "2-3 point regressions on other evaluations."

On its Coding Agent Index, Artificial Analysis found "GPT-6 Astra equals Fable 5 at less than half the cost," with "Fable 5.1 in Claude Code" leading at 70. A later index revision moved Astra up: Artificial Analysis raised its Intelligence Index to version 4.3, where GPT-6 Astra (max) and Claude Fable 5.1 both score 53 and share the top of the ranking.

On availability at the time the post was published, the post overstated day-one access. OpenAI's announcement says Astra "is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API, Microsoft Azure, and AWS Bedrock."

Trade press reported that only organizations enrolled in OpenAI's Daybreak cybersecurity program could initially access the model, while "Plus, Pro, Business, and Enterprise ChatGPT subscribers, along with developers using the OpenAI API, were left out."

Altman acknowledged: "First, sorry for the messy rollout." Access widened the following day: Altman posted that "GPT-6 Astra is now available to all Pro, Enterprise, and Business Premium users in Work/Codex, and is available in the API. We will start rollout to Plus and Business users next."

On the AGI framing, the quote is real but the attribution of the judgment to OpenAI as an institution is not supported. Axios reports that Brockman called Astra a "generational leap" and said it could eventually be seen as the arrival of AGI, and that he "personally believes OpenAI has reached AGI, while leaving users to decide."

He closed the press briefing with "Welcome to the AGI era." The benchmark operator whose score anchored the AGI narrative declined to endorse it: ARC Prize's own position is that "the benchmark's authors say they are not claiming AGI."

Chollet does not call the result proof of AGI, though he describes progress running "twice as fast" than expected and is moving up his forecast.

§

Findings

✓ What's accurate 7

  • GPT-6 Astra exists and was released by OpenAI on September 3, 2026. This is confirmed on OpenAI's announcement page, safety overview, system card, and API model documentation
  • It is OpenAI's current flagship, succeeding GPT-5.6 Sol, and OpenAI describes it as its most capable model
  • OpenAI does claim computer and browser use, multi-step workflow execution, and creation of documents, spreadsheets, and presentations matching user templates and style
  • The three caption numbers are genuinely OpenAI's published figures, not invented. 100% on ExploitBench and 99.9% on ARC-AGI-3 appear verbatim in OpenAI's announcement
  • The rollout list of Plus, Pro, Business, Enterprise plus OpenAI API, Microsoft Azure, and AWS Bedrock matches OpenAI's announcement exactly
  • Brockman did say "Welcome to the AGI era," at a press briefing on launch day
  • Independent confirmation exists for a real, large gain in agentic and computer-use work. ARC Prize records genuine SOTA on ARC-AGI-3 and human-beating action efficiency, and Artificial Analysis records large token-efficiency gains

≈ What's misleading 8

  • Marketing as evidence: the post credits "Source: OpenAI" and reproduces the vendor's own superlatives and vendor-run scores as settled fact. For existence and pricing, OpenAI is the primary source. For "most powerful" and every comparative score, it is an interested party's self-report, and two independent aggregators reached opposite conclusions about it.
  • Harness mismatch: the 99.9% ARC-AGI-3 figure was produced under OpenAI's Provider Adapter harness, which preserves reasoning state between calls. The benchmark's own operator, using its standard harness, scored the same model at 62.7%. Both numbers were published the same day, and the post carries only the favorable one. The widely repeated jump from 7.8% to 99.9% compares two different harnesses.
  • Benchmark cherry picking: the caption selects the three saturated scores and omits the regressions in the same evidence base, including Humanity's Last Exam with tools at 57.2% against Fable 5.1's 65.0%, an approximately 80 Elo drop on GDPval-AA v2, and DeepSWE v1.1 at 74.1% where Meta's Muse Spark 1.3 is reported ahead at 75.4%.
  • Capability extrapolation: "building software autonomously" adds a word OpenAI did not use. OpenAI's claim is "best model for software engineering to date," which is a quality claim about assisted work, not an autonomy claim. Independent coding indices show Astra tied with or behind Claude Fable 5.1 rather than operating unsupervised.
  • Cost compute omission: the headline numbers were run at maximum reasoning effort, and the ARC-AGI-3 runs cost between $19,000 and $26,000. Presenting those figures as what the model does omits the spend that produced them and does not describe default product behavior.
  • Unreleased as released: on September 4, when the post was published, access was limited to organizations in the Daybreak cybersecurity program. Plus, Pro, Business, Enterprise, and API users did not yet have it, and Altman apologized for a "messy rollout." The post's own wording "is rolling out to" is defensible, but the headline "has released... capable of" invites a reader to think the described product was in their hands, which it was not.
  • Misattribution: the image card presents "Welcome to the AGI era" alongside a formal definition of AGI in a layout that reads as an OpenAI institutional position. The reporting shows Brockman spoke to his personal belief and explicitly left the judgment to users, and the operator of the benchmark that anchored the claim declined to endorse it.
  • Marketing as evidence, second instance: two of the three headline benchmarks are not arm's-length. ExploitBench is an internal OpenAI evaluation, and Epoch AI discloses that OpenAI funded part of FrontierMath's development and holds access to part of the set. The post presents all three as neutral scoreboards.

? What's uncertain 6

  • The exact FrontierMath Tier 4 figure. OpenAI's prose says 98% while the circulating table figure is 97.6%. Both trace to OpenAI. I did not retrieve the announcement's benchmark table directly, so I cannot state which is the canonical row
  • Whether the ARC-AGI-3 headline is 99.9% or 98.6%. Both appeared in launch coverage, and OpenAI is reported to have revised five metrics after launch. I did not retrieve the revision log
  • Current tier-by-tier availability as of 2026-09-09. Pro, Enterprise, Business Premium, and the API were confirmed live on September 4. Plus and Business were "next," and I found no official confirmation that rollout to those tiers is complete
  • The $10 / $50 per million token price is reported by trade analysis. I did not open OpenAI's pricing page, so it is unconfirmed here
  • Whether the shipped product performs the described document, presentation, and software tasks reliably for ordinary users. No independent test of production-effort behavior was found, and Artificial Analysis reported lower presentation-quality Elo in one agentic test
  • Whether the system card resolves the harness question. Analyses point to it as the place to check production effort levels, and I did not read the card's benchmark appendix
Distortion flags marketing as evidence harness mismatch benchmark cherry picking capability extrapolation cost compute omission unreleased as released misattribution
§

Sources

14 of 15 linked to records
[1]

OpenAI, "GPT-6 Astra: A new generation of intelligence", Sep 3 2026

primary vendor announcement of record
https://openai.com/index/gpt-6-astra/ ↗
[2]

ARC Prize Foundation, "OpenAI's GPT-6 Astra on ARC-AGI-3", Sep 3 2026

primary independent benchmark operator, published methodology
https://arcprize.org/blog/astra ↗
[3]

OpenAI GPT-6 Astra System Card / Deployment Safety Hub

primary vendor safety artifact of record
https://deploymentsafety.openai.com/gpt-6-astra ↗
[4]

OpenAI API docs, model page "gpt-6-astra"

primary vendor documentation of record
https://developers.openai.com/api/docs/models/gpt-6-astra ↗
[5]

OpenAI, "Safety overview: GPT-6 Astra"

primary vendor safety artifact
https://openai.com/index/safety-overview-gpt-6-astra/ ↗
[6]

Artificial Analysis, "Benchmarking GPT-6 Astra"

primary independent evaluator
https://artificialanalysis.ai/articles/benchmarking-gpt-6-astra ↗
[7]

Sam Altman, X post on staged availability, Sep 4 2026

primary vendor principal, informal
https://x.com/sama/status/2095973658867171733 ↗
[8]

Axios, "OpenAI releases new model GPT-6 Astra, says it may represent AGI"

secondary named-outlet journalism
https://www.axios.com/2026/09/03/openai-astra-gpt-6-agi-brockman ↗
[9]

Washington Post, "OpenAI's Greg Brockman says its new model Astra is AGI"

secondary named-outlet journalism
https://www.washingtonpost.com/technology/2026/09/03/openai-greg-brockman-says-its-new-model-astra-is-agi/ ↗
[10]

CNBC, "OpenAI announces rollout of GPT-6 Astra model"

secondary named-outlet journalism
https://www.cnbc.com/2026/09/03/open-ai-astra-gpt-6-cyber.html ↗
[11]

The New Stack, "GPT-6 Astra's score of 98.6% looked like AGI. Then researchers read the fine print."

secondary trade press
https://thenewstack.io/astra-arc-agi-benchmark/ ↗
[12]

CSO Online / Computerworld, "Sam Altman calls GPT-6 Astra rollout 'messy'"

secondary trade press
https://www.csoonline.com/article/4219249/ ↗
[13]

The Decoder, "Benchmarks disagree on GPT-6 Astra..."

secondary trade press
https://the-decoder.com/benchmarks-disagree-on-gpt-6-astra ↗
[14]

Vellum, DataCamp, MindStudio benchmark breakdowns

tertiary vendor-blog analysis
This citation could not be independently verified.
[15]

Wall St Engine (X), benchmark list

tertiary social aggregation, cited only to show the number that circulated
https://x.com/wallstengine/status/2095578934452818082 ↗
How links are chosen. A source is linked only when the address comes from the investigation's own retrieval or from a registry lookup (PubMed, Crossref) that matches the citation's title and year. Author lists shown as registry-verified come from the registry record, not from the report text. Citations that cannot be matched are labeled, never guessed.
This is one case on the record See the full case, browse the archive, and search every checked claim on TrueSeeker AI Open on ai.trueseeker.com →