TrueSeeker AI · Verified claim report Case 3750102e15 · 2026-09-23

§ Claim under review · Benchmark

"97.6% on FrontierMath Tier 4" / "97.6 on Frontier Math, Tier 4, the hardest math benchmark that exists" (Rowan Cheung, TikTok, sponsored video, 2026-09-17)

Circulating claim, as submitted.

Verdict

Mostly accurate

Confidence

High
§

Summary

The number is real. OpenAI's GPT-6 Astra did score 97.6% on FrontierMath Tier 4, and this is not just a company talking point: Epoch AI, the organisation that actually built and runs the benchmark, tested the model itself, got 98%, and has formally declared the benchmark saturated. Astra also solved the last Tier 4 problem no AI had cracked. Three things the video leaves out. First, the score is on a rewritten version of the test, issued in June 2026 after Epoch found errors in 42% of the problems and removed several. Second, Epoch openly discloses that OpenAI funded FrontierMath and has access to most of the problems and answers, which is worth knowing before treating any score on it as clean. Third, calling Tier 4 "the hardest math benchmark that exists" is already out of date: Epoch's newer Erdős benchmark of unsolved problems saw Astra solve just 2 of 68, a 3% score, and Astra actually trails Anthropic's Claude on Humanity's Last Exam. The specific claim holds up. The "benchmarks are saturated" story around it is selective, and the video is a disclosed paid promotion.

§

The readings

key figures from the evidence
97.6 %

GPT-6 Astra on FrontierMath Tier 4 v2, OpenAI table

98 %

Epoch AI's independent run of Astra on Tier 4, saturated

2 of 68

Astra solved on Epoch's newer FrontierMath Erdos set

§

Why this verdict

As of 2026-09-22, the specific number checks out. 97.6% on FrontierMath Tier 4 v2 appears in OpenAI's launch announcement table, is reproduced consistently by several independent analysts, and is corroborated by the benchmark's own operator, Epoch AI, which ran the model with pre-release access, reported 98%, and declared the benchmark saturated. I considered and rejected "Accurate" because the claim drops the v2 qualifier on a benchmark that was substantially rebuilt three months earlier and omits that the evaluator was funded by the vendor it was scoring. I considered and rejected "Partially accurate but misleading" because no source I found contradicts or bounds away the operative proposition. The missing context is real but it does not change what the number is. The surrounding video framing, specifically "the hardest math benchmark that exists" and "benchmarks are basically already saturated," is weaker than the number itself and is addressed above, but it is not the claim submitted for verification. Confidence is High because the benchmark operator independently confirms the result, which is the strongest evidence class available for a benchmark claim.
§

Evidence

The number is real and it is close to correct. OpenAI's own launch page states in prose that Astra "saturates FrontierMath Tier 4 with a 98% score." The more precise 97.6% figure comes from the Academic comparison table in that same announcement, which multiple independent analysts reproduce consistently: Vellum, DataCamp, The Decoder and Yotta Labs all record FrontierMath Tier 4 v2 at 97.6% for Astra against 83.0% for GPT-5.6 Sol and 87.8% for Claude Fable 5.1.

Critically, this is not a vendor-only number. Epoch AI, which builds and operates FrontierMath, ran the model itself with pre-release access from OpenAI and reports its own result. Epoch states that Tier 4 launched on 2025-07-11 with a top score of 5%, that less than 14 months later the top score is 98%, and that it now considers the benchmark saturated. Epoch separately reports that GPT-6 Astra solved the final previously-unsolved Tier 4 problem, authored by Jay Pantone, and notes that unlike many Tier 4 problems this one was not solved via an unintended shortcut. Epoch's 98% and OpenAI's 97.6% are the same result at different rounding: Tier 4 v2 contains 43 problems, and 42 of 43 is 97.7%.

So the operative proposition, that Astra scored 97.6% on FrontierMath Tier 4, is supported by the vendor's table and independently corroborated by the benchmark operator's own run.

§

Findings

✓ What's accurate 5

  • GPT-6 Astra exists. OpenAI released it on 2026-09-03, and it is rolling out via ChatGPT paid plans and the API, matching the video's rollout statement.
  • The 97.6% FrontierMath Tier 4 figure is genuinely in OpenAI's announcement table. It is not invented and it is not a misread of another row.
  • The figure is independently corroborated. Epoch AI, the benchmark's operator, tested the model itself and reports 98% and formal saturation.
  • Astra did solve the last unsolved Tier 4 problem, per Epoch, and Epoch confirms that solution was not a shortcut.
  • The prime numbers claim in the video also checks out at the artifact level. OpenAI states Astra "improved a term in a bound on these gaps that had remained unchanged for more than 80 years" and published the proofs.

≈ What's misleading 6

  • Omitted qualifier: the claim says "FrontierMath Tier 4" with no version. The score is on Tier 4 v2, a 43-problem set issued in June 2026 after Epoch corrected errors in 42% of the original problems, with 7 problems removed entirely. A viewer hears a stable yardstick. The yardstick was rebuilt three months before the score.
  • Marketing as evidence: the 97.6% digits come from OpenAI's own announcement table in a disclosed sponsored video. The independent corroboration exists and is genuine, but the video presents the vendor's number as simply the fact, with no indication of who ran it. The saturation finding belongs to Epoch. The 97.6% belongs to OpenAI.
  • Omitted qualifier: the evaluator is not a disinterested party. Epoch discloses that OpenAI funded FrontierMath and has access to problem statements and solutions outside a 50-question holdout. That disclosure is material to any Tier 4 score and appears nowhere in the video.
  • Capability extrapolation: "the hardest math benchmark that exists" is contradicted by the same operator's own newer artifact. Epoch launched FrontierMath Erdős, 68 historically unsolved problems requiring Lean solutions, and Astra scored 3%, solving 2 of 68. Every other model scored zero. Tier 4 was the hardest. It is not any more.
  • Benchmark cherry picking: the video says frontier intelligence benchmarks are "basically already saturated." Astra scores 57.2% on Humanity's Last Exam with tools, behind Claude Fable 5.1 at 65.0%, Fable 5 at 63.8% and Opus 5 at 63.6%. On the independent Artificial Analysis Intelligence Index it sits at 61.2, effectively tied with its own predecessor and behind Fable 5.1 at 65.7. The saturated benchmarks are the ones shown.
  • Internal inconsistency on the companion number: the spoken transcript says 99.9% on ARC-AGI-3 while the caption of the same post says 98.6%. OpenAI's page says 99.9%. The Decoder notes that 99.9% was achieved "under its own test conditions," and other coverage reports that the figure depends on a stateful OpenAI harness while independent stateless runs land far lower. This is a harness mismatch on the secondary number, not on the FrontierMath claim.

? What's uncertain 4

  • I did not retrieve OpenAI's Academic comparison table directly. The exact digits 97.6 reach me through multiple independent secondary reproductions of that table, which agree with each other. The primary page I did reach states 98% in prose.
  • The Tier 4 harness protocol is not established in the sources I reached: attempts per problem, whether the score is a single run or a mean across runs, and the reasoning-effort setting used for the specific 97.6% entry are not published in what I found. 42 of 43 is 97.7%, so 97.6% may be an average across runs rather than a single pass, but I will not reconstruct a protocol I did not see.
  • Whether any contamination affected the result. OpenAI's access to Tier 4 problem statements and solutions outside the holdout is disclosed fact. Whether it bears on this score is not something the public evidence resolves either way.
  • The mathematical results, including the prime gap bound, are published as Lean proofs. Commentary flags that the large proof file has not yet had independent human semantic review. I could not verify the review status.
Distortion flags omitted qualifier marketing as evidence capability extrapolation benchmark cherry picking harness mismatch
§

Sources

9 of 9 linked to records
[1]

OpenAI, "GPT-6 Astra: A new generation of intelligence" (launch announcement of record)

primary vendor
https://openai.com/index/gpt-6-astra/ ↗
[2]

Epoch AI, "AI Benchmarks & Capabilities" hub and GPT-6 Astra model page (benchmark operator of record)

primary independent evaluator with disclosed OpenAI funding relationship
https://epoch.ai/benchmarks ↗
[3]

Epoch AI, FrontierMath Tier 4 (v2) benchmark page, including the 2026-06-12 correction notice and conflict-of-interest disclosure

primary benchmark operator
https://epoch.ai/benchmarks/frontiermath-tier-4-v2 ↗
[4]

Epoch AI Research on X, saturation statement

primary benchmark operator
https://x.com/EpochAIResearch/status/2098103831502708864 ↗
[5]

Epoch AI, "Clarifying the creation and use of the FrontierMath benchmark" (conflict-of-interest statement)

primary benchmark operator
https://epoch.ai/latest/openai-and-frontiermath ↗
[6]

Epoch AI, "Announcing FrontierMath Erdős"

primary benchmark operator
https://epoch.ai/latest/announcing-frontiermath-erdos ↗
[7]

Vellum, "GPT-6 Astra Benchmarks Explained" (reproduces the announcement's Academic table row by row)

secondary vendor-adjacent analyst tooling
https://www.vellum.ai/blog/gpt-6-astra-benchmarks-explained ↗
[8]

DataCamp, "GPT-6 Astra: Features, Benchmarks, and Pricing"

secondary tech education publisher
https://www.datacamp.com/blog/gpt-6-astra ↗
[9]

The Decoder, GPT-6 Astra benchmark coverage

secondary named tech outlet
https://the-decoder.com/gpt-6-astra-is-the-first-model-making-openai-willing-to-declare-the-agi-era/ ↗
How links are chosen. A source is linked only when the address comes from the investigation's own retrieval or from a registry lookup (PubMed, Crossref) that matches the citation's title and year. Author lists shown as registry-verified come from the registry record, not from the report text. Citations that cannot be matched are labeled, never guessed.
This is one case on the record See the full case, browse the archive, and search every checked claim on TrueSeeker AI Open on ai.trueseeker.com →