§ Claim under review · Benchmark
"GPT-5.4 xHigh tops Scale's standardized SEAL public board at 59.1% — the highest score any model reaches when every model runs the same harness on the public Pro set."
Verdict
Superseded
Confidence
HighSummary
This was true in June 2026 and is out of date now. Scale AI's standardized SWE-bench Pro public leaderboard does list GPT-5.4 at xHigh effort with 59.1%, so the number itself is real and comes from an independent runner rather than from OpenAI. But the same board now shows Meta's Muse Spark 1.1 at 61.5%, added after that model launched on July 9 2026, so 59.1% is not the highest standardized score any more. The board's error bars are about plus or minus 3.5 points and its own ranking ties the two models at first place, so neither should be described as a clear winner. The phrase "any model" is also broader than the board is: many current frontier models, including the newest Claude, Gemini and GPT-5.6 releases, have no entry on it at all. One further thing changed since the claim was written. In July 2026 OpenAI published an audit finding roughly 30% of this benchmark's public tasks were broken and withdrew its recommendation of the benchmark, which weakens any ranking claim built on it. One detail I could not check: some top entries on the board carry an asterisk whose meaning I was unable to read, and if it flags differently run submissions it would qualify the "same harness" premise.
The readings
key figures from the evidenceGPT-5.4 (xHigh) SWE-bench Pro public set resolve rate
Muse Spark 1.1 SWE-bench Pro public set score, now higher
SWE-bench Pro public tasks found broken in OpenAI audit
Why this verdict
Evidence
The deciding artifact is Scale's own public-set board, and its current state does not match the claim. The live public-set ranking reads Muse Spark 1.1 (NEW) at 61.50±3.10, gpt-5.4 (xHigh) at 59.10±3.56, Muse Spark at 55.00±3.60, claude-opus-4-6 (thinking) at 51.90±3.61, gemini-3.1-pro (thinking) at 46.10±3.60, claude-opus-4-5-20251101 at 45.89±3.60 , descending through gpt-5-2025-08-07 (High) at 41.78, qwen3-coder-480b-a35b at 38.70, and gemma-3-27b-it at 11.38 . Scale's leaderboard index shows the same top of board for the public open-source-repository track: Muse Spark 1.1 (NEW) 61.50±3.10, gpt-5.4 (xHigh) 59.10±3.56, Muse Spark 55.00±3.60 .
So the 59.1% figure is real and correctly attributed. The superlative attached to it is not: a higher standardized score exists on the same board. The rank numbering on the board interleaves as 1, 1, 3, 3, 5, 5, 5, indicating the board assigns tied ranks where confidence intervals overlap, which means GPT-5.4 (xHigh) is co-ranked first rather than the sole leader, and the model with the higher point score is Muse Spark 1.1.
The change is datable. Meta released Muse Spark 1.1 on July 9, 2026, alongside a public preview of the Meta Model API , after the claim's June 2026 vintage.
The claim's own origin chain confirms the staleness. The Morph LLM page that carries the sentence dates its table to June: "Scores below are from the public set (731 tasks), Pass@1, as of June 2026. GPT-5.4 (xHigh) leads at 59.1%, 4.1 points ahead of Meta's Muse Spark (55.0%) and 7.2 ahead of the best Claude run (Opus 4.6 thinking, 51.9%)." The same publisher's sibling page already reflects the newer board: "On Scale's standardized SWE-bench Pro board the newest frontier models are not run yet, so the top entries there are Muse Spark 1.1 (61.50%) and gpt-5.4 (59.10%). Scale SEAL public set (standardized scaffolding, July 2026): Muse Spark 1.1 61.50%..." One publisher, two pages, two different leaders.
The claim's "same harness" premise is broadly faithful to how the board works. Scale states it ran frontier models on Pro using the SWE-Agent scaffold , and third-party analysis describes the standardized board as the place where "Scale's standardized SEAL public board is the only place every model runs the same harness" . Run conditions are published: results are described as initial runs subject to change pending official announcement, models are run with uncapped cost and a turn limit of 250, and the ± column is a 95% binomial confidence interval over 730 problems .
Separately, the benchmark's standing has deteriorated since the claim was written. OpenAI said it audited the Scale AI-developed benchmark and found a series of issues , and retracted earlier support after the audit found almost 30% of its tests were "broken" . OpenAI's own announcement states it found 30% of SWE-Bench Pro tasks to be broken and is retracting its previous recommendation that the research community use it as a leading coding eval, noting that some correct solutions fail because of hidden requirements, contradictory instructions, overly strict tests, or incomplete grading criteria . The audit used model-based investigator agents alongside independent reviews from five experienced software engineers .
Findings
✓ What's accurate 6
- The number is real and correctly attributed. GPT-5.4 (xHigh) is listed at 59.10 on Scale's standardized SWE-bench Pro public-set board.
- GPT-5.4 (xHigh) is still co-ranked first on that board under the board's own confidence-interval tie rule.
- The claim correctly distinguishes the standardized board from vendor-scaffold numbers, which is the distinction most citations of "SWE-bench Pro" get wrong.
- "Same harness for every model" is an accurate description of what the standardized board does for the models it contains.
- Calling it a SEAL board is loose but not wrong. SWE-bench Pro is a Scale AI evaluation product, hosted alongside the SEAL leaderboards; the artifact of record is titled the SWE-Bench Pro Public Dataset leaderboard.
- The claim was accurate as a description of the June 2026 board, where GPT-5.4 (xHigh) led at 59.1% ahead of Muse Spark at 55.0%.
≈ What's misleading 5
- **Date context mismatch:** the claim presents a June 2026 board state as the present one. As of 2026-08-12 the board's highest public-set score is Muse Spark 1.1 at 61.50, added after Meta's July 9 2026 release. "The highest score any model reaches" is no longer the board's arithmetic.
- **Omitted qualifier:** "any model" means "any model Scale has run." The standardized board contains no entry for Claude Fable 5, Opus 4.8, Opus 5, GPT-5.5, the GPT-5.6 line, GLM-5.2 or Kimi K3. The claim's universal quantifier describes a partial and lagging roster, not the model population.
- The claim treats a rank as a clean win when the board itself does not. GPT-5.4's ±3.56 interval overlaps Muse Spark 1.1's, and the board's numbering ties them at rank 1. Reading 59.1 as a distinct summit reads more precision off the board than the board publishes.
- **Temporal overreach:** the underlying runs are labelled by the benchmark operator as initial and subject to change, and the claim converts that provisional snapshot into a settled fact about model capability.
- Not a defect in the claim itself, but a change in what it means: since the claim was written, the frontier lab whose model it flatters audited this benchmark, found roughly 30% of the public tasks broken, and withdrew its recommendation of it. A leader-of-the-board statement about this set now carries much less weight than it did in June.
? What's uncertain 5
- Several top entries, including both Muse Spark rows, gpt-5.4 (xHigh), claude-opus-4-6 (thinking) and gemini-3.1-pro (thinking), carry an asterisk that older entries such as claude-opus-4-5-20251101 do not. I read the board's ranking table through the search index's rendering of the page and could not retrieve the footnote legend, so I cannot say what the asterisk denotes. If it marks provider-submitted or differently scaffolded runs, it would qualify the claim's "every model runs the same harness" premise directly. I am not asserting that it does.
- Whether the Muse Spark 1.1 entry of 61.50 is a Scale-run figure or a submitted one. Meta's own launch materials also report 61.5 on SWE-Bench Pro, an exact coincidence I could not resolve either way.
- The scaffold's precise identity. Scale's page text says SWE-Agent, while third-party descriptions and some Scale Labs board labels reference mini-SWE-agent. Which variant produced the 59.10 figure is not pinned.
- Whether Scale has revised, re-run, or annotated the public board in response to OpenAI's broken-task findings.
- The exact date Muse Spark 1.1 was added to the board, as distinct from its July 9 2026 model release date.
Sources
11 of 11 linked to recordsScale Labs SWE-Bench Pro Public Dataset leaderboard, the board of record for this claim, current ranking table
Scale Labs leaderboard index, showing the same public-set top three
SWE-Bench Pro project page and harness notes (turn limit, cost, CI construction)
SWE-Bench Pro paper, Deng et al., arXiv:2509.16941
Scale AI launch blog for SWE-Bench Pro, Sept 19 2025, set composition
OpenAI, "Separating signal from noise in coding evaluations," July 8 2026, plus OpenAI's own post announcing it
The Stack, OpenAI retraction reporting, July 9 2026
Morph LLM, "SWE-bench Pro Leaderboard (2026)," updated Aug 10 2026, the apparent origin of the claim's wording
Morph LLM, "Best AI Model for Coding (2026)," ~July 2026, same publisher, contradicting figure
Digital Applied, "SWE-bench in 2026," June 16 2026, carries the claim's sentence nearly verbatim
Kingy.ai and Digital Applied coverage of Muse Spark 1.1 release, July 9 2026