TrueSeeker AI · Verified claim report Case 3013b5743a · 2026-08-15

§ Claim under review · Benchmark

"GLM-5.3 uses the same base model as GLM-5.2, with performance gains coming almost entirely from post-training: Z.ai Code Bench improved from 20.9% to 31.4%, Terminal-Bench 3.0 from 4.6% to 28.3%, DeepSWE from 46.2% to 66.9%, and AutomationBench from 26.2% to 48.2%."

Circulating claim, as submitted.

Verdict

Mostly accurate

Confidence

Medium
§

Summary

The numbers in this post are accurately copied from Z.ai's own GLM-5.3 launch materials, published 14 August 2026. Z.ai does say GLM-5.3 reuses the GLM-5.2 base model unchanged and that all gains come from expanded post-training, and Bloomberg reported the same. The four score pairs check out against Z.ai's documentation and launch table. The important missing context is that Z.ai ran all four of these evaluations itself, inside its own harness, and no independent party has reproduced any GLM-5.3 result, partly because the model weights are being held back for about two weeks over cybersecurity concerns. Two smaller caveats: the Z.ai Code Bench figures of 20.9 to 31.4 are the "High effort" setting on a chart with several settings, and that benchmark is private and cannot be rerun by anyone outside the company. The Terminal-Bench jump also starts from a near-zero baseline of 4.6, which makes the multiple look larger than the underlying progress. Z.ai's same table also shows GLM-5.3 losing to GPT-5.6 Sol and Claude Fable 5 on several of these benchmarks, which the post does not mention.

§

The readings

key figures from the evidence
4.6 to 28.3 %

Terminal-Bench 3.0, Z.ai self-run, near-floor baseline

46.2 to 66.9 %

DeepSWE v1.1, Z.ai's own harness, vs public leaderboard 44%

20.9 to 31.4 %

Z.ai Code Bench, High-effort pair only, private benchmark

§

Why this verdict

All four score pairs and the same-base-model premise match Z.ai's own launch materials as of 2026-08-15, and the claim's "almost entirely" is slightly more cautious than Z.ai's own "all we did was scale post-training." I considered and rejected "Accurate," because the figures are entirely vendor-run and the Code Bench pair silently picks the High-effort point on a multi-point curve. I also considered and rejected "Source exists but framing is misleading" and "Partially accurate but misleading," because no cited source contradicts or bounds away the operative proposition: the numbers are transcribed correctly, the comparison is GLM-5.3 versus GLM-5.2 rather than a manufactured win over rivals, and the post's own caption labels Code Bench as internal and attributes the cyber figures to Z.ai. Confidence is capped at Medium because a benchmark claim resting only on vendor-run numbers cannot go higher, no independent reproduction of any GLM-5.3 figure exists, and I read the vendor's table through documentation and reporting rather than by loading the launch blog itself. ---
§

Evidence

Z.ai released GLM-5.3 on 2026-08-14. Z.ai's own developer documentation states, in the vendor's words, that Terminal-Bench 3.0 increased from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and Agents' Last Exam from 23.8 to 28.5, and that GLM-5.3 improves by 50% over GLM-5.2 on Z.ai Code Bench.

Multiple named outlets report the same table. VentureBeat records the AutomationBench pair as 26.2 to 48.2 and the same Terminal-Bench and DeepSWE pairs. Trade coverage records the Z.ai Code Bench pair as 20.9 to 31.4 at High effort, with a separate Max-effort pair of 23.4 to 34.5, and a token-efficiency claim of roughly 50,000 output tokens per task at High effort versus roughly 120,000 for Claude Opus 4.8 at 29.5%.

On the architectural claim, Bloomberg reports that GLM-5.3 is built on the same roughly 700-billion-parameter base as its predecessor, and Z.ai's launch text is quoted by The Stack and Decrypt as saying that scaling post-training is all the company did for this release. Reported base specification is a 743B-parameter mixture-of-experts model with roughly 40B active parameters, unchanged from GLM-5.2.

Every figure in the claim is Z.ai-run. The announcement footnotes, mirrored on OpenLM.ai, show that CyberGym, ExploitGym and the coding rows were evaluated by Z.ai inside Claude Code 2.1.207 at max reasoning effort, temperature 1.0, top_p 1.0, 128K max new tokens, single-run pass@1, unlimited timeout, with a domain whitelist. AutomationBench was run on v1.0.6 with a specific PR fix applied. Only GDPval-AA v2 (evaluated by Artificial Analysis) and Toolathlon Verified (official evaluation service) were scored outside Z.ai, which is two rows out of roughly nineteen.

One provenance discrepancy is documented: the independent DeepSWE leaderboard, which runs all models on mini-swe-agent, listed GLM-5.2 at 44% plus or minus 2 at the time of the launch, not the 46.2 that appears in Z.ai's table. Z.ai's footnote also uses mini-swe-agent but specifies temperature 0.95, six-hour timeouts, 400K context and its own run.

Weights are not public. Axios reports Z.ai is delaying the public release of the model weights for two weeks while it tests and strengthens safety controls, tied to the model's vulnerability-finding capability. As of the claim date no third party has run GLM-5.3 independently, because only API and Coding Plan access exist.


§

Findings

✓ What's accurate 6

  • Z.ai does state that GLM-5.3 uses the same base model as GLM-5.2 and that the gains come from scaled post-training. Bloomberg independently reports the same-base-model characterization, sourced to the company.
  • Terminal-Bench 3.0 4.6 to 28.3 and DeepSWE v1.1 46.2 to 66.9 appear verbatim in Z.ai's own developer documentation.
  • AutomationBench 26.2 to 48.2 matches Z.ai's published table as reported by VentureBeat and other outlets.
  • Z.ai Code Bench 20.9 to 31.4 matches Z.ai's High-effort figures and is consistent with the vendor's stated "50% improvement" headline.
  • The claim's wording "almost entirely from post-training" is, if anything, weaker than what Z.ai asserts. The vendor says the gain is entirely from post-training. The claim does not overstate the vendor here.
  • GLM-5.3 exists and was released on 2026-08-14 through the Z.ai API and GLM Coding Plan. This is not a fabricated announcement.

≈ What's misleading 4

  • **Marketing as evidence:** the claim presents four benchmark deltas as results, without stating that Z.ai produced all four itself inside its own harness. Only two rows in the roughly nineteen-row launch table were scored by named third parties, and none of the four cited here is among them. A reasonable reader takes these as measured facts about the model rather than as an interested party's self-report.
  • **Omitted qualifier:** the Z.ai Code Bench pair 20.9 to 31.4 is the High-effort pair specifically. The same chart also shows a Max-effort pair of 23.4 to 34.5 and a Low-effort pair. The claim reports one point on a curve as if it were the benchmark result. It also omits that Z.ai Code Bench is private and cannot be rerun by anyone outside Z.ai.
  • **Benchmark cherry picking:** the four rows selected are among the largest jumps in the table. The Terminal-Bench 3.0 figure rises from a near-floor 4.6, which mechanically produces the most dramatic multiple in the set. The same official table shows GLM-5.3 trailing GPT-5.6 Sol and Claude Fable 5 on Terminal-Bench 3.0 (28.3 versus 34.6 and 33.7), on DeepSWE v1.1 (66.9 versus 72.7 and 69.7), and badly on ExploitBench and ExploitGym. Nothing in the claim is false because of this, but the selection makes the release look like a clean sweep when the vendor's own chart does not.
  • **Harness mismatch:** the DeepSWE baseline of 46.2 is Z.ai's own run under its own settings. The independent DeepSWE leaderboard listed GLM-5.2 at 44 plus or minus 2 using mini-swe-agent with standard settings. Same benchmark name, different evaluation record. The delta is therefore internally consistent within Z.ai's harness and is not directly comparable to the public board.

? What's uncertain 4

  • Whether the base model is genuinely unchanged cannot be independently verified. The GLM-5.3 weights have not been released, so the "same base model" statement rests entirely on Z.ai's word. This is the load-bearing premise of the whole post and it is currently unauditable.
  • No independent reproduction of any GLM-5.3 score exists as of 2026-08-15. Third parties have API access only, and the weights are held back for roughly two weeks pending safety hardening.
  • I could not load the Z.ai launch blog page itself. The vendor figures here were read from Z.ai's developer documentation and from a mirror of the announcement footnotes, plus consistent reporting across named outlets. The AutomationBench and Z.ai Code Bench pairs specifically were confirmed through reporting rather than through a vendor page I retrieved directly.
  • The claim writes all four figures as percentages. GDPval-AA v2 is an Elo score and ExploitGym is a task count, so the post's wider table mixes units, but the four cited figures do appear to be percentage metrics.
Distortion flags marketing as evidence harness mismatch omitted qualifier benchmark cherry picking
§

Sources

10 of 11 linked to records
[1]

Z.ai developer documentation, GLM-5.3 overview page (vendor channel of record) | primary for what the vendor asserts, interested-party self-report for quality/comparison | vendor | https://docs.z.ai/guides/llm/glm-5.3

unknown
This citation could not be independently verified.
[2]

Z.ai launch blog, "GLM-5.3: Frontier Coding with Emergent Cyber Capabilities"

unknown vendor
https://z.ai/blog/glm-5.3 ↗
[3]

OpenLM.ai mirror of the GLM-5.3 announcement table and evaluation footnotes

secondary aggregator
https://openlm.ai/glm-5.2/ ↗
[5]

Bloomberg, "Z.ai to Rival Anthropic, OpenAI in Coding With New AI Model"

secondary named-outlet journalism
https://www.bloomberg.com/news/articles/2026-08-14/z-ai-aims-to-catch-anthropic-openai-in-coding-with-new-ai-model ↗
[7]

Kingy AI launch breakdown (effort-level split and DeepSWE leaderboard discrepancy)

secondary trade blog
https://kingy.ai/blog/glm-5-3-specs-benchmarks-api-how-to-use/ ↗
[8]

Decrypt, launch coverage with Max-effort Code Bench figures

secondary named-outlet journalism
https://decrypt.co/375684/china-z-ai-glm-5-3-top-open-weight-coding-model ↗
[9]

Axios, weights-delay coverage

secondary named-outlet journalism
https://www.axios.com/2026/08/14/china-open-source-ai-glm-53 ↗
[10]

Interconnects analysis of the release

secondary named expert commentary
https://www.interconnects.ai/p/glm-53-how-chinese-labs-keep-stride ↗
[11]

digitalapplied analysis (count of third-party-evaluated rows)

tertiary commentary blog
https://www.digitalapplied.com/blog/glm-5-3-launch-post-training-scaling-coding-agents ↗
How links are chosen. A source is linked only when the address comes from the investigation's own retrieval or from a registry lookup (PubMed, Crossref) that matches the citation's title and year. Author lists shown as registry-verified come from the registry record, not from the report text. Citations that cannot be matched are labeled, never guessed.
This is one case on the record See the full case, browse the archive, and search every checked claim on TrueSeeker AI Open on ai.trueseeker.com →