§ Claim under review · Benchmark
"GLM-5.3 uses the same base model as GLM-5.2, with performance gains coming almost entirely from post-training: Z.ai Code Bench improved from 20.9% to 31.4%, Terminal-Bench 3.0 from 4.6% to 28.3%, DeepSWE from 46.2% to 66.9%, and AutomationBench from 26.2% to 48.2%."
Verdict
Mostly accurate
Confidence
MediumSummary
The numbers in this post are accurately copied from Z.ai's own GLM-5.3 launch materials, published 14 August 2026. Z.ai does say GLM-5.3 reuses the GLM-5.2 base model unchanged and that all gains come from expanded post-training, and Bloomberg reported the same. The four score pairs check out against Z.ai's documentation and launch table. The important missing context is that Z.ai ran all four of these evaluations itself, inside its own harness, and no independent party has reproduced any GLM-5.3 result, partly because the model weights are being held back for about two weeks over cybersecurity concerns. Two smaller caveats: the Z.ai Code Bench figures of 20.9 to 31.4 are the "High effort" setting on a chart with several settings, and that benchmark is private and cannot be rerun by anyone outside the company. The Terminal-Bench jump also starts from a near-zero baseline of 4.6, which makes the multiple look larger than the underlying progress. Z.ai's same table also shows GLM-5.3 losing to GPT-5.6 Sol and Claude Fable 5 on several of these benchmarks, which the post does not mention.
The readings
key figures from the evidenceTerminal-Bench 3.0, Z.ai self-run, near-floor baseline
DeepSWE v1.1, Z.ai's own harness, vs public leaderboard 44%
Z.ai Code Bench, High-effort pair only, private benchmark
Why this verdict
Evidence
Z.ai released GLM-5.3 on 2026-08-14. Z.ai's own developer documentation states, in the vendor's words, that Terminal-Bench 3.0 increased from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and Agents' Last Exam from 23.8 to 28.5, and that GLM-5.3 improves by 50% over GLM-5.2 on Z.ai Code Bench.
Multiple named outlets report the same table. VentureBeat records the AutomationBench pair as 26.2 to 48.2 and the same Terminal-Bench and DeepSWE pairs. Trade coverage records the Z.ai Code Bench pair as 20.9 to 31.4 at High effort, with a separate Max-effort pair of 23.4 to 34.5, and a token-efficiency claim of roughly 50,000 output tokens per task at High effort versus roughly 120,000 for Claude Opus 4.8 at 29.5%.
On the architectural claim, Bloomberg reports that GLM-5.3 is built on the same roughly 700-billion-parameter base as its predecessor, and Z.ai's launch text is quoted by The Stack and Decrypt as saying that scaling post-training is all the company did for this release. Reported base specification is a 743B-parameter mixture-of-experts model with roughly 40B active parameters, unchanged from GLM-5.2.
Every figure in the claim is Z.ai-run. The announcement footnotes, mirrored on OpenLM.ai, show that CyberGym, ExploitGym and the coding rows were evaluated by Z.ai inside Claude Code 2.1.207 at max reasoning effort, temperature 1.0, top_p 1.0, 128K max new tokens, single-run pass@1, unlimited timeout, with a domain whitelist. AutomationBench was run on v1.0.6 with a specific PR fix applied. Only GDPval-AA v2 (evaluated by Artificial Analysis) and Toolathlon Verified (official evaluation service) were scored outside Z.ai, which is two rows out of roughly nineteen.
One provenance discrepancy is documented: the independent DeepSWE leaderboard, which runs all models on mini-swe-agent, listed GLM-5.2 at 44% plus or minus 2 at the time of the launch, not the 46.2 that appears in Z.ai's table. Z.ai's footnote also uses mini-swe-agent but specifies temperature 0.95, six-hour timeouts, 400K context and its own run.
Weights are not public. Axios reports Z.ai is delaying the public release of the model weights for two weeks while it tests and strengthens safety controls, tied to the model's vulnerability-finding capability. As of the claim date no third party has run GLM-5.3 independently, because only API and Coding Plan access exist.
Findings
✓ What's accurate 6
- Z.ai does state that GLM-5.3 uses the same base model as GLM-5.2 and that the gains come from scaled post-training. Bloomberg independently reports the same-base-model characterization, sourced to the company.
- Terminal-Bench 3.0 4.6 to 28.3 and DeepSWE v1.1 46.2 to 66.9 appear verbatim in Z.ai's own developer documentation.
- AutomationBench 26.2 to 48.2 matches Z.ai's published table as reported by VentureBeat and other outlets.
- Z.ai Code Bench 20.9 to 31.4 matches Z.ai's High-effort figures and is consistent with the vendor's stated "50% improvement" headline.
- The claim's wording "almost entirely from post-training" is, if anything, weaker than what Z.ai asserts. The vendor says the gain is entirely from post-training. The claim does not overstate the vendor here.
- GLM-5.3 exists and was released on 2026-08-14 through the Z.ai API and GLM Coding Plan. This is not a fabricated announcement.
≈ What's misleading 4
- **Marketing as evidence:** the claim presents four benchmark deltas as results, without stating that Z.ai produced all four itself inside its own harness. Only two rows in the roughly nineteen-row launch table were scored by named third parties, and none of the four cited here is among them. A reasonable reader takes these as measured facts about the model rather than as an interested party's self-report.
- **Omitted qualifier:** the Z.ai Code Bench pair 20.9 to 31.4 is the High-effort pair specifically. The same chart also shows a Max-effort pair of 23.4 to 34.5 and a Low-effort pair. The claim reports one point on a curve as if it were the benchmark result. It also omits that Z.ai Code Bench is private and cannot be rerun by anyone outside Z.ai.
- **Benchmark cherry picking:** the four rows selected are among the largest jumps in the table. The Terminal-Bench 3.0 figure rises from a near-floor 4.6, which mechanically produces the most dramatic multiple in the set. The same official table shows GLM-5.3 trailing GPT-5.6 Sol and Claude Fable 5 on Terminal-Bench 3.0 (28.3 versus 34.6 and 33.7), on DeepSWE v1.1 (66.9 versus 72.7 and 69.7), and badly on ExploitBench and ExploitGym. Nothing in the claim is false because of this, but the selection makes the release look like a clean sweep when the vendor's own chart does not.
- **Harness mismatch:** the DeepSWE baseline of 46.2 is Z.ai's own run under its own settings. The independent DeepSWE leaderboard listed GLM-5.2 at 44 plus or minus 2 using mini-swe-agent with standard settings. Same benchmark name, different evaluation record. The delta is therefore internally consistent within Z.ai's harness and is not directly comparable to the public board.
? What's uncertain 4
- Whether the base model is genuinely unchanged cannot be independently verified. The GLM-5.3 weights have not been released, so the "same base model" statement rests entirely on Z.ai's word. This is the load-bearing premise of the whole post and it is currently unauditable.
- No independent reproduction of any GLM-5.3 score exists as of 2026-08-15. Third parties have API access only, and the weights are held back for roughly two weeks pending safety hardening.
- I could not load the Z.ai launch blog page itself. The vendor figures here were read from Z.ai's developer documentation and from a mirror of the announcement footnotes, plus consistent reporting across named outlets. The AutomationBench and Z.ai Code Bench pairs specifically were confirmed through reporting rather than through a vendor page I retrieved directly.
- The claim writes all four figures as percentages. GDPval-AA v2 is an Elo score and ExploitGym is a task count, so the post's wider table mixes units, but the four cited figures do appear to be percentage metrics.
Sources
10 of 11 linked to recordsZ.ai developer documentation, GLM-5.3 overview page (vendor channel of record) | primary for what the vendor asserts, interested-party self-report for quality/comparison | vendor | https://docs.z.ai/guides/llm/glm-5.3
Z.ai launch blog, "GLM-5.3: Frontier Coding with Emergent Cyber Capabilities"
OpenLM.ai mirror of the GLM-5.3 announcement table and evaluation footnotes
VentureBeat, GLM-5.3 launch coverage
Bloomberg, "Z.ai to Rival Anthropic, OpenAI in Coding With New AI Model"
MarkTechPost launch analysis
Kingy AI launch breakdown (effort-level split and DeepSWE leaderboard discrepancy)
Decrypt, launch coverage with Max-effort Code Bench figures
Axios, weights-delay coverage
Interconnects analysis of the release
digitalapplied analysis (count of third-party-evaluated rows)