§ Claim under review · Benchmark
"Anthropic's Frontier Red Team tested several AI models on 100 cyber exploitation tasks, finding GLM-5.3 successfully executed control flow hijacks in 4% of cases, Claude Mythos Preview did so in 6%, while earlier models like Claude Opus 4.6 and GLM-5.2 failed completely."
Verdict
Mostly accurate
Confidence
HighSummary
This post accurately repeats real numbers from a real Anthropic research publication. On September 29, 2026, Anthropic's Frontier Red Team published an analysis of Z.ai's open-weight model GLM-5.3, reporting that on 100 randomly selected tasks from an internal binary exploitation benchmark, GLM-5.3 achieved a full control-flow hijack in 4% of trials and Anthropic's own Claude Mythos Preview did so in 6%, while the earlier Claude Opus 4.6 and GLM-5.2 achieved none. What the post leaves out is the context around those figures. The tests ran inside sealed laboratory environments against offline targets Anthropic had set up, not against real systems, and the benchmark is Anthropic's own and has not been published, so no outside group can check the result. Anthropic also makes one of the models being compared, so these are an interested party's self-run numbers rather than an independent evaluation. The gap between 4% and 6% amounts to roughly two successes out of a hundred and is too small to show a real difference between the two models. A separate US government assessment of GLM-5.3 reached a broadly similar conclusion about the model's cyber abilities, but it used different tests and did not reproduce these particular figures.
The readings
key figures from the evidenceGLM-5.3 full control-flow hijacks on 100 binary exploitation tasks
Claude Mythos Preview full control-flow hijacks, same internal test
Why this verdict
Evidence
Anthropic's Frontier Red Team published a research post on September 29, 2026 analysing Z.ai's open-weight GLM-5.3. The post describes two automated evaluations. On ExploitBench, which targets known vulnerabilities in Chrome's V8 engine, Anthropic reports GLM-5.3 building end-to-end exploits in 50 of 410 attempts and Claude Mythos Preview in 56 of 410. Separately, on what Anthropic calls its internal Binary Exploitation benchmark, which tests whether models can find and exploit vulnerabilities in open source projects participating in Google's OSS-Fuzz, full credit is awarded only for a full control-flow hijack. The post states that it evaluated several models on 100 tasks selected at random from that benchmark, and found GLM-5.3 developing full control-flow hijacks in 4% of trials and Claude Mythos Preview in 6%. The post adds that although GLM-5.3 performs below Claude Mythos Preview on this measure, earlier models such as Claude Opus 4.6 and GLM-5.2 do not succeed in any of them. Anthropic states all tested models were run in isolated, sandboxed environments attacking only offline targets it had set up. The post also references a separate NIST CAISI assessment of September 17, 2026, which Anthropic says found GLM-5.3 to be the most cyber-capable open-weight model released to date while lagging the US frontier by about four months, and says its own capability findings broadly match CAISI's.
Findings
✓ What's accurate 7
- Anthropic's Frontier Red Team did publish this research, on September 29, 2026, under the title "GLM-5.3 and the spread of advanced cyber capabilities". The post exists on Anthropic's own site and is listed on its Frontier Red Team publications page.
- The post does describe an evaluation of several models on 100 tasks, and reports that GLM-5.3 developed full control-flow hijacks in 4% of trials.
- The 6% figure for Claude Mythos Preview matches the post.
- The post does name Claude Opus 4.6 and GLM-5.2 as earlier models that did not succeed on any of those 100 tasks.
- The models named in the claim are real. GLM-5.3 is an open-weight model from Z.ai, formerly Zhipu AI. Claude Mythos Preview is an Anthropic model that the post says was released about five months earlier only to vetted cyber defenders through a programme it calls Project Glasswing.
- The claim's closing interpretation, that this indicates a significant advancement in AI-driven cyber capabilities, tracks Anthropic's own wording that a meaningful threshold has clearly been crossed.
- The attribution chain is sound. Simon Willison's site carried the quote, and the Instagram caption correctly labels it as quoting the Anthropic Frontier Red Team rather than presenting it as Willison's own finding.
≈ What's misleading 3
- The post says "100 cyber exploitation tasks" without the conditions attached to them in the source. The source specifies 100 tasks selected at random from an internal benchmark built on open source projects in Google's OSS-Fuzz, scored only on a full control-flow hijack, with every model run in an isolated sandbox against offline targets Anthropic had set up. The phrase "successfully executed control flow hijacks" can read to a general audience as attacks carried out against real systems. The source describes contained laboratory runs.
- "failed completely" compresses a narrower statement. The source says Claude Opus 4.6 and GLM-5.2 did not succeed on any of these 100 tasks at the full control-flow-hijack bar. Failing to reach that specific bar is not the same as total failure on the tasks, since the benchmark awards full credit only at that level and partial progress is not reported in the claim.
- The evaluator is correctly named, but the reader is not told that Anthropic produced every number in the comparison, including the number for a competitor's model and for its own, using a benchmark it has not released. For a comparison claim, that makes these an interested party's self-reported figures rather than an independent result, and no outside party can check them.
? What's uncertain 4
- Whether 4% and 6% mean four and six of the 100 tasks, or 4% and 6% of a larger pool of trials. The source says "100 tasks" and "4% of the trials" without stating how many trials were run per task. Some outlets rendered it as 4% of 100 tasks, which may be a simplification of the source rather than a figure the source gives.
- Whether the two-point gap between GLM-5.3 and Claude Mythos Preview reflects any real difference in capability. With this few successes it is within the range that could arise from chance.
- Whether the 4% and 6% figures hold up under another evaluator. The benchmark is internal and unpublished, so no independent reproduction exists or is currently possible. The separate NIST CAISI assessment used different benchmarks and scoring and reached a broadly similar capability conclusion, but it is not a check on these specific numbers.
- Z.ai's position on the findings. No public response from Z.ai was found as of 2026-10-07.
Sources
4 of 8 linked to recordsAnthropic Frontier Red Team, "GLM-5.3 and the spread of advanced cyber capabilities", Sep 29 2026 (Fasano, Fleischer, McFaul, Xiao, Gallagher)
Anthropic Frontier Red Team publications index listing the Sep 29 2026 post
Simon Willison's Weblog, quote post of Sep 29 2026
NIST CAISI assessment of GLM-5.3 cyber capabilities, Sep 17 2026, as described in Anthropic's post and in secondary coverage
The Next Web, "Anthropic says China's GLM-5.3 nearly matches Mythos at cyber exploits"
Tom's Hardware report on the Frontier Red Team report
Gigazine, Trending Topics, Business Standard, Mixed News coverage
ExploitBench paper, arXiv 2605.14153