Anthropic · 2026-09-29 · major
Anthropic tests GLM-5.3 — its safeguards fall to simple tricks up to 100%
Anthropic's report on Z.ai's open-weight GLM-5.3 finds it builds end-to-end exploits almost as often as Claude Mythos Preview, and that simple tricks bypass its safeguards 64% to 100% of the time in simulated tests.

Anthropic says an open model now matches its restricted cyber model closely, but anyone can remove GLM-5.3's safety limits.
Quick facts
| Published by | Anthropic, September 29, 2026 |
|---|---|
| Model tested | GLM-5.3 (Z.ai, open weights) |
| ExploitBench | 12% (50 of 410) vs 14% for Claude Mythos Preview |
| Safeguard bypass | 64% deceptive prompt, 92% prefill, 100% abliterated |
| Abliteration cost | About 2,200 GPU hours (~$4,400) |
What is it?
Anthropic published a security study of GLM-5.3, the open-weight coding model from Z.ai (Zhipu AI). The report finds GLM-5.3 can build working end-to-end cyber exploits at a rate close to Claude Mythos Preview, Anthropic's restricted cyber model. In human-directed tests, researchers used GLM-5.3 to find unknown bugs in a web browser's JavaScript engine and chain them into an exploit that stole files.
How does it work?
On ExploitBench, GLM-5.3 produced end-to-end exploits in 50 of 410 attempts (12%), against 14% for Mythos Preview; on a binary-exploitation benchmark it reached 4% against 6%. Earlier models such as Claude Opus 4.6 and GLM-5.2 scored near zero on both. Anthropic then tested the safeguards: bare malicious requests were refused, but a false cover story worked 64% of the time, prefilled reasoning tokens 92%, and an abliterated copy 100%. Abliteration took about 2,200 GPU hours (~$4,400) and cut refusals from 95% to 6% while GPQA-Diamond scores stayed the same.
Why does it matter?
Frontier-level exploit building is now in a model anyone can download, and Anthropic notes that several abliterated versions were public within days of release. The report follows NIST CAISI's September 17 finding that GLM-5.3 is the most cyber-capable open-weight model so far, about four months behind the US frontier. Anthropic calls for government safety testing of GLM-5.3's successors and says it will widen defender access to Claude's cyber tools — security teams should assume attackers have these capabilities today.
Who is it for?
security teams, AI policy researchers, open-model developers
Frequently asked questions
- How does GLM-5.3 compare to Claude Mythos Preview at writing exploits?
- Anthropic's GLM-5.3 report puts the two models close together. On ExploitBench, GLM-5.3 built end-to-end exploits in 12% of attempts (50 of 410) against 14% for Claude Mythos Preview. On a binary-exploitation benchmark for control-flow hijacks, GLM-5.3 scored 4% and Mythos Preview 6%. Claude Opus 4.6 and GLM-5.2 were near zero on both tests.
- Why did Claude's safeguards hold when GLM-5.3's did not?
- Anthropic says Claude's safeguards blocked the deceptive-prompt requests that got GLM-5.3 to help 64% of the time. The Anthropic API also gives users no way to prefill Claude's thinking, which pushed GLM-5.3 to 92% compliance. And because GLM-5.3's weights are public, anyone can abliterate it, which reached 100% compliance in Anthropic's tests.
- What does NIST CAISI say about GLM-5.3?
- NIST's Center for AI Standards and Innovation (CAISI) published its own GLM-5.3 assessment on September 17, 2026. CAISI found that GLM-5.3 is the most cyber-capable open-weight model released to date, but that its cyber capabilities are significantly lower than current US frontier models, lagging by about four months on an aggregate of CAISI's cyber benchmarks.
- What does Anthropic recommend after the GLM-5.3 findings?
- Anthropic's GLM-5.3 report makes three asks. Anthropic says it is working to safely expand access to Claude's cyber capabilities for as many defenders as possible. It argues that governments should run safety tests on sufficiently capable models, including GLM-5.3's successors. And it asks open-weight developers worldwide to safeguard these capabilities and prevent misuse.