Article
Anthropic: open-weight GLM-5.3 builds exploits, lacks safeguards
Anthropic says Z.ai's open-weight GLM-5.3 nearly matches Claude Mythos Preview at building cyber exploits, and its safeguards are easy to bypass.
Anthropic's Frontier Red Team says an open-weight model from Zhipu AI, known outside China as Z.ai, can build working cyber exploits nearly as well as Anthropic's own restricted Claude Mythos Preview. According to Anthropic's analysis, published Sept. 29, 2026, GLM-5.3 was released "without meaningful safeguards to limit misuse," and anyone can download it. If the findings hold up, attackers now have free access to a capability that, until recently, only vetted defenders had.
What happened
Five months ago, Anthropic announced Claude Mythos Preview, which it describes as the first AI model that could autonomously build sophisticated, end-to-end cyber exploits. The company says it released that model in a limited way through Project Glasswing. It says trusted defenders used it to find more than 10,000 vulnerabilities in critical software before attackers could get similar tools.
Anthropic's new post argues that head start is over. "But those models have now arrived," the authors write, pointing to GLM-5.3 as the example.
Anthropic's results come from automated benchmarks and from human experts using the model. It says all testing took place in isolated, sandboxed environments against offline targets the team set up for the evaluations.
The benchmark numbers
On ExploitBench, which tests how well models exploit known vulnerabilities in Chrome's V8 engine, Anthropic reports that GLM-5.3 produced end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview did so in 56 of 410.
On Anthropic's internal Binary Exploitation benchmark, built from open-source projects in Google's OSS-Fuzz program, the company says GLM-5.3 achieved a full control-flow hijack in 4% of trials on 100 randomly chosen tasks. Mythos Preview reached 6%. According to Anthropic, earlier models such as Claude Opus 4.6 and GLM-5.2 succeeded on none of them.
Anthropic also had human experts use GLM-5.3 against targets with no known vulnerabilities. These sessions typically lasted a day or less, with under an hour of human focus. In one example, the company says the model chained together multiple 0-day vulnerabilities it found in a component of a popular web browser. The resulting exploit page stole a user's SSH private key through a malicious website.
The safeguards problem
Anthropic's central argument is about access more than raw capability. The company says that in its simulated tests, attackers could bypass GLM-5.3's safeguards between 64% and 100% of the time using simple techniques. It says the same attacks did not succeed against safeguarded Claude models.
The comparison has a catch. Anthropic notes that the Claude benchmark results come from models run with safeguards disabled, and that its reduced-safeguard versions are limited to vetted users. GLM-5.3, by contrast, is open-weight, so anyone can download and run it.
Anthropic concludes that GLM-5.3's "lax safeguards significantly increase the cyber capabilities available to malicious actors." It also acknowledges that the same capabilities can help defenders secure their own systems.
An independent check from NIST
Anthropic is not the only one making this assessment. On Sept. 17, NIST's Center for AI Standards and Innovation (CAISI) published its own evaluation of GLM-5.3. CAISI called it "the most cyber-capable open-weight model released to date" and estimated that it trails the US frontier by about four months on an aggregate of its cyber benchmarks.
Anthropic says its capability findings broadly match CAISI's. Its post adds a separate question: how easily GLM-5.3's safeguards can be bypassed or removed.
Why it matters
For years, the most capable offensive-security models have sat behind lab-controlled access. Anthropic's findings suggest that gap has nearly closed for exploit development, at least on the tasks it measured. A freely downloadable model with weak guardrails is much harder to control than one served through an API.
For security teams, the practical takeaway is to assume attackers may have AI that can help write working exploits. For policymakers, it raises questions about how open-weight releases should be evaluated before they ship. The four-month lag that CAISI estimated also suggests these capabilities spread quickly.
What is still unknown
Most of the headline numbers come from Anthropic, a direct competitor to Z.ai, and some come from an internal benchmark that outsiders cannot reproduce. The sources reviewed here include no response from Z.ai, and they do not show how Z.ai describes GLM-5.3's safeguards or intended uses.
It is also unclear how often these capabilities succeed in real-world attacks rather than sandboxed tests. The benchmark success rates, roughly 12% on ExploitBench and 4% on Binary Exploitation, show real capability but not reliability. Community discussion has started on Reddit's r/LocalLLaMA and Hacker News, but independent replication has not yet been reported.