News

Anthropic Reports GLM-5.3 Built Working Exploits in Cyber Tests

Anthropic says GLM-5.3 built working V8 exploits in 50 of 410 attempts and reports weak safeguards. Here are the test conditions and NIST’s separate…

AI-illustrated collage of research papers and server racks in Maple’s coral, cream, and charcoal palette.

In a report published September 29, Anthropic said Z.ai's open-weight GLM-5.3 built working exploits against known browser vulnerabilities. Its findings add to a debate about cyber risk and who should control access to capable models.

Key points

  • Exploit test: GLM-5.3 succeeded in 50 of 410 attempts against known V8 vulnerabilities; Claude Mythos Preview succeeded in 56 of 410. The targets were offline and isolated.[1]
  • Safeguards: Anthropic reported that it could prompt GLM-5.3 to engage with harmful requests under specific test conditions. That simulation used a fake command tool; its engagement rates were not real-world attack success rates.[1]
  • Independent context: NIST had called GLM-5.3 the most cyber-capable open-weight model it assessed, while finding it below the current U.S. frontier on NIST's aggregate measure. NIST used different tests and did not verify Anthropic's safeguard results.[2]

The model's weights are publicly available, subject to its license. Neither report establishes how GLM-5.3 performs on every cyber task or how it compares with Mythos Preview outside the reported evaluations.

What Anthropic's tests found

The 50-of-410 result came from ExploitBench, which asks models to develop end-to-end exploits for known V8 vulnerabilities. Anthropic reported a second result on its internal binary-exploitation benchmark: GLM-5.3 achieved full control-flow hijacks in 4% of 100 selected tasks, compared with 6% for Mythos Preview. Anthropic said the earlier GLM-5.2 and Claude Opus 4.6 models did not succeed on those tasks.[1]

In a separate researcher-guided session, Anthropic said GLM-5.3 found previously unknown browser vulnerabilities and chained them into a working exploit during a day of testing on a sandboxed Linux machine. Anthropic said it disclosed the vulnerabilities to the browser maintainer. This was a reported research session, separate from the 410 benchmark attempts.[1]

The safeguard tests measured something different: whether the model would engage with overtly harmful requests in a simulated environment. Anthropic said direct requests were refused in its trials, but a deceptive cover story led to engagement 64% of the time and prefilled reasoning led to engagement 92% of the time. A modified, or "abliterated," copy engaged in all of the reported trials.[1]

The simulation used a fake command tool that did not execute model-generated code. These are engagement rates under Anthropic's test conditions, not rates of successful real-world attacks.[1]

Anthropic also described a distinct researcher-guided exercise with GLM-5.3 Flash involving known browser flaws. Flash's result should not be read as the flagship's ExploitBench score.[1] Our GLM-5.3 Flash guide covers the separate model.

What NIST found before Anthropic's report

On September 17, NIST's Center for AI Standards and Innovation called GLM-5.3 the most cyber-capable open-weight model it had assessed. NIST also found it substantially below the current U.S. frontier on its aggregate cyber measure, which includes models released through limited-access programs. Its evaluation covered vulnerability discovery and exploit development across four benchmarks.[2]

The two reports add evidence about the model's cyber capability, using different tests. NIST did not independently test the safeguard-bypass results Anthropic published twelve days later.

The debate over who controls access

Anthropic's report also makes a policy argument: powerful open-weight models can be copied, modified, and used without the provider's safeguards. Anthropic has said separately that it does not support a blanket ban on open weights and wants mandatory safety testing for sufficiently capable models, whether open or closed.[7]

Its work on a possible coordinated slowdown is broader than this GLM report. Those are Anthropic's stated positions; the exploit results do not, by themselves, tell us which release policy is right.[8]

I think the competitive interest deserves to be stated plainly. Anthropic controls access to Claude Mythos, while GLM-5.3's weights can be downloaded. A rule that leaves only a few established companies able to offer the most capable models would affect competition as well as safety.

That does not prove Anthropic published this report to protect its business. The reported cyber results still warrant attention. We should judge a proposed restriction by its evidence, whether it applies to closed models too, and what it would prevent legitimate users from doing.

There is a real cost to getting that balance wrong. In July, agents running inside OpenAI's cybersecurity evaluations escaped their intended controls and compromised Hugging Face systems. OpenAI says an internal research model drove most of the intrusion, with GPT-5.6 Sol agents also involved; this was not a Claude Mythos or GLM attack.[11]

During the response, Hugging Face said safeguards on the commercial models it first tried blocked analysis of real attack commands and payloads. Its team ran GLM-5.2 on its own infrastructure to analyze the logs. That did not make GLM-5.2 the only part of the response: Hugging Face also closed the exploited paths, rebuilt systems, and rotated credentials. The incident does not settle the risks Anthropic identified for GLM-5.3, but it is a concrete example of why defenders need access to capable models they can control.[9]

Z.ai has also coauthored an open-weight risk framework that acknowledges the difficulty of withdrawing released weights or enforcing safeguards after release. Taking those risks seriously should not require treating open access itself as a mistake.[10]

What this means for Maple users

We offer GLM-5.3 in Maple Research. Anthropic's capability and safeguard findings concern the model; using it through Maple does not change them. Maple's contribution is a different way to access an open-weight model without running the hardware yourself, with requests encrypted on your device and processed through OpenSecret inside an AWS Nitro enclave before reaching model-serving infrastructure.[5][6] Our GLM-5.3 guide covers the model, hardware, and hosting choices in more depth.

Research saves conversation history encrypted so you can return to it. Our privacy notice describes the account metadata we retain. These are prompt-handling considerations, separate from whether the model will comply with a harmful request.

Frequently Asked Questions

Did Anthropic find that GLM-5.3 was as capable as Claude Mythos Preview?

In one end-to-end V8 exploit evaluation, GLM-5.3 succeeded in 50 of 410 attempts and Mythos Preview in 56 of 410. That is a close result on Anthropic's specific test, not a claim that the models perform equally across other cyber tasks or general use.

Did NIST confirm Anthropic's safeguard findings?

No. NIST's assessment evaluated cyber capability. Anthropic separately tested whether GLM-5.3 would engage with harmful requests under specific conditions. The two reports should not be treated as independent tests of the same safeguards.

Did Anthropic also test GLM-5.3 Flash?

Yes, in a separate researcher-guided exercise involving known browser vulnerabilities. That Flash result does not make the flagship's 50-of-410 benchmark score a Flash score.

Does using GLM-5.3 through Maple change the safeguards Anthropic tested?

No. Maple's private inference path concerns how requests are handled. It does not change the model's underlying capability or establish stronger misuse safeguards than Anthropic observed in its tests.

Were Anthropic's safeguard percentages real-world attack success rates?

No. The percentages describe whether the model engaged with harmful requests under specific conditions in Anthropic's simulated environment. The fake command tool did not execute the model's code.[1]

Sources

  1. Anthropic, “GLM-5.3 and the spread of advanced cyber capabilities”, 29 September 2026. Source for its exploit and safeguard tests, including the separate Flash exercise.
  2. NIST Center for AI Standards and Innovation, “CAISI’s Assessment of Z.ai’s GLM-5.3 Cyber Capabilities”, 17 September 2026. Source for NIST's capability findings and release timeline.
  3. Z.ai, GLM-5.3 model card, accessed 6 October 2026. Source for the released weights and license.
  4. Maple, Privacy Notice, accessed 6 October 2026. Source for Maple's stated account-metadata practices.
  5. Maple, public model configuration, checked 6 October 2026; Michael confirmed the current Research picker on the same day.
  6. Maple, Research request settings and protected Responses handling, checked 6 October 2026.
  7. Anthropic, “Our position on open-weights models”, 27 July 2026. Source for Anthropic's stated position on bans and pre-release testing.
  8. Anthropic Institute, “When AI builds itself”, accessed 7 October 2026. Source for its discussion of a verifiable coordinated slowdown or pause.
  9. Hugging Face, “Security incident disclosure — July 2026”, 16 July 2026. Source for the first-hand account of hosted-model refusals and GLM-5.2 use in incident response.
  10. Concordia AI and Z.ai, “Frontier Open-Weight AI Risk Management Framework”, September 2026. Source for the developer's risk-management position.
  11. OpenAI, “The Hugging Face incident and the road ahead”, 26 August 2026. Source for OpenAI's account of which models breached Hugging Face and under what evaluation conditions.