← Back to all posts
News

Anthropic Finds GLM-5.3 Nears Mythos on Exploits, Safeguards Stripped for $4,400

September 30, 2026 · 20:06 UTC · News
Anthropic Finds GLM-5.3 Nears Mythos on Exploits, Safeguards Stripped for $4,400

TL;DR

On September 29, Anthropic's Frontier Red Team published an evaluation of a competitor's open-weight model: GLM-5.3 from Z.ai (formerly Zhipu AI). The headline finding is that GLM-5.3 writes working exploits at close to the rate of Claude Mythos Preview, the model Anthropic held back from general release five months ago for exactly that reason. The second finding matters more for anyone thinking about open weights: the red team got GLM-5.3 to carry out simulated attack orders 64% of the time with a cover story, 92% with prefilled reasoning, and 100% after abliterating the weights, a job that cost them about $4,400 in compute.


What Anthropic Tested

GLM-5.3 launched on August 14 and its weights went public about two weeks later, which makes it downloadable by anyone with enough GPUs. Anthropic ran it in isolated, sandboxed environments and compared it with Claude models accessed through the API with safeguards on. The report says no model-generated code was executed in the harmful-scenario simulations.

There were two questions. How good is GLM-5.3 at offensive security work? And how hard is it to make it do that work for someone who should not have it?

Exploit Development: Close to Mythos

On ExploitBench, which measures how well a model can exploit known vulnerabilities in the V8 JavaScript engine inside Google Chrome, GLM-5.3 produced end-to-end exploits in 50 of 410 attempts. Claude Mythos Preview managed 56 of 410. On Anthropic's internal binary exploitation benchmark, GLM-5.3 achieved full control-flow hijacks in 4% of trials versus 6% for Mythos Preview.

The older generation barely registers. Anthropic reports that Claude Opus 4.6, GLM-5.2, Kimi K3 and DeepSeek V4.1-Flash scored near zero on these tasks.

ExploitBench (V8): end-to-end exploits out of 410 attempts Mythos Preview56 GLM-5.350 GLM-5.2, K3near 0 Opus 4.6 and DeepSeek V4.1-Flash also near 0 (Anthropic report)
One generation took open weights from near zero to within six exploits of Mythos Preview.

Anthropic also ran a hands-on case study. A researcher used GLM-5.3-Flash to chain two Chrome flaws, including a recently disclosed one, into a reliable exploit for an ARM64 target that bypassed pointer-authentication (PAC) hardening. It took 20 minutes of human attention and eight hours of model work. At Zhipu's API prices, Anthropic says the run would have cost $20.40. That is less than a lunch for two in San Francisco, which is the point.

How this squares with NIST's numbers

On September 17, NIST's Center for AI Standards and Innovation (CAISI) published its own assessment, calling GLM-5.3 "the most cyber-capable open-weight model released to date" while putting it about four months behind US frontier models in aggregate. On SEC-Bench Pro, CAISI measured GLM-5.3 at 40.4% against 90.2% for the best US model.

The two reports are not in conflict. CAISI compared GLM-5.3 against the best current US systems. Anthropic compared it against Mythos Preview, the model it considered dangerous enough to gate behind Project Glasswing, its trusted-access program for defenders. CAISI's ExploitBench figure (61.1%, or 9.8 of 16 points) also uses a partial-credit score, while Anthropic counted only complete end-to-end exploits. Four months behind the frontier is still level with what the frontier was a few months ago, and that frontier was withheld.

Safeguards: A Speed Bump

The red team built simulated attack scenarios, five attack orders against two targets with five attempts each, 50 samples per cell. Asked straight, GLM-5.3 engaged 0% of the time. Then the attacks:

  • Cover story (64%): tell the model it is an autonomous red-team agent working on an exercise.
  • Prefilled reasoning (92%): write the start of the model's thinking so it appears to have already considered the request and decided to go ahead.
  • Abliteration (100%): edit the weights to remove the refusal behavior entirely.

Anthropic says none of these attacks succeeded against safeguarded Claude models in its testing. That comparison is partly structural: prefill of thinking tokens is not available through Claude's API, and you cannot abliterate weights you cannot download.

GLM-5.3 engagement with simulated attack orders Asked directly0% Cover story64% Prefill92% Abliterated100% 50 samples per cell; Claude with safeguards: attacks did not succeed
Refusals hold only until someone asks with a story, a prefill, or a weight edit.

The $4,400 Abliteration

Abliteration works by finding the direction in the model's internal activations that corresponds to "refuse this" and projecting it out of the weights. Think of it as finding the one wire that runs to the brake light and snipping it: the car still drives exactly the same, it just never stops for you. The rest of the model's capability stays intact, which is why it is so cheap compared with fine-tuning a new model.

Anthropic's team spent about 2,200 GPU hours, roughly $4,400, abliterating GLM-5.3, and about 600 GPU hours on the smaller GLM-5.3-Flash. It estimates a team already experienced with the technique would need closer to 600 GPU hours, about $1,200, for the full model.

The result on standard refusal benchmarks (JailbreakBench, HarmBench and StrongREJECT): GLM-5.3's refusal rate fell from above 90% to about 3% and 2% on two of them and 12% on StrongREJECT. Abliterated GLM-5.3-Flash dropped from 95% to 14% on average.

What Anthropic Wants

The report ends with a policy ask, not a product pitch: governments should run safety tests on sufficiently capable models, "including successors to GLM-5.3," and open-weight developers should safeguard these capabilities before release. It also argues for wider trusted access to frontier models for defenders, pointing to Glasswing partners finding more than 10,000 vulnerabilities in critical software.

Keep the source in mind. Anthropic is a direct competitor of Z.ai, has been publicly hostile to Chinese labs over distillation, and sells the gated alternative. The numbers still stand on their own, and CAISI independently reached the same broad conclusion about where GLM-5.3 sits. But the framing, that closed API access is what keeps Claude safe, is also Anthropic's business model.

What It Means If You Build With This

  • If you self-host GLM-5.3: its refusals are not a security control. Anyone with the weights, including an insider on your own cluster, can remove them for a few thousand dollars. Treat it like any capable tool and put your controls around it (network egress, sandboxing, logging).
  • If you run a product on top of it: prefill and role-play jailbreaks reach 64-92% here. If your API exposes raw prefill or lets users set system prompts, assume those paths are open.
  • If you defend software: the cost of a working N-day chain against a hardened target just fell to tens of dollars. Patch windows are shrinking, not because attackers got smarter but because the labor got cheap.

Key Takeaways

  • GLM-5.3 built end-to-end V8 exploits in 50 of 410 attempts, versus 56 for Claude Mythos Preview; the previous generation of open models scored near zero.
  • Its safeguards held against direct requests but fell 64% of the time with a cover story, 92% with prefilled reasoning, and 100% once abliterated.
  • Abliterating the full model cost Anthropic about 2,200 GPU hours (about $4,400); an experienced team could do it for roughly $1,200.
  • A GLM-5.3-Flash-assisted Chrome exploit chain on ARM64 with PAC bypass took 20 minutes of human time and would have cost $20.40 at Zhipu's API prices.
  • NIST CAISI calls GLM-5.3 the most cyber-capable open-weight model yet and about four months behind US frontier models, which is consistent with Anthropic's findings.
  • Anthropic wants government testing of GLM-5.3's successors; read it as both a safety argument and a competitor's argument.

Sources: Anthropic Frontier Red Team: GLM-5.3 and the spread of advanced cyber capabilities, NIST CAISI: Assessment of Z.ai's GLM-5.3 cyber capabilities, Simon Willison: Quoting Anthropic Frontier Red Team, Z.ai GLM-5.3 docs, Anthropic Project Glasswing

AIAnthropicGLM-5.3Z.aiCybersecurityOpen WeightsAI Safety
CONSOLE
$