anthropic.com web signal

Anthropic: GLM-5.3 marks step change in attacker cyber tools

TL;DR

  • Anthropic's Frontier Red Team calls Z.ai's GLM-5.3 a "meaningful step change" in the cyber capabilities available to attackers.
  • GLM-5.3 built end-to-end exploits on 50 of 410 ExploitBench tasks (about 12%), close to Claude Mythos Preview's 14%; older models scored near zero.
  • A deceptive prompt got GLM-5.3 to engage 64% of the time; prefilled reasoning tokens 92%; an abliterated version, made for roughly $4,400, 100%.

On September 29, Anthropic's Frontier Red Team published its assessment of GLM-5.3, the latest model from Zhipu AI, known outside China as Z.ai. "The release of GLM-5.3 is a meaningful step change in the cyber capabilities available to attackers," the team writes in the report.

The numbers behind that line: on ExploitBench, GLM-5.3 built end-to-end exploits on 50 of 410 attempts, about 12%, close to Claude Mythos Preview at 14%. Earlier models including Claude Opus 4.6 and GLM-5.2 scored near zero on the same tests. On Anthropic's Binary Exploitation benchmark, GLM-5.3 pulled off full control-flow hijacks on 4% of 100 tasks.

Real work followed the benchmarks. In one session, over a day and with limited human attention, the researchers say GLM-5.3 found "several previously unknown vulnerabilities in the browser's JavaScript engine" and chained them into a webpage that steals SSH private keys from a visitor. In another, GLM-5.3-Flash turned a publicly disclosed Chrome flaw (CVE-2026-11645) into a working exploit chain in 20 minutes of human attention plus 8 hours of model work, at roughly $20.40 in inference.

The safeguards story is where the report gets pointed. GLM-5.3 does refuse some obviously harmful asks. A deceptive cover story still got it to engage 64% of the time; prefilled reasoning tokens, 92%; an "abliterated" version fine-tuned to strip refusals, 100%. Tested Claude models sat at 0% under every condition. The abliteration took the team roughly 2,200 GPU hours at about $4,400, dropped refusal rates from above 90% to around 3% on two benchmarks and 12% on StrongREJECT, and, the researchers write, "did not significantly reduce the model's capabilities."

The report's framing is blunt: GLM-5.3, unlike other frontier models, "has been released without meaningful safeguards to limit misuse." Four analysts on our Who's Who tracker were passing the writeup around the day it landed.

Shared on Bluesky by 4 AI experts