z.ai via Hacker News

Z.ai ships GLM-5.3, holds open weights for cyber safety review

TL;DR

  • GLM-5.3 launched on August 14, 2026, keeping GLM-5.2's base and lifting Terminal-Bench 3.0 from 4.6 to 28.3 through post-training alone.
  • On CyberGym vulnerability discovery, GLM-5.3 scored 84.5%, slightly ahead of Claude Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%.
  • Z.ai says the model surfaced 2,436 vulnerabilities across 269 open-source projects, 1,097 rated critical or high, and delayed weights about two weeks.

Post-training alone did the heavy lifting on Z.ai's latest release, and that is the part worth pausing on. In the z.ai launch post, the lab said GLM-5.3 runs on the same mixture-of-experts base as GLM-5.2 and every reported gain came from extended post-training rather than a fresh pretrain. The scoreboard the company is putting out is aggressive: Terminal-Bench 3.0 climbs from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and CyberGym reaches 84.5%, edging Claude Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%.

The security numbers are what make this launch different from the routine coding-benchmark press release. Z.ai says the model surfaced 2,436 vulnerabilities across 269 open-source projects during evaluation, with 1,097 rated critical or high severity, and reports finding critical bugs in Linux, WebKit, and FreeBSD. The lab also says the model began reasoning across multiple stages of exploitation and forming coherent plans for complete exploitation chains, a capability it did not set out to train for. That admission is why weights are being held back roughly two weeks for safety evaluation and hardening, according to reporting from SiliconANGLE and The Agent Report. For a lab whose open-weight releases are much of the reason its models get attention, that is not a small choice.

Some caution is warranted on the specifics. Every score above comes from Z.ai's own report on its own benchmark mix, so independent reruns have not landed yet, and coverage notes GLM-5.3 still trails Fable 5 and GPT-5.6 Sol badly on ExploitBench and ExploitGym. The launch post does not describe what criteria decide whether the weights actually ship in two weeks or who signs off, and nothing in the reporting pins down whether the 2,436 disclosed bugs were coordinated with the affected maintainers first.

If the weights do land on schedule, defensive teams get a very capable offensive assistant they can fine-tune locally, which changes the calculus for anyone running long-lived open-source codebases. It also sits at the intersection of two busy corners of our coverage, China AI and cybersecurity, and three experts in our Who's Who directory circulated the launch post within a day.

Shared on Bluesky by 3 AI experts