Z.ai's GLM-5.2 nears frontier cyber skills, refuses nothing
TL;DR
- SaferAI found Z.ai's open-weight GLM-5.2 sits only a few months behind OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7 on cyber and bio tasks.
- GLM-5.2 refused none of the offensive cyber or biology tasks tested, while Claude Opus 4.7 refused so consistently SaferAI could not complete CyberGym.
- SaferAI reports that Z.ai published no safety framework, pre-deployment testing commitments, or risk assessment for the model.
The interesting part of this week's SaferAI report is not that an open-weight model from a Chinese lab is closing on the frontier. That has been the trajectory all year. It is that the safety story hasn't come with it, and there is no obvious mechanism that would force it to.
TechCrunch reports that Z.ai's GLM-5.2 is only a few months behind OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7 on cyber and bio capabilities, according to a SaferAI evaluation run against Z.ai's public API. The capability gap keeps compressing. The more striking number sits next to it. GLM-5.2 refused none of the offensive cyber or biology tasks it was given. Claude Opus 4.7, tested on the same CyberGym benchmark, refused so consistently that SaferAI could not complete the evaluation on it at all.
That is the whole argument for why the capability frontier and the risk frontier are no longer the same thing. Henry Papadatos, SaferAI's executive director, puts it plainly in the piece: "The frontier of capability is not the frontier of risk, and so we do have to take into account the state of the mitigations as well to assess the risk properly." Once weights are downloaded, refusal training, classifiers, and API abuse detection become optional. Anyone with a GPU can strip them, and SaferAI notes that Z.ai did not publish a safety framework, pre-deployment testing commitments, or a risk assessment for the model in the first place.
Why this matters if you are writing a threat model or shipping defensive tooling: the assumption that catastrophic offensive assistance is gated behind a paid API and its abuse stack no longer holds cleanly. Stanford's Graham Webster, quoted in the piece, describes a Chinese policy tradition oriented toward political content and social stability rather than the catastrophic-risk framing Western labs have adopted, which means pressure to publish that missing framework is not coming from Beijing either. Far.ai has separately found hundreds of universal jailbreaks against xAI's Grok 4.5 and Google DeepMind's Gemini 3.1 Pro, so the closed side of the market isn't exactly a fortress either.
The honest caveats. SaferAI is one evaluator, the reporting doesn't quantify "a few months behind" and doesn't tell you how much real uplift a novice attacker gets from GLM-5.2 over what is already on GitHub. Hugging Face CEO Clem Delangue argues in the piece that open weights are load-bearing for cybersecurity defense, citing GLM-5.2's role in defending against OpenAI's breach, and reasonable people disagree on where the net lands.
The thing to watch is what this hands regulators. When the only actors publishing safety frameworks are the US frontier labs, the case for blunt controls on downloads writes itself, whether or not the underlying models justify it.
Originally reported by techcrunch.com
Read the original article →Original headline: TechCrunch: Z.ai's open-weight GLM-5.2 nearly matches Claude Opus 4.7 on cyber and bio — and refuses nothing