techcrunch.com web signal

OpenAI, Anthropic Guardrails Frustrate Offensive Security Pros

TL;DR

  • TechCrunch reports offensive security researchers at firms like NCC Group and Crowdfense say AI guardrails are inconsistent and slow their vulnerability work.
  • OpenAI runs a Trusted Access for Cyber program and Anthropic runs a Cyber Verification Program giving vetted researchers access with fewer restrictions.
  • Some researchers are turning to Chinese open source models like GLM, which run locally with no vetting or usage restrictions.

Offensive security researchers spend their days looking for bugs before criminals do, and lately they say the biggest thing slowing them down is not obfuscation or exploit mitigations but the AI models themselves. TechCrunch reports that researchers at firms including NCC Group and Crowdfense are hitting inconsistent guardrails when they ask commercial models to help analyze code or confirm a bug.

Chris Anley, chief scientist at NCC Group, framed the dual-use problem plainly, saying 'fix this code' as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base. He compared the tool to a hammer that you can't build a house without but that is 'also irreducibly a weapon as well.' Chris Thompson, chief executive of RemoteThreat and founder of Offensive AI Con, told the outlet the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program.

Both OpenAI and Anthropic have vetted-access lanes for exactly this use case. OpenAI runs a Trusted Access for Cyber program, and Anthropic runs a Cyber Verification Program, with both providing vetted cybersecurity researchers access to models with fewer security restrictions. The article notes Anthropic's Fable 5 returned on July 1 after export-control restrictions in June, with Mythos 5 reintroduced to vetted U.S. organizations. Mark Dowd, a zero-day finder, said it is 'not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what's not.'

The honest caveat is that this is a single piece leaning on a small handful of named researchers, and none of them argue guardrails have no purpose. Giuseppe Cali told the reporter he'd still 'want to own the actual bug discovery and weaponization myself' with or without them, and even said, 'I am jealous of my bugs, and I like this game too much to let models play it for me.' What the reporting doesn't spell out is how many researchers have actually been admitted to the vetted programs, or the criteria the labs use to gate them.

The forward-looking piece worth watching is where researchers go when they give up on vetted access. The article says some are already being pushed toward Chinese open source models like GLM, which run locally with no vetting or usage restrictions. If U.S. labs want to be the platform of choice for offensive research, the vetted-programs pitch will have to be more predictable than the guardrails they replace.

Shared on Bluesky by 2 AI experts