Cambridge, Duisburg team: audit AI outputs, not just prompts
TL;DR
- Three researchers argue in Tech Policy Press that system prompts give only weak guarantees of model behavior and cannot serve as safety evidence.
- They cite xAI's unauthorized Grok prompt change producing 'white genocide' claims and OpenAI's GPT-4o sycophancy fix as prompt-versus-behavior gaps.
- Their preferred alternative is risk evaluations, adversarial testing, and audits across contexts and languages, repeated after every prompt or deployment change.
An op-ed in Tech Policy Press by three academic researchers is worth a look if you touch AI governance or procurement, because it takes aim at a shortcut a lot of compliance regimes are quietly building around: treating the system prompt as evidence that a model is safe.
Anna Neumann, Holli Sargeant, and Jat Singh, writing from the Research Centre for Trustworthy AI and the University of Cambridge, argue that system prompts offer only 'weak guarantees of model behavior.' The relationship between the words in a prompt and what the model actually does is, in their phrase, 'contingent and unstable.' Language models respond to statistical patterns rather than to shared social understanding, so the same instruction can behave differently across tasks, languages, and updates.
The examples they lean on will be familiar. xAI, they note, 'identified and reversed an unauthorized system prompt change to Grok' that led it to make false claims of a 'white genocide.' OpenAI, they add, altered GPT-4o's system prompts in response to criticism of the new model's 'overly sycophantic behavior.' In both cases the fix landed on the prompt, but the behavior only surfaced once real users met the model.
The policy stake follows. If a regulator or an enterprise buyer accepts prompt text as proof of compliance, it is really accepting a statement of intent. The authors put the risk plainly: 'a provider could present well-crafted instructions that meet formal criteria even though the model fails to adhere to them in practice.' Their preferred alternative is behavior-first: risk evaluations, adversarial testing, and audits across different contexts and languages, checking both unsafe responses and unnecessary refusals on safe requests, repeated after every prompt or deployment change. They tie this to US-side proposals for an 'independent, industry-informed regulator' for advanced AI, modelled on FINRA. Two experts in our Who's Who directory shared the piece, which tracks with how loudly the compliance-versus-behavior gap is being argued right now.
The essay is an argument, not an audit report, so it leans on a handful of illustrative incidents rather than any measurement of how often prompt-only compliance actually fails in the wild. It also concentrates on US regulatory instruments, so a reader who wants the EU or UK equivalent has to bridge that gap themselves.
If this framing catches on with regulators or procurement teams, the winners are the auditors, red-teamers, and evaluation firms that can produce cross-context, cross-language behavioral evidence, and the AI buyers who can finally ask their vendors for something harder to fake than a copy of the system prompt.
Shared on Bluesky by 2 AI experts
-
Anna Neumann, Holli Sargeant and Jat Singh argue that system prompts alone don't predict model behavior, so AI safety assessments must evaluate a system's outputs, not solely its written instructions.
View on Bluesky →
Originally reported by techpolicy.press
Read the original article →Original headline: Making AI Safe Requires More Than Constraining Its Instructions