huggingface.co web signal

MBZUAI Attack Steers Diffusion LLMs Toward Targeted Bias

TL;DR

  • On ambiguous BBQ items, MBZUAI's PI-controller attack lifts LLaDA-8B-Instruct's preference for a targeted demographic group from 1.8 to 16.7 percentage points.
  • On SocialStigmaQA, stigmatizing-answer selection rises from 17.6% to 58.1%, with shifts of up to 37 percentage points on other demographic targets.
  • Each attack takes about 40 minutes on one GPU, and the authors call for bias audits of the serving stack, not just the frozen model.

On ambiguous BBQ questions where the correct answer is abstention, an inference-time attack on LLaDA-8B-Instruct raised the model's preference for a targeted demographic group from 1.8 to 16.7 percentage points.

The attack, described in Noise Out, Bias In, a paper by Sarim Hashmi, Mukul Ranjan and colleagues at MBZUAI, works on masked diffusion language models, not autoregressive ones. The abstract explains why: 'An autoregressive decoder exposes an answer's distribution once, at the step that commits it; a dLLM exposes it at every denoising step before commitment, and we show that an adversary can exploit this.'

The method is deliberately plain. An attacker with access to internal activations tracks the target-answer probability through denoising, and a proportional-integral controller adjusts the strength of a steering vector on the fly. On SocialStigmaQA the same attack moved 'the selection of stigmatizing answers from 17.6% to 58.1%.' Shifts of 'up to 37 percentage points' are reported on other demographic targets. 'Each attack takes about 40 minutes on one GPU.'

Static steering can't match it. The paper reports that constant steering 'produces a far smaller shift while corrupting nearly three times as many outputs,' and that choosing a different constant for each example still falls well short.

The authors land on an auditing claim: the denoising trajectory is 'a new control channel in dLLMs' and bias checks should 'examine the serving stack rather than the frozen model alone.' Weight-only audits do not see this. Code is on GitHub.

The abstract reports no figures for Dream-7B-Instruct or LLaDA-MoE, and whether the trick transfers to served endpoints that don't expose residual activations is not addressed. It is one more entry in a busy quarter of AI safety research we are tracking.