huggingface.co web signal

MBZUAI: Deepfake Detectors Slide 99.5% to 76%, Break Under Attack

TL;DR

  • Across 20 deepfake detectors and 10 image generators since 2022, MBZUAI measured accuracy sliding from 99.5% to 76%, with every detector falling below 2% under adversarial perturbation.
  • In 3,000 Reddit images, Stable Diffusion 2.1 failed to reproduce 1,116 in 2022; by 2024 only 55 to 79 images resisted reproduction, a 14x erosion.
  • The authors' reconstruction-based certifier holds a 1% false-positive bound against adaptive attackers in the bounded-perturbation space, but not against arbitrary transformations.

Twenty deepfake detectors tested against ten image generators released since 2022 showed 'accuracy decreasing over time, from near-perfect 99.5% to 76%,' the MBZUAI team reports. Under a bounded perturbation attack at epsilon=8/255, every one of the 20 detectors fell below 2% accuracy. OmniAID, the strongest in its default setting at 93.25%, dropped to 0.00% post-attack.

The paper, posted to Hugging Face on October 6, argues the problem runs deeper than engineering. 'More fundamentally, authenticity cannot be decided from content alone, since a more capable generator may reproduce authentic content exactly, e.g., through memorization,' the authors write. Their alternative is reconstruction-based: certify an image as authentic only when no known generator can faithfully invert and rebuild it. 'If the reconstruction is faithful, then anyone with that generator could have created the sample, its authenticity is plausibly deniable, and our method abstains.'

To stress-test the method, Sarim Hashmi and co-authors ran it against 3,000 public Reddit images. In 2022, Stable Diffusion 2.1 failed to reproduce 1,116 of them, so 37.2% could still be certified real. By 2024, SD3.5 Medium failed on only 64 images, FLUX.1 Dev on 79, and FLUX.1 Dev with a Realism LoRA on just 55. The authors call this a 14x erosion in two years, and warn that 'Every generator release therefore shrinks the set of content that anyone can still certify.'

Calibrated at a 1% false-positive rate, the reconstruction detector holds that bound even against an adaptive attacker searching within the bounded-perturbation space; at the same operating point, the paper reports, most baselines reach near-zero recall. The calibration 'does not cover arbitrary adversarial transformations.' The result lands in the middle of a run of provenance stories we've tracked this week, alongside OpenAI's EU ChatGPT watermarking plan and UCLA's optical-chip deepfake detector.