DF26 benchmark: humans and detectors near chance on AI video
TL;DR
- On the DF26 benchmark, humans identified synthetic public-speaking video at 52.6 percent, versus 74.5 and 69.8 percent on two earlier deepfake datasets.
- The best detector tested, GenD-PE, reached only 69.7 AUROC, while two temporal detectors dropped to 48.2 and 61.6 AUROC on DF26.
- DF26 covers 271 real and 2,420 synthetic clips from seven generators including Veo 3.1, Kling 3.0, Wan 2.6, and HunyuanVideo 1.5.
On a new benchmark of 271 real and 2,420 synthetic public-speaking clips, human viewers identified AI-generated video from real at 52.6 percent, a coin flip. The number comes from DF26, an arxiv paper posted this month, and it sits well below the 74.5 percent and 69.8 percent humans scored on two earlier deepfake datasets.
The setup is narrow on purpose. DF26 covers "single-person public-speaking scenarios, spanning direct-to-camera recordings, official statements, and studio interviews." Fakes were produced by seven text-to-video and image-to-video systems: Kling 3.0, Veo 3.1, Wan 2.6, Grok Imagine 1.0, plus the open-source Wan 2.2, HunyuanVideo 1.5, and LTX 2.3 distilled.
The automated stack does not save it. "The highest AUROC of 69.7 is achieved by GenD-PE," the paper reports of the best detector on the full benchmark. Two temporal detectors that scored well on legacy data collapsed on DF26: "their performance drops substantially on DF26, to 48.2 and 61.6 AUROC, respectively." One of them "achieved 92.9 AUROC on Wan 2.2 T2V, but on HunyuanVideo 1.5 it performs no better than chance."
The authors are blunt: "human performance in detecting AI-generated videos, as well as state-of-the-art deepfake detectors, is close to random chance."
The paper does not break out per-generator human accuracy for the four commercial systems, and does not test detectors fine-tuned on DF26 generators.
Shared on Bluesky by 1 AI expert
Originally reported by paper
Read the original article →Original headline: Human and SOTA-Detector Deepfake Detection Collapses to Near Coin-Flip on Modern AI Video