AI Video Fools Humans at 52.6%, Near Chance; Best Detector Tops Out at 69.7%

Found first: a primary source the press has not covered yet.

A new benchmark called DF26, published on arXiv, finds that people correctly identify AI-generated public-speaking video as fake 52.6% of the time, compared to 74.5% on older synthetic video. The best algorithmic detector reaches 69.7% AUROC on the same dataset, and a second detector falls to 48.2%, effectively at chance.

What the source says

The benchmark was produced by Severyn Shykula (Ukrainian Catholic University), Andrii Yermakov, Ivan Samarskyi, and Jan Cech (Czech Technical University in Prague), and Dmytro Mishkin and Anastasiia Mishchuk (Hover Inc. and the National Academy of Sciences of Ukraine). DF26 contains 271 real videos and 2,420 synthetic videos from seven generators (Wan 2.6, Veo 3.1, Grok Imagine 1.0, Kling 3.0, Wan 2.2 A14B, HunyuanVideo 1.5, LTX 2.3), all showing single-person public-speaking in direct-to-camera, official statement, or studio interview settings. A human study of 232 labeling sessions found 52.6% accuracy at identifying DF26 fakes as synthetic, against 74.5% on CelebDF++ and 69.8% on DeepSpeak v2. The strongest detector, GenD-PE, scored 69.7% AUROC on DF26, down from 89.9% on CelebDF++, while DFD-FCG scored 48.2%. Human accuracy on LTX 2.3 output specifically fell to 25.2%.

Why it matters

DF26 tests the scenario types most common in political messaging and institutional communication: direct address to camera, official statements, and studio interviews. The 20.2-point AUROC drop for the best-performing detector, from CelebDF++ to DF26, gives a concrete measure of how quickly detection capability has degraded against models released between late 2025 and early 2026. The authors frame this as a distribution shift problem: detectors trained on earlier, lower-quality synthetic video do not transfer to current generators. With US midterm elections in November 2026, the specific, quantified failure on this video category is directly relevant to platform trust and safety teams and election security practitioners.