huggingface.co web signal

OmniConfess Scores Tokens to Catch Omni-Modal Hallucinations

TL;DR

  • OmniConfess is a training-free inference-time method that freezes a candidate response and re-scores it token-by-token under channel-wise evidence interventions.
  • The paired OmniHalluBench holds 3,540 examples drawn from six public datasets covering text, image, audio and video, in judgment and free-form generation.
  • On Qwen2.5-Omni-7B, Qwen3-Omni-30B-A3B and Nemotron-3-Nano-Omni-30B, OmniConfess lifts average F1 over the base model by 10.30, 6.34 and 7.50 points.

Researchers from Beijing University of Posts and Telecommunications and Nanyang Technological University, Singapore have put out OmniConfess, a training-free inference-time method that tries to pin down which modality is carrying each generated token in omni-modal LLMs, and to overwrite the ones leaning on the wrong channel.

The mechanics are three steps the paper calls commitment anchoring, evidence interrogation and confession-guided correction: generate a candidate response, freeze it, then re-score the same tokens with individual evidence channels removed so each token gets a per-channel dependence score. The authors motivate the token-level framing with a measurement from their own analysis: "the top 20% of tokens capture 82% of the dependence mass on average," with the first third of a response alone holding 49.8% of it.

To evaluate, they ship OmniHalluBench, built from six public datasets including PubMedQA, HaloQuest, AVHBench and CMM, and spanning text, image, audio and video with both judgment and free-form tasks. The abstract reports it is a "3,540-example benchmark" and that "OmniConfess outperforms the strongest relevant baseline across all six datasets on the primary backbone," Qwen2.5-Omni-7B. Per-dataset gains cited in the paper include +10.20 F1 on PHD, +9.35 F1 on CMM and +16.84 Token-F1 on RAGTruth.

An ablation puts the full method at 71.63 average F1, with removing anchoring, interrogation or correction dropping that to 67.06, 63.96 and 58.66 respectively; the authors write that "correction's removal has the largest effect." Across Qwen2.5-Omni-7B, Qwen3-Omni-30B-A3B and Nemotron-3-Nano-Omni-30B, OmniConfess improves average F1 over the base model by 10.30, 6.34 and 7.50 points.

Code and the benchmark are on GitHub. It arrives into a thick stream of related work: this is our 36th hallucinations story in the last 90 days, alongside 68 on the multimodal side.