paper web signal

Paper Traces Audio-Video Prompt Overrides to 'Attention Triangle'

TL;DR

  • The audio-video edge inside these diffusion models is bidirectional: sound can shape video and video can shape sound during generation.
  • When prompts fight learned priors, cross-modal attention can override the conditioning and steer output toward 'visually canonical but incorrect outcomes.'
  • The authors use attention-derived signals as both a diagnostic and an inference-time intervention that improves grounding without retraining.

Audio-video diffusion models can override the prompt on their own. In a new preprint on arXiv, Sagi Polaczek and co-authors argue that when a text prompt is in tension with what the model has learned as canonical, cross-modal attention reroutes the generation toward "visually canonical but incorrect outcomes."

The paper frames the mechanism as an "attention triangle": the three cross-attention edges connecting the text, audio, and video streams inside these models. The authors report that "routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation." The sound channel and the visual channel do not just each obey the prompt; they also negotiate with each other, and that negotiation is shaped by priors baked into model weights.

The leakage, they say, is not random. "These effects suggest that semantic artifacts arise not merely from attention spreading beyond its intended target, but from structured, bias-driven interactions along specific pathways," the authors write.

They then use the same attention signals as a fix. By reading how semantics are distributed across the three edges, they say they can "guide inference-time interventions that encourage more consistent cross-modal alignment," reporting improved grounding while preserving generation quality without retraining the base model.

The abstract names no specific models, benchmarks, or datasets. Magnitude of the reported improvement and compute overhead of the intervention are not quantified there.

Shared on Bluesky by 1 AI expert