Multi-turn LLMs act on changes users rejected, paper finds
TL;DR
- A new preprint from a team led by Junle Chen argues that multi-turn language models often keep executing changes a user raised and then rejected.
- The authors label the pattern 'mentioned-as-in-effect confusion' and probe it with Intent-Eval, a benchmark spanning tool actions, code, databases, and mathematics.
- Their proposed fix, Intent-OPSD, is a Teacher-Student self-distillation setup initialized from the same base model and trained to track active user intent.
Multi-turn language models routinely act on changes the user has already rejected, treating a proposal that was mentioned and then withdrawn as if it were still in force. "Merely mentioning a rejected change can derail task execution, even when the user's final intent remains unchanged," write the authors of a new arXiv preprint led by Junle Chen.
They label the pattern "mentioned-as-in-effect confusion," and probe it with a benchmark they call Intent-Eval, "a controlled benchmark spanning tool actions, code, databases, and mathematics." The paper's claim is that the drift is not self-correcting: "Accuracy degradation can deepen or persist as interaction continues," it reports.
The proposed fix, Intent-OPSD, is "a decision-conditioned on-policy self-distillation framework with Teacher and Student initialized from the same model," in which the frozen Teacher supervises the Student toward the active-intent trajectory consistent with the user's decision.
The abstract names no specific models and publishes no per-task accuracy numbers, so the magnitude of the regression and which systems are worst exposed will have to wait for the full paper.
Originally reported by paper
Read the original article →Original headline: LLM Agents Silently Execute Changes Users Rejected Mid-Conversation