arxiv.org web signal

ProMediConv benchmarks LLM agents on 972 legal mediation cases

TL;DR

  • ProMediConv models mediation as a proactive, multi-stage, party-aware dialogue with 11 strategies and four party-behavior-pattern states.
  • The dataset draws on 972 complete real-world legal cases annotated at the utterance level for strategies and behavior-pattern states.
  • The authors introduce Mean Attribute Difference (MAD) to measure how an agent's turns shift party behavior across a mediation dialogue.

ProMediConv, a new benchmark posted to arXiv, treats legal-dispute mediation as a proactive, multi-stage dialogue task and uses 972 real-world cases to test how well conversational agents can steer disputing parties. The paper was submitted on 10 Sep 2026 by Zesheng Wei, Mengfan Li, Wenhao Liu, Yixin Zhang, Zilei Wang and Yang Deng.

The framework encodes 11 mediation strategies and four "party behavior pattern (BP) states," and layers utterance-level annotations over the case transcripts. To measure whether an agent's turn actually moves a disputant, the authors propose Mean Attribute Difference (MAD), described as "a fine-grained metric that captures BP shifts throughout the dialogue." They also introduce a tailored baseline they call ProMediAgent and evaluate a range of models against it.

The finding, in the paper's own careful phrasing, is not encouraging. Extensive empirical analyses "reveal critical behavioral phenomena and underscore the persistent challenges current models face in dynamic, multi-party mediation." The abstract does not publish per-model scores or name the "diverse models" tested; those numbers are left to the body of the paper. Two of the researchers we follow flagged it the day it went up.

Shared on Bluesky by 2 AI experts