huggingface.co web signal

Meta's MIMESIS User Simulators Lift Agent Scores Nearly 5 Points

Meta Agents Synthetic Data ai-business

TL;DR

  • MIMESIS-9B posts a SOUL-Index of 65.7, edging past Claude-Opus-5 (64.9) and GPT-5.5 (64.1) on user-simulation fidelity.
  • A fixed GPT-5.5 agent on τ-bench succeeds 63.6% with real users but 82.4-84.4% with frontier API simulators playing the user.
  • Agents trained against MIMESIS plus Coached Self-Distillation score 31.09 vs 26.10 for a GPT-5.5 practice user across nine unseen evaluators.

Meta Superintelligence Labs published MIMESIS, a purpose-built user simulator that it says exposes a hidden flaw in how most agent benchmarks are run: the 'user' in the loop is usually another assistant LLM, and assistant LLMs are too agreeable. On τ-bench, a fixed GPT-5.5 agent succeeds 63.6% of the time with real human users but climbs to 82.4-84.4% when the user is a frontier API model, according to the paper.

The authors argue those off-the-shelf simulators are "overly cooperative, explicit, and behaviorally homogeneous compared with real users." MIMESIS, built on Qwen3.5 backbones at 4B and 9B, is trained to imitate 13 patterns drawn from real conversations — hidden evaluation criteria, incremental goalpost shifting, delayed information disclosure, false premises — with explicit reasoning supervision the team calls ThoughtTrace.

By the paper's own numbers, MIMESIS-9B posts a SOUL-Index of 65.7, nudging past Claude-Opus-5 at 64.9 and GPT-5.5 at 64.1. On RealUserSim PT3 fidelity it reports 94.0 against Claude-Opus-5's 80.6, a gap of 13.4 points, and on SimulatorArena a 3.6-point reduction in Turing distance.

The downstream claim matters more than the fidelity one. Agents trained against MIMESIS score higher when later evaluated by nine unseen user simulators. Under a GRPO baseline, swapping GPT-5.5 for MIMESIS-9B as the practice user moves the mean task score from 26.10 to 29.54; adding the paper's Coached Self-Distillation step takes it to 31.09. "Across eight environments, training with MIMESIS yields better agent performance than training with GPT-5.5 under all nine unseen user simulators," the authors write.

Code is on GitHub. The release lands in a dense month of AI agent coverage and alongside a thinner stream of synthetic-data releases that are quietly reshaping how agent evaluation gets run.