paper web signal

Trace2Env rebuilds agent environments from interaction traces

TL;DR

  • Trace2Env converts historical interaction traces into a reusable 'environment worldbook' without any model training.
  • The framework was tested across nine environments and beat prompt-based language world models on next-observation fidelity.
  • Actions generated against the simulator stayed valid more often when replayed in the real underlying system.

The claim is narrow and specific. Across nine environments, Trace2Env, a learning-free framework from a team led by Quanyu Long, reconstructs historical interaction traces into what the authors call an "environment worldbook," then lets a world-model agent read that artifact and stand in for the real system during agent training and evaluation.

The target problem is familiar to anyone trying to train agents against systems they cannot clone. "Realistic environment replicas are increasingly valuable for training and evaluating LLM agents, yet the original systems may be inaccessible or impractical to reproduce," the authors write. Rather than rebuild an executable version, they let the world-model agent consult a schema plus "grounded evidence, and induced behavioral knowledge," together with persistent episodic state, to infer each action's observation and lasting effects.

The headline results are two. Trace2Env "improves both next-observation fidelity and long-horizon interaction consistency over conventional prompt-based LWMs," the paper reports, and task-agent actions generated against the simulator remain valid more often when replayed in the real environment.

The abstract publishes no per-environment numbers, no baseline deltas, and does not name the nine systems. One preprint, posted this month, no outside replication yet.