arxiv.org web signal

Terminal-Agent Survey Argues for Runtime-Aware Evaluation

TL;DR

  • A new arXiv survey groups AI "terminal agents" around a common lens: systems whose main loop is mediated by command execution and textual feedback.
  • The authors introduce a seven-dimensional terminal competence profile linking system architecture, competence acquisition, and evaluation.
  • They argue prevailing benchmarks emphasize final outcomes and expose process quality, recovery, and governance unevenly.

Terminal-mediated AI agents are jointly shaped by "the model, interface, harness, runtime, and environment," according to a 52-page survey posted to arXiv on August 20 by Yi Bin and co-authors. The paper reframes a fragmented literature, spread across software engineering, tool use, and computer-use research, around a single organizing lens: agents whose progress-bearing loop runs through "terminal command execution, textual feedback, and stateful environment interaction."

Existing evaluations, the authors argue, "emphasize final outcomes and expose process quality, recovery, and governance unevenly." To address that, the survey introduces a seven-dimensional terminal competence profile linking system architecture, competence acquisition, and evaluation, and calls for "explicit reporting of system and runtime conditions, supported by replayable traces and process-level evidence."

The paper also cautions that "matched system comparisons reveal benchmark-dependent performance and limits of component attribution." Two of the researchers on our Who's Who list posted the link within days of the arXiv submission.

Shared on Bluesky by 2 AI experts