paper web signal

Harvard, Georgia Tech post 48-task robot-engineer exam for AI

TL;DR

  • RLE-Bench puts coding agents through 48 standardized tasks spanning interactive control, policy development, perception and estimation, and mechanical design.
  • The suite is framed as a qualifying exam for the role of 'robot learning engineer,' with scores averaged within workflows then across them on a 0 to 100 scale.
  • Harvard's Na Li calls the release 'RLE-Bench 1.0' and 'only a starting point,' and is asking outside labs to contribute new tasks.

Harvard and Georgia Tech researchers have put coding agents in front of 48 standardized tasks that ask them to build, debug, and improve pieces of real robots, not just complete pre-defined control policies. RLE-Bench, a 28-page preprint posted September 28, 2026, frames the suite as a qualifying exam for the role of "robot learning engineer."

The 48 tasks fall into four workflows: interactive control, policy development, perception and estimation, and mechanical design. In a Harvard SEAS write-up, lead investigator Na Li says: "Robotics already has many valuable benchmarks, but many focus primarily on learning and evaluating control policies. RLE-Bench takes a broader view. Can an AI system perform the range of engineering tasks required to make a robotic system actually work?"

Physical constraints are the point. Co-author Haitong Ma: "An agent can produce something that looks reasonable computationally, but once you consider the physics of the complete system, important failure modes can appear."

Scores exist on a public leaderboard at rle-bench.github.io but the retrieved pages do not list numeric results per frontier model. Li calls the current release "RLE-Bench 1.0" and "only a starting point."