Paper: Coding Agents Beat Classical Planners on Generalized Task-and-Motion Problems, 56-95% vs 47%
Summary
A new paper evaluates Claude- and GPT-based coding agents on generalized task-and-motion planning (TAMP) by having them synthesize programs against 28 simulated environments from KinDER and PDDLStream. Across 980 generated programs and 98,000 evaluation episodes, three agent configurations scored 56% to 95% on 16 held-out environments versus 47% for classical planners, while using substantially less compute per instance. Agent traces show them calibrating physical models, testing edge cases and iterating — a strong new baseline for robot planning that had previously required hand-engineered solvers.
Originally reported by huggingface.co
Read the original article →Original headline: Paper: Coding Agents Beat Classical Planners on Generalized Task-and-Motion Problems, 56-95% vs 47%