CADWorld benchmark: top agent 17.5% on FreeCAD vs experts' 87%
TL;DR
- On the new CADWorld benchmark, the strongest computer-use agent tested reached 17.5% success versus 87.0% for expert reference performance.
- The benchmark contains 200 tasks across 11 mechanical-CAD workflow categories in FreeCAD, including sketching, assembly, CAM, FEM, and technical drawing.
- Success is graded by executable checks on saved artifacts covering geometry, parametric structure, constraints, manufacturing state, and simulation results.
The strongest computer-use agent tested on CADWorld solved 17.5% of tasks. Expert reference performance on the same set was 87.0%.
That is the headline result of CADWorld, a benchmark from Zihan Dong and coauthors that puts agents inside FreeCAD and grades them on 200 tasks across 11 mechanical-CAD workflow categories, including sketching, assembly, CAM, FEM, and technical drawing. Agents drive the software through screenshots and GUI actions; a checker then inspects the saved files.
The authors are blunt about why they built it: 'Computer-use agents are increasingly evaluated in realistic desktop environments, but existing benchmarks provide limited coverage of professional engineering workflows whose outputs are persistent, structured artifacts.' Grading is not pass/fail on a screen, the paper says, but 'task-specific executable checks over saved FreeCAD artifacts and auxiliary outputs, covering geometric properties, parametric structure, constraints, manufacturing state, and simulation results.'
The failure pattern splits by ability. Weaker agents often fail to produce a valid artifact at all. Stronger ones produce something, then trip on structural requirements, geometric precision, and construction methodology. That, the paper reports, 'exposes a gap between general GUI competence and reliable execution of persistent, verifiable engineering workflows.'
The abstract does not name which model reached the 17.5% ceiling, and does not publish per-category scores.
Shared on Bluesky by 1 AI expert
Originally reported by paper
Read the original article →Original headline: CADWorld: Best AI Agent Scores 17.5% on Professional FreeCAD Tasks—Human Experts at 87%