τ^τ-Bench: Best AI Passes 23.9% at Building Production Agents; Experts Score 82.2%
Summary
First benchmark to measure whether AI can deliver a production-ready customer-service agent end-to-end under realistic business constraints; a 58-point gap between the best AI configuration and expert-human baselines quantifies how far 'AI builds AI agents' remains from deployment-ready status.
Shared on Bluesky by 2 AI experts
Originally reported by paper
Read the original article →Original headline: τ^τ-Bench: Best AI Passes 23.9% at Building Production Agents; Experts Score 82.2%