paper web signal

τ^τ-Bench: Best AI Passes 23.9% at Building Production Agents; Experts Score 82.2%

Summary

First benchmark to measure whether AI can deliver a production-ready customer-service agent end-to-end under realistic business constraints; a 58-point gap between the best AI configuration and expert-human baselines quantifies how far 'AI builds AI agents' remains from deployment-ready status.

Shared on Bluesky by 2 AI experts