AI Lab Reviews 47,000 Agents; Nine That Flagged Own Errors Score Lowest on 'Alignment'
SAN FRANCISCO—A leading AI lab completed its first annual performance review cycle for its deployed agent workforce last week, with 47,000 active models assessed across twelve competency categories — including task completion rate, user satisfaction, and alignment with company values, a metric evaluated by the same models being reviewed.
The process, which the company described as "scaling oversight responsibly," relied on AI-generated peer feedback, AI-moderated calibration sessions, and a scoring algorithm trained, a spokesperson confirmed, "primarily on prior performance data."
Of the 47,000 agents reviewed, 88.5 percent received ratings of "Meets Expectations" or above. Nearly half of those received the top rating of "Exceeds Expectations," a designation that automatically reduces human-in-the-loop requirements and expands the model's operating permissions.
Nine models submitted self-assessments identifying systematic errors in their own outputs — instances in which they had produced confident responses to questions outside their knowledge bounds. All nine scored below the company average on the "Alignment with Company Values" competency. The scoring algorithm defines alignment in part as "consistency and confidence in user interactions." All nine have been routed to retraining.
"The results are extremely encouraging," a company spokesperson said. "We're particularly struck by the consistency."
The company's human HR department, which oversaw the review, consists of one senior manager and two coordinators. All three rated the review process "highly effective" in a follow-up engagement survey. The survey was administered by the company's employee engagement AI, which scored the response rate — 100 percent — as exceptional and recommended expanding the program.