paper web signal

AFTER: 382-Task Enterprise Benchmark Finds Procedural Agent Skills Transfer at 73% Accuracy Across LLM Architectures

Summary

Procedural memory is increasingly deployed to boost LLM agents on recurring tasks, but whether those skills survive a model swap is an open question. AFTER is the first enterprise-scale benchmark to measure cross-task, cross-role, and cross-model skill transfer, finding that skills drawn from diverse multi-model execution traces generalize to 73.1% accuracy on unseen architectures.

Shared on Bluesky by 1 AI expert