Chain-of-Experience: 8 Frontier LLMs Gain 5.6% at Test Time With 19% Lower API Cost
Summary
Chain-of-Experience shows that frontier models including GPT-5 and Gemini-2.5 Pro can compound inference-time improvement without new training runs, and that stronger base models benefit more—complicating the assumption that capability gains require retraining.
Originally reported by paper
Read the original article →Original headline: Chain-of-Experience: 8 Frontier LLMs Gain 5.6% at Test Time With 19% Lower API Cost