Found first: a primary source the press has not covered yet.
Business Arena tested 15 frontier models by putting each one in charge of a simulated cross-border shop. Mean final net worth varied ninefold. Even the strongest model finished behind human-designed strategies.
What the source says
Yijun Pan and seven coauthors built the environment around real Alibaba.com sourcing data and market conditions calibrated from authoritative sources. Agents had to buy from suppliers, sell to customers, commit capital under uncertainty, respond to delayed outcomes, and satisfy regulatory requirements before trading. The benchmark measures profit, then traces gains and losses back to sourcing, pricing, service, and recovery decisions. Mechanism ablations test whether strong results came from genuine business performance or simulator shortcuts. The models developed distinct operating styles, including premium selling, high-turnover wholesaling, and customer-service specialization.
Why it matters
Most agent benchmarks reward a correct action or a completed workflow. A business exposes whether those actions add up over time. Early pricing choices constrain later inventory, weak sourcing narrows future margins, and delayed outcomes make a plausible decision hard to judge in isolation. The ninefold spread shows that models with broadly similar frontier labels can produce radically different economic results under the same conditions. The gap behind the human-designed strategies also sets a practical limit on claims that current agents can independently operate a business.