paper web signal

Business Arena finds 9x net-worth gap across 15 frontier LLMs

TL;DR

  • A new benchmark called Business Arena had 15 frontier LLMs run cross-border shops on Alibaba.com data and found a ninefold spread in final net worth.
  • Even the best-performing model finished behind human-designed strategies, which the authors say shows business operation remains challenging for LLM agents.
  • The evaluation goes beyond profit to skill-level metrics and action-level attribution, tracing which sourcing and pricing decisions created or destroyed value.

Chat leaderboards do not tell you whether an LLM can actually run a business. A new arXiv paper introducing Business Arena tries to answer that question directly, by dropping 15 frontier models into a simulated cross-border shop and measuring what they end up with in the till.

The setup, per the abstract, is that "an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon," with market conditions "grounded in real Alibaba.com sourcing data and market conditions calibrated from authoritative sources." The authors evaluate 15 frontier models and report "a ninefold difference in mean final net worth" across the field. That is roughly an order of magnitude between the best and worst operator, on a task where every model in the pool can plausibly write a fluent product listing.

The harder result is that "even the best model falls behind human-designed strategies," which the authors read as evidence that "business operation remains challenging for LLM agents." Alongside profit, the benchmark reports skill-level metrics and action-level attribution, tracing which specific sourcing, pricing, and recovery moves created or destroyed value across a run. That is the part that makes this more interesting than another leaderboard: it isolates where a model loses money, not just how much.

The abstract holds a lot back. It does not name which 15 frontier models were tested, so the reader cannot yet see who topped or tanked the ranking, and it does not describe how the human-designed strategies baseline was built or by whom. The setup is also anchored to Alibaba.com wholesale conditions, so whether the same ninefold spread would appear in other verticals or geographies is genuinely unknown from this paper alone.

For anyone selling or buying agentic commerce tooling, this is the kind of evaluation to point at instead of a chat benchmark. If a vendor claims their model can operate a store, ask for the profit curve.