X-Tree lifts WebArena agent scores up to 4.5% via BPE merge
TL;DR
- X-Tree recovers hierarchical structure from agent trajectories by merging frequent, successful action spans into a reusable skill tree, no LLM calls required.
- The method lifts success rate by up to 4.5% on WebArena, 5.8 points on ScienceWorld, and 4.1% on WebShop at matched data and compute budgets.
- The authors plug X-Tree into three training settings: offline RL, online RLVR with an adaptive skill bonus, and on-policy self-distillation.
A new training recipe called X-Tree improves web-agent success rates by up to 4.5% on WebArena, 5.8 percentage points on ScienceWorld, and 4.1% on WebShop at matched data and compute budgets, according to the paper on arXiv from Sitao Cheng, Xunjian Yin, Victor Zhong and co-authors, submitted September 26, 2026.
The pitch is that current agent training throws away structure. "Multi-step agents are trained on flat action streams: SFT and RLVR weight every token uniformly and ignore the sub-procedures that recur across tasks, the hierarchy that lets humans plan top-down from reusable routines," the authors write. X-Tree instead borrows from byte-pair encoding: it scores adjacent action spans by how often they recur, how long they are, and how often they appear in successful episodes, then merges the best pairs into tree nodes, and merges nodes into larger nodes. No LLM is called to build the vocabulary.
The tree then plugs into three training regimes. In offline RL, "each X-Tree node acts as one training instance," with the policy rolled out from the gold prefix. In online RLVR, "X-Tree adds a bonus for each X-Tree skill a rollout executes with an adaptive weight that supplies signal while verifier signal is scarce." In on-policy self-distillation, X-Tree replaces an LLM-written skill bank as the teacher's privileged context.
The abstract reports deltas only, not absolute success rates, and does not name which base models were tested at the three scales. The gains come from a single paper with no third-party replication yet.
Originally reported by paper
Read the original article →Original headline: X-Tree Gains +4.5% on WebArena by Building a Reusable Skill Vocabulary From Agent Trajectories