Alibaba Paper Debuts PoS Belief-State Framework for LLM Agents
TL;DR
- Alibaba's PoS maintains explicit belief states at inference time, with no additional training required, for long-horizon LLM agents.
- The paper reports a 22.68% relative gain on ALFWorld and a 37.89% relative gain on RCA-100 joint accuracy over same-backbone baselines.
- PoS flags a failure mode called 'Belief Trapping' — stagnation, cycles, or drift — and applies recovery tailored to the pattern.
An Alibaba Group paper posted to Hugging Face argues that LLM agents fail at long tasks because keeping a chat history is not the same as understanding the world around them.
The proposal is a framework called PoS, short for Progression of States. It runs at inference time and holds, as the agent's decision context, an explicit belief that pairs a running estimate of the current world state with the task requirements still unresolved. The authors give the characteristic failure mode a name: "Belief Trapping, where the agent continues to act without making meaningful progress toward the goal." PoS tries to catch three flavors of it — stagnation, cycles, drift — and apply recovery sized to the pattern and the kind of requirement that is blocked.
The reported numbers are relative gains over the strongest same-backbone baseline: a 22.68% lift on ALFWorld, a 37.89% lift on RCA-100 joint accuracy, and, per the abstract, "the highest overall performance on every benchmark with all three LLM backbones" across four benchmarks spanning execution and diagnosis. No additional training is involved.
The abstract does not name the three backbones or the two benchmarks beyond ALFWorld and RCA-100, and the gains are reported in relative rather than absolute terms. The paper sits in a busy week of agent coverage on our tracker, with Alibaba continuing to publish inference-time additions rather than training recipes.
Code is on the authors' GitHub.
Originally reported by huggingface.co
Read the original article →Original headline: HF Paper 'Beyond Memory' Introduces Progression-of-States Framework, Lifts ALFWorld 22.68%