Ant Group's UniSkill Hits 98.4% ALFWorld, 84.7% WebShop
TL;DR
- UniSkill reports 98.4% success on ALFWorld and 84.7% on WebShop with a shared policy that both acts and proposes edits to its own skill bank.
- Candidate skills are scored by how swapping them in shifts action log-likelihoods on prior successful versus failed trajectories, avoiding extra environment rollouts.
- Training stayed stable over 250 steps where the closest joint-training baseline, Evolving-RL, collapsed; the method still exceeds 90% ALFWorld success at 3B parameters.
UniSkill, from The Chinese University of Hong Kong and Ant Group's Qiantang Credit unit, trains one shared policy to both execute tasks and propose edits to its own skill bank, reaching 98.4% success on ALFWorld and 84.7% on WebShop in a paper posted to Hugging Face.
The trick is in how it scores a proposed skill without paying for extra environment rollouts. The authors hold the current actor and recorded trajectories fixed, then measure how swapping the retrieved skill for the proposed one changes the token-normalized action log-likelihood on a prior successful trajectory versus a prior failed one. They call the resulting signal "contrastive action feedback," and only edits that pass both a skill critic and this check are written back into the bank.
Head-to-head, the paper reports UniSkill improves ALFWorld success over GRPO by 20.8 percentage points and over SkillRL's WebShop success rate by 12.0 pp. Against Evolving-RL, the closest joint-training baseline (which does run extra rollouts), UniSkill is ahead by 5.4 pp on ALFWorld. The authors write that UniSkill "sustains stable joint training over 250 steps" where Evolving-RL "subsequently undergoes performance collapse" after strong early gains.
The method also holds up at a smaller scale: with Qwen2.5-3B-Instruct as the shared backbone, UniSkill-3B keeps improving late in training, "exceeding 90% success, and finishing above GRPO-7B." Code is at github.com/LimOkii/UniSKill. One dependency the paper does not price: the skill critic is a DeepSeek-V4-Pro model, called at every proposal.
Originally reported by huggingface.co
Read the original article →Original headline: Ant Group's UniSkill Co-Evolves Agent Policy and Skill Bank via Contrastive Action Feedback