huggingface.co web signal

Muon optimizer lifts agentic RL success from 0.29 to 0.55 on ALFWorld, new arXiv paper finds

Summary

A new arXiv paper studies vanilla Muon versus AdamW for sparse-reward agentic RL post-training on Qwen2.5-0.5B-Instruct in ALFWorld. Under Group-in-Group Policy Optimization, applying Muon only to hidden weight matrices raises final-window validation success from 0.290 to 0.546 (+88%), while high-rate AdamW controls retain no post-update success. At a 1e-5 learning rate, GraphGPO-Muon reaches 0.901 success and hits 0.5/0.75 milestones 30–60 updates earlier, though multi-seed and cross-task validation remain open.