UI-Venus-2 Scales GUI Agent Across 170+ Multilingual Apps
TL;DR
- UI-Venus-2 targets mobile, web, and desktop through what the abstract calls a 'unified closed-loop reasoning-action framework.'
- The report claims coverage of more than 170 multilingual mobile apps and native desktop operating systems.
- RL training relies on trace-level and sample-level evaluators using visual keypoints and multi-model voting to score agent runs.
The Venus Team's new technical report describes UI-Venus-2 as a "general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework," according to the abstract posted on arXiv.
The paper frames the shift from benchmark models to real deployment as a reward-verification problem. The authors write that GUI agents suffer from "limited environment coverage, brittle task construction, and unreliable reward verification," and respond by scaling three axes at once: environments, tasks, and verification.
On environments, the paper claims coverage of "more than 170 multilingual mobile apps and native desktop operating systems." On verification, it describes "trace-level and sample-level evaluators with visual keypoints and multi-model voting to ensure reliable RL signals for training." The report also says it integrates "safety-aware mechanisms to ensure controlled execution of consequential actions."
The abstract names no model sizes, no benchmark scores, and no baselines.
Originally reported by paper
Read the original article →Original headline: UI-Venus-2: 9B/27B Foundation GUI Agent Spans 170+ Multilingual Apps With Keypoint-Verified RL