paper web signal

UI-Venus-2 Scales GUI Agent Across 170+ Multilingual Apps

TL;DR

  • UI-Venus-2 targets mobile, web, and desktop through what the abstract calls a 'unified closed-loop reasoning-action framework.'
  • The report claims coverage of more than 170 multilingual mobile apps and native desktop operating systems.
  • RL training relies on trace-level and sample-level evaluators using visual keypoints and multi-model voting to score agent runs.

The Venus Team's new technical report describes UI-Venus-2 as a "general-purpose foundation GUI agent designed to operate across mobile, web, and desktop environments through a unified closed-loop reasoning-action framework," according to the abstract posted on arXiv.

The paper frames the shift from benchmark models to real deployment as a reward-verification problem. The authors write that GUI agents suffer from "limited environment coverage, brittle task construction, and unreliable reward verification," and respond by scaling three axes at once: environments, tasks, and verification.

On environments, the paper claims coverage of "more than 170 multilingual mobile apps and native desktop operating systems." On verification, it describes "trace-level and sample-level evaluators with visual keypoints and multi-model voting to ensure reliable RL signals for training." The report also says it integrates "safety-aware mechanisms to ensure controlled execution of consequential actions."

The abstract names no model sizes, no benchmark scores, and no baselines.