Qwen-UI-Agent Report Claims Lead on Real-Device GUI Benchmarks
TL;DR
- Qwen-UI-Agent reports 82.1% on MobileWorld, 92.2% on MobileWorld-Real, and 97.5% on AndroidDaily, framed as state-of-the-art mobile performance.
- Desktop and web results land at 79.5% on OSWorld-Verified, 73.6% on WebArena, and 81.5% on ScreenSpot-Pro grounding.
- The report positions the agent against Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol on real-device execution tasks.
A new technical report posted to arXiv from the Qwen-UI-Agent team makes a bold claim: their foundation GUI agent beats leading frontier closed models across real-world computer-use benchmarks spanning mobile, desktop, and web. The reported numbers are aggressive on mobile, with 82.1% on MobileWorld, 92.2% on MobileWorld-Real, and 97.5% on AndroidDaily.
Away from mobile, the picture is more mixed. Desktop use is reported at 79.5% on OSWorld-Verified, with only 40.0% partial-progress on the harder OSWorld-v2. Web browsing lands at 73.6% on WebArena, and GUI grounding at 81.5% on ScreenSpot-Pro. The report frames these as competitive with frontier models on those domains, while positioning the mobile results as state-of-the-art.
If the story is less about any single benchmark and more about who is doing the beating, the report positions Qwen-UI-Agent against Opus 4.8, Gemini 3.1 Pro, and GPT-5.6 Sol on the same real-device execution tasks. For a foundation agent from this lineage to reportedly match or surpass the frontier proprietary stacks on agent-style computer use, not just chat, is the sort of parity claim that changes what teams evaluating automation agents look at first.
The honest caveat is that this is a self-reported technical report, not an independent evaluation, and the OSWorld-v2 figure is a partial-progress score rather than full task success. Real-world workflows tend to break in exactly those long-horizon desktop and web scenarios where the gap between partial progress and full success is wide. The report also does not spell out cost per task, on-device latency, or how the agent behaves on out-of-distribution GUIs, which is where paper leaderboards and shipped products usually diverge.
The upside for teams building agents is optionality. If the Qwen-UI-Agent line is genuinely competitive on execution across platforms, there is now a credible non-Western foundation option for cross-platform automation pipelines that do not want to be single-sourced on OpenAI, Anthropic, or Google.
Originally reported by paper
Read the original article →Original headline: Qwen-UI-Agent Beats Opus 4.8, Gemini 3.1 Pro, GPT-5.6 Sol Across GUI Benchmarks