DeskForge trains computer-use agents on 1.2M desktop scenes
TL;DR
- DeskForge builds a 1.2M-observation desktop dataset with 159.7M element instances by composing and exploring real applications.
- Fine-tuning Qwen3.5-4B on 200K grounding examples lifted ScreenSpot-Pro by 11.51 points and OSWorld-G by 10.11 points.
- On WebArena-Infinity the fine-tuned Qwen3.5-4B solved 50 of 119 tasks, up from 31, with similar gains on OpenApps.
A new arXiv preprint from A. Said Gurbuz, Ahmed Nassar, Sunghwan Hong, Marc Pollefeys and Peter W. J. Staar reports that fine-tuning Qwen3.5-4B on 200K grounding examples drawn from a 1.2M-observation desktop corpus lifts ScreenSpot-Pro accuracy by 11.51 percentage points and OSWorld-G by 10.11 points.
The corpus, DeskForge, contains 159.7M element instances and is built by a system that "composes and explores real applications to generate large-scale supervision for computer-use agents," according to the paper. Four vision-language models were tested. Beyond pure grounding benchmarks, the same fine-tuned Qwen3.5-4B solved 50 of 119 WebArena-Infinity tasks, up from 31, and 15 of 100 OpenApps tasks, up from 3.
The authors say framework code, dataset and fine-tuned models are available via a project page. The abstract does not break out per-model deltas for the three other VLMs tested, nor does it report numbers on closed systems.
Originally reported by paper
Read the original article →Original headline: DeskForge-1M: 1.2M Annotated Desktop Screenshots Push Computer-Use Agents +11pp on ScreenSpot-Pro