paper web signal

DeskForge trains computer-use agents on 1.2M desktop scenes

TL;DR

  • DeskForge builds a 1.2M-observation desktop dataset with 159.7M element instances by composing and exploring real applications.
  • Fine-tuning Qwen3.5-4B on 200K grounding examples lifted ScreenSpot-Pro by 11.51 points and OSWorld-G by 10.11 points.
  • On WebArena-Infinity the fine-tuned Qwen3.5-4B solved 50 of 119 tasks, up from 31, with similar gains on OpenApps.

A new arXiv preprint from A. Said Gurbuz, Ahmed Nassar, Sunghwan Hong, Marc Pollefeys and Peter W. J. Staar reports that fine-tuning Qwen3.5-4B on 200K grounding examples drawn from a 1.2M-observation desktop corpus lifts ScreenSpot-Pro accuracy by 11.51 percentage points and OSWorld-G by 10.11 points.

The corpus, DeskForge, contains 159.7M element instances and is built by a system that "composes and explores real applications to generate large-scale supervision for computer-use agents," according to the paper. Four vision-language models were tested. Beyond pure grounding benchmarks, the same fine-tuned Qwen3.5-4B solved 50 of 119 WebArena-Infinity tasks, up from 31, and 15 of 100 OpenApps tasks, up from 3.

The authors say framework code, dataset and fine-tuned models are available via a project page. The abstract does not break out per-model deltas for the three other VLMs tested, nor does it report numbers on closed systems.