Modern computer-use agents break down when action targets are ambiguous on complex multi-window desktops. DeskForge-1M attacks this bottleneck by programmatically composing real applications into controlled scenes and producing dense, per-element annotations plus recorded click transitions, yielding supervision that targets both visual grounding and action effects.
Key Features
- Scale and controlled diversity — 1.21M annotated observations, 159.7M retained element instances, and 917K recorded click transitions across about 323K scenes. Scenes span 19 real Linux applications, seven appearance presets and seven display resolutions, with scene-level splits that hold out apps, themes, or resolutions for robust generalization testing.
- Dense, accessibility-rooted annotations — every visible element is reconciled from accessibility trees, screenshots and window geometry and annotated with type, role, visible text and fragments, occlusion state, hierarchy, owning window, interaction properties and a stable identifier. This enables detectors and VLMs to learn fine-grained screen structure, not just sparse targets.
- Action supervision and instructions — click transitions store before/after frames, target geometry, measured effects (appeared, moved, text/state changes) and eligibility flags. 663k recorded clicks have natural-language instructions synthesized for grounding training, supporting both single-step grounding and longer-horizon planning research.
- Engineering-friendly distribution — data served as WebDataset shards for streaming, with Parquet index tables (observations, transitions, instructions, scenes) for selective access and random lookup; a compact transitions_preview subset aids quick browsing and demos.
Who it's for and trade-offs
Great fit if you need large-scale, structured supervision for GUI grounding, screen parsing, or training embodied/agent models that act on desktop UIs — especially when per-element interaction properties and action effects matter. The dataset is crafted to evaluate generalization across unseen apps, themes, and resolutions. Look elsewhere if you require native Windows/macOS captures (the backend is a single Linux/Xfce compositor and Windows/macOS styles are emulated), goal-directed human demonstrations (actions are from undirected exploration), or fully human-written instructions (many instructions are synthesized from recorded clicks). Annotation audits report very high automated quality (sampled element correctness ~99.8% and instruction soundness ~97.8%), but text rendered by applications is included verbatim and may reflect live web content or proprietary assets.