AIAny
Icon for item

DeskForge-1M

Provides 1.21M densely annotated desktop screenshots and 159.7M element instances for training and evaluating GUI grounding and screen-parsing models. Includes per-element accessibility-derived annotations, 917K recorded click transitions, multi-application scenes across seven appearance presets and resolutions; distributed as WebDataset shards with Parquet indexes.

Introduction

Modern computer-use agents break down when action targets are ambiguous on complex multi-window desktops. DeskForge-1M attacks this bottleneck by programmatically composing real applications into controlled scenes and producing dense, per-element annotations plus recorded click transitions, yielding supervision that targets both visual grounding and action effects.

Key Features
  • Scale and controlled diversity — 1.21M annotated observations, 159.7M retained element instances, and 917K recorded click transitions across about 323K scenes. Scenes span 19 real Linux applications, seven appearance presets and seven display resolutions, with scene-level splits that hold out apps, themes, or resolutions for robust generalization testing.
  • Dense, accessibility-rooted annotations — every visible element is reconciled from accessibility trees, screenshots and window geometry and annotated with type, role, visible text and fragments, occlusion state, hierarchy, owning window, interaction properties and a stable identifier. This enables detectors and VLMs to learn fine-grained screen structure, not just sparse targets.
  • Action supervision and instructions — click transitions store before/after frames, target geometry, measured effects (appeared, moved, text/state changes) and eligibility flags. 663k recorded clicks have natural-language instructions synthesized for grounding training, supporting both single-step grounding and longer-horizon planning research.
  • Engineering-friendly distribution — data served as WebDataset shards for streaming, with Parquet index tables (observations, transitions, instructions, scenes) for selective access and random lookup; a compact transitions_preview subset aids quick browsing and demos.
Who it's for and trade-offs

Great fit if you need large-scale, structured supervision for GUI grounding, screen parsing, or training embodied/agent models that act on desktop UIs — especially when per-element interaction properties and action effects matter. The dataset is crafted to evaluate generalization across unseen apps, themes, and resolutions. Look elsewhere if you require native Windows/macOS captures (the backend is a single Linux/Xfce compositor and Windows/macOS styles are emulated), goal-directed human demonstrations (actions are from undirected exploration), or fully human-written instructions (many instructions are synthesized from recorded clicks). Annotation audits report very high automated quality (sampled element correctness ~99.8% and instruction soundness ~97.8%), but text rendered by applications is included verbatim and may reflect live web content or proprietary assets.

Information

  • Websitehuggingface.co
  • OrganizationsETH Zurich, IBM Research Zurich, Microsoft
  • AuthorsA. Said Gurbuz, Ahmed Nassar, Sunghwan Hong, Marc Pollefeys, Peter W. J. Staar
  • Published date2026/08/25

Categories

More Items

Hugging Face

De-identified longitudinal multimodal CT dataset for multicancer screening that pairs ~24k CT volumes with radiology reports and voxel-wise tumor annotations across 13 cancer types. Designed for longitudinal disease modeling, detection/segmentation and vision–language research; CC BY‑NC‑ND 4.0 for non-commercial use.

Hugging Face

Provides 98,877 historical newspaper page images (1700s–1940s) paired with ALTO OCR, per-word confidences and line/word bounding boxes in image pixels — ready for line-level OCR training, OCR quality estimation, re‑OCR comparisons and layout analysis. OCR is library-produced (silver); image resolutions and OCR quality vary.

Hugging Face

Generates humanoid robot motion references that preserve object contact locations/timing by solving windowed trajectory optimizations against contact targets in the object frame. Releases retargeted trajectories for two Unitree robots across 75 objects (≈13.9k robot–motion pairs); CC BY‑NC‑SA 4.0.