AIAny
Icon for item

Gaming Dataset (gaming-1)

Provides ~494.7 hours of trimmed native PC/console gameplay screen recordings organized by game, with per-session clips plus input and per-frame event annotations. Each workflow includes clip.mp4, events.json, frame_events.json, and metadata — suitable for training vision-action, behavior-cloning, and gameplay understanding models.

Introduction

Large, focused gameplay corpora are valuable because noisy streaming footage (overlays, desktop, launchers) confounds model learning. This dataset supplies nearly 495 hours of gameplay-only clips trimmed to remove launchers, desktop screens, and viewing/streaming artifacts while keeping in-game menus, lobbies, loading, and cutscenes — making it directly usable for vision-action research and gameplay analysis.

What Sets It Apart
  • Session-oriented layout: each workflow folder contains a 30fps H.264 clip.mp4 plus events.json (input/app events rebased to clip timeline), frame_events.json (per-frame event view), and a metadata.json summarizing duration and event counts — so you get aligned video+event traces out of the box.
  • Wide game coverage with concentrated scale: 776 workflows across 168 distinct games totaling 494.7 hours; top titles include Valorant (102.2 h), Minecraft (41.3 h), and GTA V (34.7 h) — useful for both breadth and per-title depth.
  • Trimmed to pure gameplay: non-gameplay content (launchers, desktop, watching/streaming) removed, reducing label noise for behavior cloning, imitation learning, and supervised vision-action training.
  • Practical formats: 30fps CFR H.264 video and NDJSON event files enable straightforward ingestion with standard tools (datasets, pandas, polars) and simple conversion pipelines.
Who it's for — and tradeoffs

Great fit if you need medium-scale, gameplay-focused video + input traces for training vision-action models, behavior cloning, imitation learning, gameplay understanding, or dataset augmentation. The dataset's per-session organization and event alignment lower preprocessing overhead.

Look elsewhere if you require an explicit open license (the Hugging Face card lists no license), ultra-high frame-rate / hardware-synchronized HID traces, or full desktop recordings (this collection is trimmed to in-game footage). Also note the dataset is Windows-heavy (~489.3h) with limited macOS coverage (~5.3h), which may bias platform-specific behaviors.

Information

Categories

More Items

Hugging Face

Provides 7,366 recorded agent trajectories from H Company’s Holo4 benchmark runs, with step-level reasoning, actions, tool results, token usage and screenshots for replay and analysis. Bundled as JSON and image files for per-trajectory inspection and automated replay; released under Apache 2.0.

Hugging Face

Provides 150,000 source‑grounded decision examples for training models that pick options, judge yes/no propositions, or assign ordered scores. Each row pairs a 'state' with JEV-style typed questions (CHOICE/NOUL/SCORE); multiple configs and train/test splits are included, license mixed/unknown.

Hugging Face

Provides 23,625 semi-structured smart-contract audit findings (title, description, PoC, recommendation, normalized severity) for defensive-security research; requires cleaning, deduplication, and PoC filtering before model training.