AIAny
Icon for item

RekaDaily-10k (raw)

Provides raw, unscripted first-person household video footage for training vision and embodied AI models. Released incrementally on Hugging Face in WebDataset shards with metadata parquets under Apache‑2.0; current raw tier contains ~7,834 hours (≈397k videos).

Introduction

Why this matters

Training models that act in the physical world needs continuous, messy, first-person views — not polished, edited clips. RekaDaily-10k (raw) supplies large-scale egocentric footage recorded by paid collectors in real homes and workplaces, preserving natural pauses, lighting changes, and interruptions that are critical for learning real-world manipulation and long-horizon behaviors.

What Sets It Apart
  • Raw, uncut sessions at scale: the release is incremental but the dataset already includes ~7,834 hours (≈397,171 videos), ~9,836 WebDataset shards and ~70 TB of media — useful when you need long continuous context or realistic temporal noise. This is not a curated highlight reel; it’s the original recordings as captured.
  • Production-ready packaging for large-scale workflows: videos are distributed as ~8 GB WebDataset tar shards with paired JSON sidecars; metadata/index and browse parquets provide a full row per video (thumbnail in browse), with fields like project, flow/activities or category/subcategory, lighting, duration, fps, resolution, collector id.
  • Open, permissive license and provenance steps: released under Apache‑2.0 for commercial use and redistribution; container metadata (GPS, device IDs, timestamps) is stripped and automated PII screening applied, though consumers are warned to validate and request takedowns if needed.
Who it's for — and tradeoffs

Great fit if you are building or fine-tuning vision-language-action models, embodied agents, or multimodal perception systems that require authentic first-person sequences and raw temporal context. The dataset’s scale and WebDataset layout make it suitable for large-batch training and automated clipping/annotation pipelines.
Look elsewhere if you need fully curated, short labeled clips out of the box or guaranteed perfect PII removal: the raw tier requires substantial preprocessing (clipping, deduplication, quality filtering, annotation) and carries residual privacy/content risks despite screening. If you want immediate caption supervision, consider the processed & captioned tier (released separately) rather than the raw tier.

Information

Categories

More Items

Hugging Face

Provides agentic instruction‑tuning trajectories for software‑engineering tasks, formatted for supervised fine‑tuning and agent training. Contains multi‑file edits, tests, docs and structured agent traces (≈5,115 records, 1.9 GiB). Intended for commercial use; licensed CC‑BY 4.0 with additional permissive licenses.

Hugging Face

Provides 1.7M+ synthetic and real infographic charts paired with their tabular data for training and evaluating multimodal models on infographic understanding, chart-to-table extraction, chart code generation, and example-based chart synthesis.

Provides a year-scale multimodal benchmark and evaluation framework for on-device long-term memory in personal assistants, built from real mobile user trajectories. Tests memory construction, retrieval, updating, temporal reasoning, and implicit preference inference, and includes a knowledge-grounded synthesis pipeline to form coherent long-horizon trajectories.