AIAny
Icon for item

RekaDaily-10k (raw)

Provides raw, unscripted first-person household video footage for training vision and embodied AI models. Released incrementally on Hugging Face in WebDataset shards with metadata parquets under Apache‑2.0; current raw tier contains ~7,834 hours (≈397k videos).

Introduction

Why this matters

Training models that act in the physical world needs continuous, messy, first-person views — not polished, edited clips. RekaDaily-10k (raw) supplies large-scale egocentric footage recorded by paid collectors in real homes and workplaces, preserving natural pauses, lighting changes, and interruptions that are critical for learning real-world manipulation and long-horizon behaviors.

What Sets It Apart
  • Raw, uncut sessions at scale: the release is incremental but the dataset already includes ~7,834 hours (≈397,171 videos), ~9,836 WebDataset shards and ~70 TB of media — useful when you need long continuous context or realistic temporal noise. This is not a curated highlight reel; it’s the original recordings as captured.
  • Production-ready packaging for large-scale workflows: videos are distributed as ~8 GB WebDataset tar shards with paired JSON sidecars; metadata/index and browse parquets provide a full row per video (thumbnail in browse), with fields like project, flow/activities or category/subcategory, lighting, duration, fps, resolution, collector id.
  • Open, permissive license and provenance steps: released under Apache‑2.0 for commercial use and redistribution; container metadata (GPS, device IDs, timestamps) is stripped and automated PII screening applied, though consumers are warned to validate and request takedowns if needed.
Who it's for — and tradeoffs

Great fit if you are building or fine-tuning vision-language-action models, embodied agents, or multimodal perception systems that require authentic first-person sequences and raw temporal context. The dataset’s scale and WebDataset layout make it suitable for large-batch training and automated clipping/annotation pipelines.
Look elsewhere if you need fully curated, short labeled clips out of the box or guaranteed perfect PII removal: the raw tier requires substantial preprocessing (clipping, deduplication, quality filtering, annotation) and carries residual privacy/content risks despite screening. If you want immediate caption supervision, consider the processed & captioned tier (released separately) rather than the raw tier.

Information

Categories

More Items

Hugging Face

Aggregated, screened corpus of 55,050 normalized Indian public-information text bodies and 65,209 source records for retrieval and question-answering. Exports include deduplicated CSV/Parquet with provenance, topic labels, extraction quality flags and a private SQLite backup.

Hugging Face

Provides 3,451 hours (2,051,810 clips) of AI‑generated 48 kHz Turkish speech with transcripts, spoken forms and per‑clip voice descriptions for TTS and ASR development. Includes 2,752 designed voices and is licensed CC BY 4.0 / CC BY‑SA 4.0 (attribution to PatientDesk AI required).

Hugging Face

Provides imagined interaction segments generated by world models for RoboTwin2.0 tasks, stored as fixed-length HDF5 chunks (21 observation frames, 20 actions, rewards and episode flags). Useful for training and evaluating world-model-based policies; currently limited to the RoboTwin2.0 subdataset.