The dataset captures the scale and temporal breadth needed to study short-video ecosystems, recommendation dynamics, and large-scale multimodal training: a single, queryable Parquet archive of social video content and engagement signals spanning 2014–Oct 2026.
What Sets It Apart
- Massive, time-indexed snapshot: 5,597,462,038 public TikTok videos (snapshot date 2026-10-03) stored as one Parquet file per month to enable efficient time-sliced queries and DuckDB-style analytics.
- Rich metadata per row: posting timestamp, country, language, duration, caption, hashtags, mentions, on-screen text, sound_id, is_ad/branded_content, TikTok Shop product/seller ids, an is_ai_generated flag (nullable), and engagement counters (views, likes, comments, shares, saves, downloads) with stats_updated_at.
- Research-oriented license and access model: CC BY-NC 4.0 for research/personal use; commercial features (creator profiles, daily live updates, scraping code) provided separately by the data publisher.
- Engineering-friendly format: parquet layout and column examples show ready integration with DuckDB, Pandas, Polars, Dask and other big-data tooling for sampling, filtering, and join workflows.
Who It's For and Trade-offs
Great fit if you need: large-scale social-video corpora for recommendation research, temporal engagement analysis, multimodal model training (video+text+audio metadata), or macroscale trend studies. Expect to run analyses on big-data infrastructure; monthly Parquet files simplify time-based experiments but still require substantial storage and compute.
Look elsewhere if: you need commercial rights to redistribute or productize creator profiles and daily-updated feeds (those are restricted and offered via the publisher's commercial service), or if you require guaranteed completeness of older archived fields (some archived videos have 0/null for size, downloads, and the AI flag).
Where It Fits
This dataset sits between academic benchmark corpora and proprietary platform exports: it's large enough for industry-scale research and modeling, but the CC BY-NC license and missing commercial features mean production or redistribution use may require a commercial agreement with the data provider.