AIAny
Icon for item

TikTok Videos (5.6 billion)

Provides a monthly Parquet snapshot of ~5.6 billion public TikTok videos (2014–Oct 2026), including captions, hashtags, sounds, engagement metrics and TikTok Shop links. Designed for large-scale querying (DuckDB/Pandas/Polars); licensed CC BY-NC 4.0 for research and personal use.

Introduction

The dataset captures the scale and temporal breadth needed to study short-video ecosystems, recommendation dynamics, and large-scale multimodal training: a single, queryable Parquet archive of social video content and engagement signals spanning 2014–Oct 2026.

What Sets It Apart
  • Massive, time-indexed snapshot: 5,597,462,038 public TikTok videos (snapshot date 2026-10-03) stored as one Parquet file per month to enable efficient time-sliced queries and DuckDB-style analytics.
  • Rich metadata per row: posting timestamp, country, language, duration, caption, hashtags, mentions, on-screen text, sound_id, is_ad/branded_content, TikTok Shop product/seller ids, an is_ai_generated flag (nullable), and engagement counters (views, likes, comments, shares, saves, downloads) with stats_updated_at.
  • Research-oriented license and access model: CC BY-NC 4.0 for research/personal use; commercial features (creator profiles, daily live updates, scraping code) provided separately by the data publisher.
  • Engineering-friendly format: parquet layout and column examples show ready integration with DuckDB, Pandas, Polars, Dask and other big-data tooling for sampling, filtering, and join workflows.
Who It's For and Trade-offs

Great fit if you need: large-scale social-video corpora for recommendation research, temporal engagement analysis, multimodal model training (video+text+audio metadata), or macroscale trend studies. Expect to run analyses on big-data infrastructure; monthly Parquet files simplify time-based experiments but still require substantial storage and compute.

Look elsewhere if: you need commercial rights to redistribute or productize creator profiles and daily-updated feeds (those are restricted and offered via the publisher's commercial service), or if you require guaranteed completeness of older archived fields (some archived videos have 0/null for size, downloads, and the AI flag).

Where It Fits

This dataset sits between academic benchmark corpora and proprietary platform exports: it's large enough for industry-scale research and modeling, but the CC BY-NC license and missing commercial features mean production or redistribution use may require a commercial agreement with the data provider.

Information

Categories

More Items

Hugging Face

Generates humanoid robot motion references that preserve object contact locations/timing by solving windowed trajectory optimizations against contact targets in the object frame. Releases retargeted trajectories for two Unitree robots across 75 objects (≈13.9k robot–motion pairs); CC BY‑NC‑SA 4.0.

Hugging Face

Collection of 1.44M unique Turkish voice‑assistant sentences (≈1,819 hours estimated), normalized for TTS and organized by service domains (appointments, banking, e‑commerce). Designed for training and evaluating TTS and text-generation models; licensed CC BY 4.0 with required attribution.

Turns solved protein structures into FoldingCorpus and Fold2Reason — a post-training recipe that supervises an LLM with discrete structural Q&A plus continuous 3D geometry to improve spatial, graph and scientific reasoning; reports consistent gains across 10 benchmarks.