Large-scale synthetic video dataset of 236,937 1080p clips (≈5,841 hours) of digital humans with per-frame metric depth and camera parameters — built as a controllable supplement for world-model pretraining, camera-motion generalization, and geometry-aware physical-AI research.
Curated monolingual Khasi sentence corpus (CSV) with under 1,000 sentences for language modeling, tokenization, and low-resource NLP experiments. Single-column structure (khasi_sentence) and CC BY‑NC 4.0 license — suitable for research and data-augmentation workflows, not for commercial use.
Contains ~1,973 distilled roleplay conversations with character-perspective chain-of-thought traces (<think> blocks) for fine-tuning persona-focused chat models. Includes teacher provenance, safety/review flags, and filters for NSFW/borderline samples — suited for SFT and character retention tests.
Provides 100 English–Khasi parallel sentence pairs with aligned studio-quality WAV recordings for ASR, TTS and translation evaluation; curated by Medharvix as a restricted public sample—full corpus available by request.
Curated multimodal training corpus for spatial intelligence: ~8.16M QA-style samples paired with ~2.72M unique images (≈1.1 TB). Provides JSONL annotations, a 1,000-sample preview, and 52 independent image archives — used to train SenseNova-SI models.
A collection of biology-focused 'mystery' tasks for benchmarking model performance on biomedical reasoning, evidence synthesis, and problem solving; curated by Anthropic and hosted on Hugging Face, designed for granular evaluation of scientific decision-making.
Provides ~85K contrastive visual question–answer pairs where each example contains an anchor and a matched counterpart (image, question, answer). Pairs span General, Reasoning, Math, Graph/Chart and OCR categories to help train and evaluate fine‑grained, faithful visual reasoning in VLMs.
Early-preview (≈1.2k rows) dataset of agentic coding prompts and unedited model responses generated by DeepSeek‑V4‑Pro, covering real-world programming tasks across many languages. Intended for research, filtering, and model evaluation rather than production training without review.
Provides the dataset and accompanying technical report for a DeepSeek project that interleaves spatial markers (points and boxes) into multimodal LLM reasoning. Includes a public subset of data and benchmarks under an MIT license; model weights are not included.
Provides 545,431 math problems with model-generated solution traces (chain-of-thought and Python tool-integrated reasoning) verified against reference answers for training and evaluating LLM mathematical reasoning. Parquet-format dataset; DeepSeek‑V4‑Pro generated traces and mixed CC BY / CC BY‑SA licensing.
Collects ML Intern coding-agent session traces as Claude‑Code‑style JSONL event streams for viewing with the Hugging Face Agent Trace Viewer. Each file is one session (messages, tool calls, outputs, timestamps); automated scrubbing is applied but no comprehensive human redaction—treat as potentially sensitive.
Instruction‑tuning dataset of 8,706 Claude Opus 4.6/4.7–generated examples where each assistant turn begins with a synthetic <think> block to emulate chain‑of‑thought. Provided as four splits (full/instruct/roleplay/code), ~17M tokens total, Apache‑2.0, not manually reviewed.