Filtered subset of the OPUS 4.6 parallel corpus that isolates reasoning-related translation examples and removes 979 refusals, providing a cleaner 3,000×-filtered dataset for training or evaluating NLP models focused on reasoning in translation.
Cleaned reasoning dataset of problem→thinking→solution triplets derived from Opus 4.6, provided in Parquet with ~2,160 cleaned rows (original 3,305). Filters remove empty/short/refusal/non‑substantive responses; hosted on Hugging Face under Apache‑2.0.
Provides 1.06M web interaction trajectories (state, action, next_state) represented primarily as A11y trees for training browser world models and web agents. Covers diverse real‑web domains, English/Chinese pages, and long contexts (up to 30K tokens); residual PII and dynamic content may limit reproducibility.
Paired brain MRI scans and radiology text annotations for multimodal vision–language research. Provides image-level labels and image–text pairs suited for VQA, classification, and image-to-text tasks; CC BY-NC-SA 4.0 and ~10K–100K samples — research/non-commercial use.
Provides 6,000 runnable, operator-level PyTorch tasks for training and evaluating CUDA kernel generation models; each sample includes executable code, operator descriptors, and provenance tags, with execution-driven filtering to ensure reproducibility and contamination control.
Provides an annotated multimodal human-motion dataset for language-to-action and robotics research, with BVH and MuJoCo files plus recordings targeted at Unitree-G1 and NVIDIA-SOMA platforms. Covers locomotion, gestures, dance and object interaction with English annotations and 100K–1M samples.
JSONL dataset of Claude Opus 4.6 chain-of-thought traces paired with high-difficulty math and logic problems for supervised fine-tuning and distillation; exposes step-by-step reasoning to teach process-oriented problem solving and improve math/logic accuracy in smaller LLMs.
Provides 10 million synchronized egocentric experience episodes with structured 3D/4D multimodal annotations — 2.88B RGB frames, 720M depth frames, 576M pose/mocap frames and ~1PB total. Designed for embodied AI, robotics, and multimodal pretraining; research-only, gated access.
Provides a 1,000-row sample user–item interaction Parquet for the TAAC2026 recommendation task, using a flat column layout with 120 top-level columns (IDs, labels, user/item int & dense features, and four-domain behavioral sequences). Updated 2026-04-10.
Transforms articulated 3D asset creation into a programmatic, LLM-driven code-generation workflow that produces objects with semantic parts, robust geometry, and physical joints. Includes CLI generation, a local viewer, and pipelines for large-scale dataset contribution.
A 228,557-example dataset of reasoning traces segmented into blocks with iterative, compressed "memento" summaries so LLMs can learn to manage long context. Includes a training-ready subset and a `full` subset with sentence/block-level annotations for research and SFT.
Large-scale mid-training corpora for multimodal models: 10,809 ~60s video shards, caption splits (30s/60s/180s/>10min), 84 spatial-reasoning shards, and CSV mappings to source YouTube IDs. Small Parquet preview configs are provided for schema inspection.