Contains 8,124 reasoning conversations (extended-thinking + final responses) generated by Anthropic Claude Opus 4.7 for distillation into open-source LLMs. Each row stores the prompt, thinking trace, final answer and usage metadata; packaged under Apache‑2.0.
Synthetic JSON dataset of model-generated prompts and step-by-step reasoning traces (≈90k rows, ~75M tokens) created with Claude Sonnet 4.6 and cross-checked by Gemini 3.1 Pro — intended for training or fine-tuning LLMs on natural reasoning, multi-domain code/math, and instruction following. Hosted on Hugging Face, MIT license.
Provides 2,405 chain-of-thought reasoning traces generated by Claude Opus 4.7 for hard math, science, and formal problems. Each record pairs a problem with the model's full <think> working and a polished answer; available as parquet splits for non-commercial research under Anthropic's usage policy.
Synthetic Korean-language persona dataset for training and evaluating conversational and generative models — 1M records (≈7M persona entries) with 26 fields aligned to South Korea’s demographic distributions. Built with NeMo Data Designer and released under CC BY 4.0.
A 1.4M image–text style dataset for text-to-image generation and style transfer, produced by mapping 170K curated style prompts to 400K content prompts via Qwen-Image to yield strong intra-style consistency. Designed for training and evaluating style-aware generative models; license: other.
Benchmark dataset for evaluating clinician-facing chat assistants: physician-authored conversations plus rubric items, use-case and difficulty labels, specialty metadata, and a built-in canary to reduce benchmark contamination. Hosted on Hugging Face under an MIT license.
Provides a large-scale multimodal embodied dataset (vision, depth, hand/arm kinematics, tactile) captured with an exoskeleton glove and egocentric sensors; organized as clip-level Zarr volumes for manipulation, imitation learning, and vision–action research. Includes both high-precision glove measurements and natural bare-hand clips; sizable storage required.
Provides instruction-based (before, after) structured 3D latents (SLAT) with aligned RGB views and natural-language edit prompts for training and evaluating instruction-following 3D editing models. Covers part-level semantic edits across seven edit types (deletion, addition, modification, scale, material, color, global) and supplies shard-based NPZ assets and loader code.
Provides 150,000 synthetic Vietnamese patient personas to condition clinical text generation. Each persona bundles demographics, socioeconomic context, health and behavior fields, and prompt-ready narratives; intended for research and simulation, not clinical decision-making.
Benchmarks ASR on long-form English call-center conversations with wide accent coverage; 128.6 hours across 14 accent groups and 16 service domains, designed for segmentation-sensitive evaluation and intended for evaluation/analysis (CC BY‑SA 4.0).
Provides a 30K+ problem multimodal, multilingual dataset of Olympiad-level math problems with expert solutions and a math-aware retrieval benchmark—includes images, hierarchical topics, provenance from official booklets, and LLM-assisted metadata (v0, CC BY 4.0).
Provides satellite image tiles paired with per-tile land-cover captions and bounding-box overlays in SFT-compatible JSONL for supervised fine-tuning. Includes RGB chips, optional Mapbox context, metadata, and train/validation/test splits derived from Sentinel‑2 and Earth Engine labels.