Supervised fine-tuning dataset of 7,716 reasoning-focused Q&A examples distilled from the DeepSeek‑V4‑Flash teacher; provided as a cleaned JSONL train split for distillation and SFT experiments.
1,000 JSONL samples containing full chain-of-thought reasoning traces and final answers produced by DeepSeek‑V4‑Pro for use in student-model distillation and quality checks. Prompts sampled from Jackrong/GLM-5.1-Reasoning-1M-Cleaned; Apache‑2.0 licensed.
Collects real-world developer–AI coding sessions with full transcripts, tool calls, agent thinking traces, Git commits, and agent vs. human code attribution. Packaged as Parquet tables (conversations, sessions, commits, checkpoints, repositories) for analysis of agent behavior and human–AI collaboration.
Provides ~55K multimodal VQA items with matched contrastive pairs and model‑generated rationales across five categories (General, Reasoning, Math, Graph/Chart, OCR), enabling research on faithful visual reasoning and robustness. Train split: 54,844 examples; license unspecified—verify before use.
Contains full chain-of-thought traces and final answers generated by DeepSeek-V4-Pro for use as distillation supervision. Key features: full CoT exposure, ~1,000 mixed-domain samples (JSONL/Parquet), Apache-2.0 license — suitable for training student models but watch for source contamination.
Cleaned dataset of reasoning-distillation examples derived from Claude Opus 4.7 outputs — 4,807 retained JSON chat rows after removing simulated-thinking, duplicates, and missing fields. Packaged for model distillation and reasoning evaluation; Apache-2.0 packaging with upstream Anthropic usage constraints.
Training dataset for byte-level language identification across 334 languages with ~2.48M paragraph samples (primarily Wikipedia and open-licensed corpora). Curated to reduce multilingual contamination, boost low-resource coverage, target frequent confusions, and preserve per-row license metadata for attribution.
Provides 1.7M agent interaction traces in terminus-2 format for training and evaluating agentic LLMs and RL agents. Compiled from 219 source datasets across code repair, shell, math, competitive programming and general tasks; produced with the Harbor harness.
Aggregates 750k+ Harbor-compatible agentic tasks from 100+ public sources (Parquet shards preserved). Includes tasks with and without verifiers for RL evaluation or SFT/datagen workflows, enabling reproducible trace generation.
Collection of 76 image-centric multimodal subdatasets (≈6.9M samples, ~39.56B estimated tokens) for training vision–language models, each published with a standardized conversation JSONL and dataset card. Media are referenced by path/URL and must be fetched separately; licensing is primarily CC-BY-4.0 with per-subdataset variations.
Parallel Khasi–English sentence pairs for machine translation research focused on low-resource NLP in Northeast India. Provided as a small CSV (sentence_id, english_text, khasi_text) under CC BY‑NC 4.0 for non-commercial research use.
Provides 104.9M curated image–text pairs with precomputed embeddings, structured annotations and pre-encoded VAE latents for text-to-image pretraining and retrieval. Combines filtered web sources and synthetic samples with multi-model re-captioning, deduplication and safety filters; Apache-2.0.