Collects real-world developer–AI coding sessions with full transcripts, tool calls, agent thinking traces, Git commits, and agent vs. human code attribution. Packaged as Parquet tables (conversations, sessions, commits, checkpoints, repositories) for analysis of agent behavior and human–AI collaboration.
Contains full chain-of-thought traces and final answers generated by DeepSeek-V4-Pro for use as distillation supervision. Key features: full CoT exposure, ~1,000 mixed-domain samples (JSONL/Parquet), Apache-2.0 license — suitable for training student models but watch for source contamination.
Cleaned dataset of reasoning-distillation examples derived from Claude Opus 4.7 outputs — 4,807 retained JSON chat rows after removing simulated-thinking, duplicates, and missing fields. Packaged for model distillation and reasoning evaluation; Apache-2.0 packaging with upstream Anthropic usage constraints.
Training dataset for byte-level language identification across 334 languages with ~2.48M paragraph samples (primarily Wikipedia and open-licensed corpora). Curated to reduce multilingual contamination, boost low-resource coverage, target frequent confusions, and preserve per-row license metadata for attribution.
Provides 1.7M agent interaction traces in terminus-2 format for training and evaluating agentic LLMs and RL agents. Compiled from 219 source datasets across code repair, shell, math, competitive programming and general tasks; produced with the Harbor harness.
Multilingual on-device translation model compressed to 1.25-bit via the Sherry quantization, supporting 33 languages and 1,056 directions in a 440MB package for offline mobile translation and demos.
Parallel Khasi–English sentence pairs for machine translation research focused on low-resource NLP in Northeast India. Provided as a small CSV (sentence_id, english_text, khasi_text) under CC BY‑NC 4.0 for non-commercial research use.
Curated monolingual Khasi sentence corpus (CSV) with under 1,000 sentences for language modeling, tokenization, and low-resource NLP experiments. Single-column structure (khasi_sentence) and CC BY‑NC 4.0 license — suitable for research and data-augmentation workflows, not for commercial use.
Contains ~1,973 distilled roleplay conversations with character-perspective chain-of-thought traces (<think> blocks) for fine-tuning persona-focused chat models. Includes teacher provenance, safety/review flags, and filters for NSFW/borderline samples — suited for SFT and character retention tests.
Provides 100 English–Khasi parallel sentence pairs with aligned studio-quality WAV recordings for ASR, TTS and translation evaluation; curated by Medharvix as a restricted public sample—full corpus available by request.
A prompt-only mixture of ~478k prompts designed to support antidoom-style generation and preference-data pipelines for reducing model repetition (doom loops). Prompts are stripped of answers and labels and sourced from many public datasets so it’s usable for FTPO/adapter generation but not for supervised QA evaluation.
Evaluates LLM-driven agents on long-horizon, policy-rich U.S. healthcare workflows using 75 clinical task fixtures and a 20-app MCP simulator; includes task fixtures, shared worlds, and leaderboard integration (Managed-Care handbook is gated).