Provides 1.7M agent interaction traces in terminus-2 format for training and evaluating agentic LLMs and RL agents. Compiled from 219 source datasets across code repair, shell, math, competitive programming and general tasks; produced with the Harbor harness.
Multilingual on-device translation model compressed to 1.25-bit via the Sherry quantization, supporting 33 languages and 1,056 directions in a 440MB package for offline mobile translation and demos.
Parallel Khasi–English sentence pairs for machine translation research focused on low-resource NLP in Northeast India. Provided as a small CSV (sentence_id, english_text, khasi_text) under CC BY‑NC 4.0 for non-commercial research use.
Curated monolingual Khasi sentence corpus (CSV) with under 1,000 sentences for language modeling, tokenization, and low-resource NLP experiments. Single-column structure (khasi_sentence) and CC BY‑NC 4.0 license — suitable for research and data-augmentation workflows, not for commercial use.
Contains ~1,973 distilled roleplay conversations with character-perspective chain-of-thought traces (<think> blocks) for fine-tuning persona-focused chat models. Includes teacher provenance, safety/review flags, and filters for NSFW/borderline samples — suited for SFT and character retention tests.
Provides 100 English–Khasi parallel sentence pairs with aligned studio-quality WAV recordings for ASR, TTS and translation evaluation; curated by Medharvix as a restricted public sample—full corpus available by request.
A prompt-only mixture of ~478k prompts designed to support antidoom-style generation and preference-data pipelines for reducing model repetition (doom loops). Prompts are stripped of answers and labels and sourced from many public datasets so it’s usable for FTPO/adapter generation but not for supervised QA evaluation.
Evaluates LLM-driven agents on long-horizon, policy-rich U.S. healthcare workflows using 75 clinical task fixtures and a 20-app MCP simulator; includes task fixtures, shared worlds, and leaderboard integration (Managed-Care handbook is gated).
A Chinese public-transit route-planning dataset for training and benchmarking LLMs that generate structured transit routes from origin–destination pairs. Releases include a large CPT corpus, SFT train/test splits, and a 30K real-world benchmark; anonymized and real testsets are provided for privacy-aware, fair evaluation.
Provides a county-harmonized corpus of U.S. municipal and county ordinance text (≈2.21M chunks) labeled for function, substantive indicator, and topic to support legal NLP, retrieval, and comparative local-law research. Includes model-assigned labels and continuous scorers (opacity, paternalism, enforcement discretion) plus coverage metadata; not exhaustive or a substitute for legal advice.
A retrieval benchmark suite focused on “oblique queries,” where relevance depends on latent attributes rather than surface keywords. Includes five tasks with large corpora, qrels (and pooled judgments), and task-specific constraints for evaluating embedding-based retrievers and reasoning-augmented retrieval.
Preview of an MoE model family (V4-Pro: 1.6T params, 49B active; V4-Flash: 284B, 13B active) built for 1M-token contexts. A hybrid attention design cuts single-token inference FLOPs to 27% and KV cache to 10% versus V3.2 at million-token length.