Generates controllable multilingual speech from text with nine predefined timbres and custom-voice control; supports voice design, quick voice cloning and low-latency streaming (first audio packet after a single character), suitable for real-time TTS and voice-design workflows.
Large-scale mathematical reasoning dataset of model-generated solution trajectories produced with and without Python Tool-Integrated Reasoning (TIR), with final answers verified against reference solutions. Contains ~3.64M JSONL training samples (~144 GB) and per-source CC-BY / CC-BY-SA licensing; intended for training and evaluating tool-augmented mathematical reasoning in LLMs.
Provides a machine-readable collection of 5,426 open and historically significant mathematical problems with LaTeX statements, structured metadata and curated per-problem AI-assisted research notes. Includes difficulty labels, canonical problem sets (Millennium, Hilbert, Erdős) and files optimized for benchmarking math reasoning.
Provides a physical reconstruction benchmark of OmniDocBench v1.5 by producing five real-world photographic variants (Scanning, Warping, Screen‑Photography, Illumination, Skew) for each of 1,355 pages, inheriting original ground-truth to enable controlled, scenario-wise evaluation of document parsing robustness.
Processed, multilingual news corpus of 1.357B articles extracted from Common Crawl CC‑News with per-article WARC provenance. Includes trafilatura-extracted bodies, language labels (GlotLID & CommonLingua), IPTC topic tags, monthly Parquet shards and a companion FM-index for sub-10ms substring queries; bulk text access is gated for academic research.
Generates anime-style and other non-photorealistic illustrations from text prompts. A 2B-parameter diffusion base preview trained on millions of anime images (and ~800k non-anime art) and released under a non-commercial license; best used in ComfyUI around ~1MP resolution.
Multimodal OCR and document-understanding toolkit for recognizing complex layouts, tables, formulas and code. Uses Multi-Token Prediction and stable RL for better training; ships as a 0.9B-parameter model with a Python SDK and deployment guides for vLLM, SGLang and Ollama.
Provides 100 real-world, open-ended research tasks paired with expert-written rubrics (around 40 weighted criteria per task) to evaluate long-form, web-browsing research agents on factual accuracy, analysis depth, presentation, and citation quality.
Turns natural-language directions into end-to-end video editing workflows: LLM-powered planning, media search/organization, ASR rough-cut, and reusable Style Skills for consistent storytelling. Integrates agent Skills (OpenClaw/Claude Code) and optional AIGC transitions.
Provides L3 refined synthetic training data by converting high-quality web corpora into Q&A pairs and multi-style rewrites; supplies 400B+ English and 200B+ Chinese tokens for late-stage LLM pretraining and decay-phase training.
Generates high‑fidelity, expressive speech and environmental sounds from text. The MOSS‑TTS Family provides specialized models for long‑form TTS, multi‑speaker dialogue, voice design and realtime streaming, plus torch‑free inference paths (llama.cpp / ONNX) and Hugging Face releases.
Filtered subset of the OPUS 4.6 parallel corpus that isolates reasoning-related translation examples and removes 979 refusals, providing a cleaner 3,000×-filtered dataset for training or evaluating NLP models focused on reasoning in translation.