Measures how well LLMs and agent-driven workflows prepare supervised training data end-to-end by jointly benchmarking data construction and data-quality evaluation across six domains, using a downstream-grounded protocol and new metrics.
Dataset of 5,000 reconstructed chain-of-thought samples produced by trace‑inversion from Claude‑opus‑4.7 summaries — packaged for SFT/DPO fine‑tuning. Key features: reconstructed CoT traces, multilingual prompts, gzip .jsonl format. Best used for reasoning distillation and model-level supervision; synthetic traces may need extra verification.
Provides 9,000 reconstructed chain-of-thought (CoT) SFT examples produced by trace inversion from Claude Opus 4.6 outputs for fine-tuning reasoning-capable LLMs. Multilingual, packaged as .jsonl.gz and SFT/DPO-ready; verify numeric/code cases before training.
RL training dataset for long-context language-model fine-tuning with ~23K samples and nine reward types, provided in Parquet with bilingual ground-truth and reward metadata for direct RL/bench evaluation.
Supervised fine-tuning dataset of instruction-style examples in English and Chinese covering generation, QA, reasoning, math and code — targeted for SFT of 10–100B-parameter LLMs. Associated with arXiv:2602.09003; first published May 21, 2026.
Provides 100,000 generated low-quality↔high-quality image pairs created with modern multi-frame/multi-modal models to boost generalization of image restoration methods; includes train/test JSONL lists, baseline training code, and pretrained checkpoints under CC BY‑NC‑ND 4.0.
Creator-centric benchmark for evaluating text-to-image models with 1,000 bilingual prompts and a 3-level, 56-facet taxonomy. Includes a trained Q-Judger judge model and leaderboard-ready evaluation scripts to surface gaps in real-world fidelity and creative generation.
Provides a 289-case (1,058-turn) multi-turn benchmark that evaluates interactive video world models across 22 metrics and five dimensions (quality, setting, interaction, consistency, physics). Includes first-/third-person and navigation splits plus a 20-model leaderboard for head-to-head comparisons.
Parallel Chinese→Vietnamese dataset of webnovel (xianxia) text provided in JSON for NMT training and teacher-student distillation. In-domain, ~100K–1M examples with CC-BY-4.0 license — useful for fine-tuning or distillation experiments but limited by narrow genre and small download footprint.
Provides raw newline-delimited JSON agent traces where assistant responses were generated by qwen/qwen3.7-max, captured with Teich; includes 47 JSONL files, an embedded tools schema snapshot, and conversion guidance for supervised fine‑tuning and distillation.
Large streaming-audio dataset for training and evaluating audio-LLMs and audio agents. About 2.28M clips grouped into multi-turn “streams” across six task subsets (ASR, speech translation, audio understanding, voice chat, proactive response, environment-aware); audio shipped as tar shards.
Benchmark for evaluating vision–language models on measurement-grounded inputs vs. RGB, emphasizing low-light, HDR, and visibility-sensitive evidence recovery. Contains 2,183 paired test examples with local image assets for controlled RAW↔RGB comparisons.