The dataset provides long, teacher-produced assistant traces from a large reasoning-capable model (qwen3.8-max-preview), making it useful for experiments in supervised fine-tuning and off-policy distillation that aim to capture chain-of-thought style reasoning. Its mixture is heavily weighted toward math and code tasks and preserves raw model outputs, including visible <think> blocks, but comes with nontrivial provenance and license constraints inherited from upstream sources and Alibaba Cloud Model Studio.
Qwen3.8-Max Distillation 50K
A curated collection of 49,772 teacher-generated chat traces from qwen3.8-max-preview for supervised fine-tuning and off-policy distillation. Preserves visible chain-of-thought blocks, emphasizes math/code/reasoning mixes, and includes provenance and licensing cautions tied to Alibaba Cloud Model Studio.
Introduction
Information
- Websitehuggingface.co
- Organizationsr0b0tlab, Alibaba Cloud, Qwen (Alibaba)
- Published date2026/07/22
Categories
More Items
Provides 30,969 action-conditioned video episodes, each with source MP4, per-frame keyboard control logs, captions, and a COLMAP sparse pose model — intended for research on action-conditioned video prediction, controllable world models, and representation learning.
Provides a dual-channel, channel-separated sample (8.9 hours) and access path to a 1,000‑hour English conversational corpus for commercial and research use. Delivers 48 kHz per-speaker audio, word-level machine transcripts, and per-speaker metadata designed for full‑duplex/turn-taking and ASR/ TTS research.
Provides 997 chain-of-thought cybersecurity reasoning records distilled from the Kimi K3 model, each with an explicit <think> trace and a technical resolution or structured tool invocation. Includes verified tool-call objects, diffs, cross-domain coverage, and token-level metadata for fine-tuning and evaluating reasoning models.