AIAny
Icon for item

GLM-5.2 Conversation

50,000 distilled conversational traces (≈120M tokens) generated from GLM-5.2 for high-reasoning text generation and QA, covering STEM, programming, creative and support dialogues; Apache-2.0 licensed.

Introduction

Most broad LLM corpora mix many casual dialogues without isolating reasoning-heavy interactions. This dataset emphasizes high-reasoning conversational traces distilled from GLM-5.2, making it a compact resource if your goal is to teach or evaluate step-by-step reasoning, chain-of-thought style answers, and applied STEM/programming dialogue behaviors.

What Sets It Apart
  • Distillation-focused: 50,000 traces condensed from GLM-5.2 outputs, totaling about 120M tokens — intended to preserve reasoning quality while reducing dataset size.
  • Reasoning & domain diversity: examples span algebra, calculus, quantum concepts, astronomy, data science, biology and chemistry, plus programming tasks and code review prompts — so model exposures combine conceptual STEM reasoning with practical coding dialog.
  • Training-oriented format: provided as JSON and labeled for text-generation and QA-style supervision, suitable for SFT/distillation workflows without heavy preprocessing.
  • Permissive license and reuse: Apache-2.0 license allows broad use, including commercial fine-tuning and distillation experiments.
Who It's For and Tradeoffs

Great fit if you need a mid-sized, reasoning-focused conversational corpus for supervised fine-tuning, distillation, or evaluation of chain-of-thought behavior in LLMs. Look elsewhere if you require massive raw human-chat logs, multimodal data, or heavily curated human-annotated labels; the traces were model-generated (prompts from GPT-OSS-120b answered by GLM-5.2) so they reflect model tendencies rather than pure human reasoning. Expect faster iteration and lower compute needs than training on multi-billion-token corpora, at the cost of inheriting generator biases and occasional synthetic artifacts.

More Items

Hugging Face

Provides a dual-channel, channel-separated sample (8.9 hours) and access path to a 1,000‑hour English conversational corpus for commercial and research use. Delivers 48 kHz per-speaker audio, word-level machine transcripts, and per-speaker metadata designed for full‑duplex/turn-taking and ASR/ TTS research.

Hugging Face

Provides 997 chain-of-thought cybersecurity reasoning records distilled from the Kimi K3 model, each with an explicit <think> trace and a technical resolution or structured tool invocation. Includes verified tool-call objects, diffs, cross-domain coverage, and token-level metadata for fine-tuning and evaluating reasoning models.

Hugging Face

Provides an unattended text-to-video-and-audio streaming toolkit built around FastH3 (a 4-step distillation of MiniMax-H3): generation/retime/HTTP push scripts, a 221-scene prompt library, checkpoint conversion and ComfyUI workflows to run a continuous local stream.