AIAny
Icon for item

IFM/Pretrain-Behaviors

Behavior-focused text corpus for LM pretraining, organized into seven Parquet-backed subsets (reasoning, planning, data-science, games, general, format-rewrites, other). Supports streaming, custom sampling, and large-scale dataset pipelines for research and model training.

Introduction

Behavioral text patterns strongly influence how language models reason, plan, and simulate decision-making. This dataset packages behavior-focused text into shardable Parquet subsets so you can stream, weight, and combine domain slices for pretraining or controlled evaluations.

What Sets It Apart
  • Focused domains: explicit subsets for reasoning, planning, data science, games, format rewrites and general/other content let you upweight behavioral genres during pretraining, rather than relying on undifferentiated web crawls.
  • Engineering-friendly format: released as Parquet shards with predictable shard prefixes and streaming support, so pipelines using datasets, polars, dask or mIcroissant-style iterators can integrate it with minimal conversion overhead.
  • Modular provenance guidance: subset-level metadata may include source-specific cleaning, deduplication, quality scores or synthetic generation flags, enabling targeted selection and risk assessment prior to model training.
Who It's For and Trade-offs

Great fit if you need to emphasize human-like reasoning/planning or behavioral scenarios in a pretraining mix, want shardable Parquet data for large-scale streaming pipelines, or need separable subsets for ablation studies. Look elsewhere if you require fully documented per-example provenance, multimodal signals (audio/video), or a curated benchmark with fixed evaluation splits — this release targets pretraining-scale text corpora rather than evaluation-only benchmarks.

Information

Categories

More Items

Hugging Face

Provides multiple Parquet-backed subsets of code problem-solving data (direct answers, chain-of-thought reasoning, and task synthesis) that are streamable and prepared for language-model training and evaluation.

Hugging Face

Provides 30,969 action-conditioned video episodes, each with source MP4, per-frame keyboard control logs, captions, and a COLMAP sparse pose model — intended for research on action-conditioned video prediction, controllable world models, and representation learning.

Hugging Face

Provides a dual-channel, channel-separated sample (8.9 hours) and access path to a 1,000‑hour English conversational corpus for commercial and research use. Delivers 48 kHz per-speaker audio, word-level machine transcripts, and per-speaker metadata designed for full‑duplex/turn-taking and ASR/ TTS research.