AIAny
Icon for item

Open-PerfectBlend

A mixed instruction dataset for SFT and RLHF research that combines chat, math, code and instruction-following samples from multiple public datasets under an Apache-2.0-compatible license; intended for instruction tuning and evaluation.

Introduction

Why this matters Open-PerfectBlend packs diverse instruction-following data into a single public dataset to support supervised fine-tuning and RLHF-style experiments. By mixing chat, math word problems, code-oriented examples and instruction-response pairs from several existing collections, it gives researchers a broad training mixture without relying on a single source.

What Sets It Apart
  • Multi-source mixture: combines large public datasets (e.g., MetaMathQA, UltraInteract_sft, ultrachat, orca-math, ultrafeedback, evol-codealpaca, AutoIF, and ShareGPT-derived preference data) to produce ~1.42M train examples after deduplication. This yields wider task coverage (conversational, mathematical reasoning, coding, instruction following).
  • Licensing and format: distributed under Apache-2.0, stored in parquet and intended for use with Hugging Face Datasets and common processing stacks.
  • Practical notes: the published split contains precise counts and some known data-quality issues (a discussion noted ~62k rows with no assistant response in the train split and a deduplication step that removed ~88.1k samples). A separate decontaminated variant exists that removes a small number of contaminated documents.
Who It's For and Tradeoffs

Great fit if you need a single, license-clear mixture for instruction tuning or RLHF research and want varied example types (chat, math, code). Look elsewhere if you require guaranteed per-source provenance at the example level, fully cleaned assistant-only rows out of the box, or a dataset that includes the paper's withheld "harmful intent" category—the original reproducer documented differences and remaining prompt-only rows. Consider running your own decontamination and assistant-response filters before assistant-only SFT.

Where It Fits

Use it as a broad pre-fine-tuning mixture or as part of a training corpus ensemble for LLM instruction alignment. For very strict benchmarking or production deployments, augment with focused, high-quality task-specific datasets and explicit provenance/cleaning steps.

Information

Categories

More Items

Hugging Face

Provides a machine-readable catalog of 117 AI/AX safety and deployment-readiness diagnostic criteria for assessing model intrinsic and serving/infrastructure risks. Includes MODEL-SCAN and AX-SCAN axes, bilingual source fields, per-item evidence guidance, severity/assurance metadata, and a CC BY-NC 4.0 release-candidate.

Hugging Face

Evaluates schema-guided structured extraction from documents: given a document and a JSON schema, systems must return a schema-valid JSON with page-and-box grounding. Covers 370 documents (4,869 pages) across 8 business domains and 67 document types; scores value accuracy, word/page grounding, and long-list completeness.

Hugging Face

Provides 10 million synchronized egocentric experience episodes with structured 3D/4D multimodal annotations — 2.88B RGB frames, 720M depth frames, 576M pose/mocap frames and ~1PB total. Designed for embodied AI, robotics, and multimodal pretraining; research-only, gated access.