AIAny
Icon for item

BEV Decision Mix

Provides 150,000 source‑grounded decision examples for training models that pick options, judge yes/no propositions, or assign ordered scores. Each row pairs a 'state' with JEV-style typed questions (CHOICE/NOUL/SCORE); multiple configs and train/test splits are included, license mixed/unknown.

Introduction

Why this matters

Most NLP datasets focus on free-form generation or single-label classification. This mixture supplies typed, bounded decision targets — choice keys, boolean probabilities, and ordinal scores — so models can be trained and evaluated to return calibrated, structured answers instead of long text. That matters if you want a model to make discrete routing, tool-selection, or scoring decisions with explicit probability distributions.

What Sets It Apart
  • Broad, source-grounded mixture: 150,000 English rows drawn from many upstream datasets, public records, and text-game environments, with 140,007 distinct normalized states and repeated states allowed new decision questions.
  • JEV-style typed questions: CHOICE (select a key from a bounded map), NOUL (P(yes) for boolean claims), and SCORE (ordinal level index and probabilities) are baked into each row’s JSON, enabling single-pass, typed outputs.
  • Multiple focused configs: default (150k), plus hard_50k, numeric_temporal, skills, and counterfactual_15k that emphasize multi-hop, temporal/numeric reasoning, calibration, and paired counterfactuals.
  • Domain diversity: sentiment, retail, spatial/logical reasoning, routing, finance, scientific judgments, game-state decisions, and more — useful for training general decision logic across tasks.
  • Label provenance and caveats: many labels are upstream annotations, executed programs, or deterministic rules; some question wordings were model-rewritten and some upstream sources are synthetic. No single blanket license is asserted for the mixed-source release.
Who it’s for — and trade-offs

Great fit if you want to train or evaluate bounded decision models that must output typed, calibrated distributions (e.g., tool routing, intent classification, ordinal scoring, or safety gating). The dataset’s structured questions_json is convenient for models that return probabilities per option rather than free text.

Look elsewhere if you need a single-source, consistently licensed corpus for redistribution or an independent benchmark free of upstream overlap; the test split is a held-out portion of the mixture and may include upstream benchmark families. Also, SCORE targets are ordinal indices (not calibrated numeric labels), and some rows include model-rewritten text or synthetic upstream content, so verify provenance and license before commercial redistribution.

Where it fits

Use this when your model must make discrete, auditable decisions with explicit probabilities (e.g., a decision-focused System One API). It complements generative or retrieval datasets by teaching models to map grounded states to constrained outputs rather than generating long answers.

Information

Categories

More Items

Hugging Face

Provides 7,366 recorded agent trajectories from H Company’s Holo4 benchmark runs, with step-level reasoning, actions, tool results, token usage and screenshots for replay and analysis. Bundled as JSON and image files for per-trajectory inspection and automated replay; released under Apache 2.0.

Hugging Face

Provides 23,625 semi-structured smart-contract audit findings (title, description, PoC, recommendation, normalized severity) for defensive-security research; requires cleaning, deduplication, and PoC filtering before model training.

Hugging Face

Provides over 1.1M hours of high-bandwidth, multichannel multilingual speech with segment- and word-level timestamps, English translations, and per-file metadata for ASR, TTS and audio-representation research. Preserves original 48kHz multichannel OPUS audio and is released under CC BY 3.0.