AIAny
Icon for item

Tiny Theory of Mind

Provides 2,000 synthetic multiple-choice items designed for continuation log-likelihood scoring to evaluate small language models' Theory of Mind (social-cognitive) abilities; 40 constructs, balanced answer positions, and easy/medium difficulty.

Introduction

Why this matters Most ToM benchmarks focus on large models or narrow paradigms; evaluating compact base models needs short, scoreable items that work with continuation likelihoods rather than instruction-following. This dataset frames a broad set of child-development–inspired ToM constructs as short continuation prompts so low-parameter models can be benchmarked reliably.

What Sets It Apart
  • Continuation-native format: items are written as unfinished scenes with four possible continuations and intended for base-model length-normalized log-likelihood scoring rather than instruction-response generation, which makes automatic, reproducible evaluation straightforward.
  • Broad, balanced coverage: 2,000 examples across 40 Theory-of-Mind constructs (50 samples each), exactly balanced answer positions, and a 50/50 split between easy and medium difficulty to reduce shortcut incentives.
  • Small-model focus and baselines: designed specifically for small models (pretrained continuation scoring) and shipped with 36 baseline results so users can place new models in context quickly.
  • Pedagogical framing: constructs map to child grade bands (PreK–6) and diverse themes (false belief, deception, perspective taking, sarcasm, etc.), which helps analyze which developmental-style concepts models capture.
Who it's for and trade-offs

Great fit if you need a compact, machine-scoreable ToM probe for low-parameter or base models, want balanced multiple-choice items, and prefer synthetic, controllable scenarios. Look elsewhere if you need naturalistic human-authored narratives, free-text explanations, chain-of-thought style prompts, or large-scale bilingual corpora—the dataset is synthetic, English-only, and optimized for likelihood-based continuation scoring rather than generative evaluation.

Where it fits

Use this as a fast, reproducible unit test for ToM-like capabilities in small LMs, as a complement to larger bilingual or human-annotated ToM benchmarks when you need lightweight diagnostics or continuous integration checks.

Information

  • Websitehuggingface.co
  • OrganizationsAxiomicLabs
  • AuthorsMmorgan-ML, Datdanboi, Amanda Long
  • Published date2026/09/23

Categories

More Items

Hugging Face

Provides 5.5K+ self-contained data-analysis RL tasks: each row bundles a real tabular dataset, a question, and a deterministically-gradable gold answer. Verified from jupyter-agent notebooks; splits for training, held-out testing, and quick eval; intended for prompting, fine-tuning, and agent RL.

Hugging Face

Provides mixed-domain, verifiable RL training environments for LLM agents (code, cyber, knowledge work, web dev, music) as Parquet datasets, with domain-specific verifiers, Docker artifacts and links to training code for reproducible agentic RL experiments.

Provides WROP: a 1.5M-sample synthetic video corpus and a 300-question exam for training and evaluating object permanence and solidity in video world models. Includes 150 Blender task generators, a human Elo benchmark across 14 models, and a fine-tuned 16B continuation model (PWM-WROP).