Why this matters Most ToM benchmarks focus on large models or narrow paradigms; evaluating compact base models needs short, scoreable items that work with continuation likelihoods rather than instruction-following. This dataset frames a broad set of child-development–inspired ToM constructs as short continuation prompts so low-parameter models can be benchmarked reliably.
What Sets It Apart
- Continuation-native format: items are written as unfinished scenes with four possible continuations and intended for base-model length-normalized log-likelihood scoring rather than instruction-response generation, which makes automatic, reproducible evaluation straightforward.
- Broad, balanced coverage: 2,000 examples across 40 Theory-of-Mind constructs (50 samples each), exactly balanced answer positions, and a 50/50 split between easy and medium difficulty to reduce shortcut incentives.
- Small-model focus and baselines: designed specifically for small models (pretrained continuation scoring) and shipped with 36 baseline results so users can place new models in context quickly.
- Pedagogical framing: constructs map to child grade bands (PreK–6) and diverse themes (false belief, deception, perspective taking, sarcasm, etc.), which helps analyze which developmental-style concepts models capture.
Who it's for and trade-offs
Great fit if you need a compact, machine-scoreable ToM probe for low-parameter or base models, want balanced multiple-choice items, and prefer synthetic, controllable scenarios. Look elsewhere if you need naturalistic human-authored narratives, free-text explanations, chain-of-thought style prompts, or large-scale bilingual corpora—the dataset is synthetic, English-only, and optimized for likelihood-based continuation scoring rather than generative evaluation.
Where it fits
Use this as a fast, reproducible unit test for ToM-like capabilities in small LMs, as a complement to larger bilingual or human-annotated ToM benchmarks when you need lightweight diagnostics or continuous integration checks.