AIAny
Icon for item

Anthropic/BioMysteryBench-full

A collection of biology-focused 'mystery' tasks for benchmarking model performance on biomedical reasoning, evidence synthesis, and problem solving; curated by Anthropic and hosted on Hugging Face, designed for granular evaluation of scientific decision-making.

Introduction

Many benchmarks measure final-answer accuracy but miss whether a model reasons and handles scientific evidence correctly. This dataset assembles multi-step, biology-focused "mystery" tasks to stress-test models' domain reasoning, evidence reconciliation, and practical research-style judgments rather than just surface-level recall.

What Sets It Apart
  • Tasks emphasize multi-step reasoning and evidence handling, so evaluations reveal whether a model reaches conclusions with domain-appropriate justification rather than guessing.
  • Curated by domain-aware designers ( Anthropic ) and provided as a reusable Hugging Face dataset, so it fits into common evaluation pipelines and can be combined with rubrics or automatic judges.
  • Broad artifact support (textual prompts plus supporting artifacts) means prompts often require interpreting auxiliary data, which better approximates real life-science workflows.
Who It's For and Trade-offs
  • Great fit if you evaluate LLMs or multimodal models intended for biomedical literature synthesis, hypothesis ranking, or decision-support in life-science workflows. It helps surface reasoning faults and evidence-misuse risks.
  • Look elsewhere if you only need large-scale classification or simple QA datasets: these tasks are smaller and intentionally harder per example, focusing on quality of reasoning and interpretability over raw throughput or scale.
Where It Fits
  • Complements numeric benchmarks (accuracy-focused) by adding a layer of scientific-validity assessment. Use it alongside automated metrics and human-expert rubrics when assessing model readiness for research-assist roles.

Information

  • Websitehuggingface.co
  • OrganizationsAnthropic, Hugging Face
  • Published date2026/04/29

Categories

More Items

Evaluates multimodal context learning across grounding, new information application, and knowledge acquisition using a 3,443-instance benchmark spanning science, finance, long documents, spatial reasoning, and web VQA; finds current multimodal models perform poorly (best score 0.2847) and analyzes failure modes.

Hugging Face

Provides 2,000 hours of synchronized, high‑fidelity robot‑free bimanual manipulation demonstrations with multi‑view video, calibrated end‑effector trajectories, gripper states, and language annotations. Curated from a 20,000+ hour corpus; features 6 camera views, ~3 mm pose accuracy, <40 µs cross‑sensor sync, and LeRobot v3‑style Parquet+MP4 export under CC BY 4.0.

Hugging Face

A small image-folder dataset for multimodal/vision model safety benchmarking, containing under 1,000 curated images with annotations to exercise safety-related model behaviours; licensed CC BY 4.0 and hosted on HuggingFace.