AIAny
Icon for item

Anthropic Insights Pilot: Partner Cluster Data

Provides aggregated, privacy-preserving cluster outputs from three external research teams' analyses of ~250k Claude/Claude Code conversations; includes per-team CSVs (Stanford, Oxford, METR) for studying human–AI collaboration and model behavior without raw conversations.

Introduction

Why this matters

Anthropic shared curated, aggregated cluster outputs produced for three independent research partners so external teams can analyze how people use Claude without accessing raw conversations. The dataset captures researcher-defined "facets" and Claude-generated cluster names/descriptions for roughly 250,000 sampled consumer conversations from a fixed April–May 2026 window, enabling replication and secondary analysis of human–AI collaboration patterns while preserving user privacy.

What Sets It Apart
  • Partner-provided cluster outputs: Contains the exact cluster CSVs that Anthropic delivered to Stanford (SALT), Oxford (Human Information Processing Lab), and METR, reflecting each partner's facet definitions and aggregation choices rather than raw logs.
  • Privacy-first aggregation: No raw conversations, user IDs, or org IDs are included; clusters were manually reviewed by Anthropic and underwent third-party reidentification attempts that did not succeed. Cross-facet columns and summary statistics enable intersectional analysis without exposing individuals.
  • Research-focused schema: Rows are clusters (cluster_id, cluster_name, cluster_description, level, num_records, ratios, mean/sum where applicable) and are cross-tabulated against other facets. Files include stanford, oxford, metr, and metr_addendum with slightly different aggregation thresholds and facet sets.
Who It's For and Trade-offs

Great fit if you are an academic or policy researcher studying human–AI collaboration, usability, or productivity claims and you need realistic, model-generated cluster summaries rather than raw text. The dataset supports reproducible cross-facet analysis and method development for cluster-based evaluation.

Look elsewhere if you need representative population-level usage statistics, raw conversation text, or enterprise/API data: the sample covers consumer Free/Pro/Max usage only, is a one-time snapshot (April–May 2026), omitted one problematic Oxford facet, and cluster labels are Claude-generated interpretations (not validated ground truth). Treat cluster names as interpretive summaries and read the accompanying guidance before drawing hard conclusions.

Information

  • Websitehuggingface.co
  • OrganizationsAnthropic, Social and Language Technologies (SALT) Lab, Stanford University, Human Information Processing Lab, University of Oxford, METR
  • AuthorsKunal Handa, Miranda Zhang, Gabriel Nicholas, Miles McCain, Ryan Heller, Saffron Huang, Thomas Millar, Suzanne Wang, Shan Carter, Mo Julapalli
  • Published date2026/08/26

Categories

More Items

Evaluates AI agents' ability to complete end-to-end scientific workflows by releasing and assessing 97 tasks from a 300-task FrontierChallenge suite across chemistry, materials, life science, and electrochemistry. Finds that top agent configurations achieved only a 20.6% pass rate despite high partial scores, revealing a gap between partial progress/confident completion claims and actual complete scientific deliverables.

Hugging Face

Provides 22.7 hours of read Amharic speech (7,405 clips, 320 speakers) for ASR, collected via a crowdsourced Telegram bot and peer-validated; speaker- and prompt-disjoint train/validation/test splits, 16 kHz audio under CC BY 4.0.

Hugging Face

Evaluates whether tool-using LLM agents reliably complete stateful business workflows via 507 executable agent–tool–user tasks across retail, travel, auto insurance, neobank, and IT/HR consulting. Provides browsable Parquet tables for tasks, scenarios, and agent instructions; v1.0 is intended for evaluation-only.