AIAny
Icon for item

Anthropic Insights Pilot: Partner Cluster Data

Provides aggregated, privacy-preserving cluster outputs from three external research teams' analyses of ~250k Claude/Claude Code conversations; includes per-team CSVs (Stanford, Oxford, METR) for studying human–AI collaboration and model behavior without raw conversations.

Introduction

Why this matters

Anthropic shared curated, aggregated cluster outputs produced for three independent research partners so external teams can analyze how people use Claude without accessing raw conversations. The dataset captures researcher-defined "facets" and Claude-generated cluster names/descriptions for roughly 250,000 sampled consumer conversations from a fixed April–May 2026 window, enabling replication and secondary analysis of human–AI collaboration patterns while preserving user privacy.

What Sets It Apart
  • Partner-provided cluster outputs: Contains the exact cluster CSVs that Anthropic delivered to Stanford (SALT), Oxford (Human Information Processing Lab), and METR, reflecting each partner's facet definitions and aggregation choices rather than raw logs.
  • Privacy-first aggregation: No raw conversations, user IDs, or org IDs are included; clusters were manually reviewed by Anthropic and underwent third-party reidentification attempts that did not succeed. Cross-facet columns and summary statistics enable intersectional analysis without exposing individuals.
  • Research-focused schema: Rows are clusters (cluster_id, cluster_name, cluster_description, level, num_records, ratios, mean/sum where applicable) and are cross-tabulated against other facets. Files include stanford, oxford, metr, and metr_addendum with slightly different aggregation thresholds and facet sets.
Who It's For and Trade-offs

Great fit if you are an academic or policy researcher studying human–AI collaboration, usability, or productivity claims and you need realistic, model-generated cluster summaries rather than raw text. The dataset supports reproducible cross-facet analysis and method development for cluster-based evaluation.

Look elsewhere if you need representative population-level usage statistics, raw conversation text, or enterprise/API data: the sample covers consumer Free/Pro/Max usage only, is a one-time snapshot (April–May 2026), omitted one problematic Oxford facet, and cluster labels are Claude-generated interpretations (not validated ground truth). Treat cluster names as interpretive summaries and read the accompanying guidance before drawing hard conclusions.

Information

  • Websitehuggingface.co
  • OrganizationsAnthropic, Social and Language Technologies (SALT) Lab, Stanford University, Human Information Processing Lab, University of Oxford, METR
  • AuthorsKunal Handa, Miranda Zhang, Gabriel Nicholas, Miles McCain, Ryan Heller, Saffron Huang, Thomas Millar, Suzanne Wang, Shan Carter, Mo Julapalli …
  • Published date2026/08/26

Categories

More Items

Hugging Face

Evaluation dataset for comparing eight text-to-image models using 8,000 generated images with source prompts and per-image scores for aesthetic quality, emotional resonance, and content integrity. Includes model labels, shared prompts, GPT-5.6 Sol automated scores, embedded images in Parquet, and an Apache-2.0 license.

Hugging Face

Provides a 1 trillion-token multimodal interleaved dataset (HTML subset updated as data_v1_1 with 742B HTML tokens) and 3.4B images drawn from HTML/PDF/ArXiv sources for multimodal pretraining; released under CC-BY-4.0 with safety and deduplication guidance.

Hugging Face

Provides 369 Harbor sandbox tasks ported from OpenAI's openai/math: each task is a Lean theorem with missing `sorry` proofs that an agent must complete, graded by a strict Comparator exact-match verifier. Includes task definitions, generator, and manifest for RL/code-agent evaluation.