Why this matters
Anthropic shared curated, aggregated cluster outputs produced for three independent research partners so external teams can analyze how people use Claude without accessing raw conversations. The dataset captures researcher-defined "facets" and Claude-generated cluster names/descriptions for roughly 250,000 sampled consumer conversations from a fixed April–May 2026 window, enabling replication and secondary analysis of human–AI collaboration patterns while preserving user privacy.
What Sets It Apart
- Partner-provided cluster outputs: Contains the exact cluster CSVs that Anthropic delivered to Stanford (SALT), Oxford (Human Information Processing Lab), and METR, reflecting each partner's facet definitions and aggregation choices rather than raw logs.
- Privacy-first aggregation: No raw conversations, user IDs, or org IDs are included; clusters were manually reviewed by Anthropic and underwent third-party reidentification attempts that did not succeed. Cross-facet columns and summary statistics enable intersectional analysis without exposing individuals.
- Research-focused schema: Rows are clusters (cluster_id, cluster_name, cluster_description, level, num_records, ratios, mean/sum where applicable) and are cross-tabulated against other facets. Files include stanford, oxford, metr, and metr_addendum with slightly different aggregation thresholds and facet sets.
Who It's For and Trade-offs
Great fit if you are an academic or policy researcher studying human–AI collaboration, usability, or productivity claims and you need realistic, model-generated cluster summaries rather than raw text. The dataset supports reproducible cross-facet analysis and method development for cluster-based evaluation.
Look elsewhere if you need representative population-level usage statistics, raw conversation text, or enterprise/API data: the sample covers consumer Free/Pro/Max usage only, is a one-time snapshot (April–May 2026), omitted one problematic Oxford facet, and cluster labels are Claude-generated interpretations (not validated ground truth). Treat cluster names as interpretive summaries and read the accompanying guidance before drawing hard conclusions.