AIAny
Icon for item

Harvey LAB

Benchmarks LLM agents on realistic legal work by packaging lawyer-style assignments with client materials and expert, per-deliverable rubrics. Includes an execution harness to run, score, and compare agents across a large, evolving task set spanning multiple practice areas.

Introduction

Legal work is long-horizon, evidence-heavy, and assessed against partner-level scrutiny. Harvey LAB frames real legal assignments as agent tasks — each task pairs an instruction with client documents and an expert rubric — so teams can measure whether an agent’s deliverable would pass a partner or client review.

What Sets It Apart
  • Task+Harness combination: LAB is both a dataset of attorney-style tasks and an execution harness that runs agents end-to-end and collects outputs for evaluation, so evaluation is repeatable rather than ad-hoc.
  • All-pass, atomic rubrics: Deliverables are scored against expert-written, binary pass/fail criteria (facts, citations, deadlines, dollar amounts, severity ratings, formatting) tied to specific files. That structure supports LLM-based judging, per-criterion signals for training, and consistent comparisons across runs.
  • Realistic scope and scale: The initial release includes hundreds-to-thousands of tasks across dozens of practice areas with tens of thousands of rubric criteria, emphasizing long-horizon workflows (e.g., M&A data-room assignments) rather than toy benchmarks.
  • Open and iterative: Released as open-source so model providers, law firms, and researchers can reproduce results, contribute tasks, and refine evaluation standards over time.
Who It's For and Tradeoffs

Great fit if you want to understand where LLM agents can safely replace, augment, or require human-in-the-loop review for legal work — e.g., law firms, legaltech teams, and researchers building long-horizon agent skills. It helps quantify ROI and identify specific failure modes.

Look elsewhere if you need a plug-and-play production legal assistant today: LAB is a research and benchmarking resource (not a deployed SaaS product), requires setup (the harness and judge configuration), and assumes human oversight for any client-facing use. The benchmark also depends on judged criteria and evolving task sets, so results are only as meaningful as the rubric and the chosen judge configuration.

More Items

Hugging Face

Provides 1,080,814 images extracted from ~65,000 digitised British Library book volumes (c.1510–c.1900), split into four algorithmic image-type configs and packaged as parquet for image–text multimodal research and retrieval.

GitHub
AI Agent2025

Transforms Claude Code into a structured development platform by injecting behavioral instructions and orchestrating workflows via 30 slash commands. Provides 20 specialized agents and optional MCP server integrations for faster, token‑efficient research and agent-driven dev workflows.

GitHub
AI Coding2024

Runs an AI coding agent that edits, tests, and manages code across terminal, desktop, web, and GitHub repositories. Uses specialized agents and a curated model catalog (DeepSeek, MiMo, MiniMax, GPT-5.6) and offers free access supported by text ads.