Legal work is long-horizon, evidence-heavy, and assessed against partner-level scrutiny. Harvey LAB frames real legal assignments as agent tasks — each task pairs an instruction with client documents and an expert rubric — so teams can measure whether an agent’s deliverable would pass a partner or client review.
What Sets It Apart
- Task+Harness combination: LAB is both a dataset of attorney-style tasks and an execution harness that runs agents end-to-end and collects outputs for evaluation, so evaluation is repeatable rather than ad-hoc.
- All-pass, atomic rubrics: Deliverables are scored against expert-written, binary pass/fail criteria (facts, citations, deadlines, dollar amounts, severity ratings, formatting) tied to specific files. That structure supports LLM-based judging, per-criterion signals for training, and consistent comparisons across runs.
- Realistic scope and scale: The initial release includes hundreds-to-thousands of tasks across dozens of practice areas with tens of thousands of rubric criteria, emphasizing long-horizon workflows (e.g., M&A data-room assignments) rather than toy benchmarks.
- Open and iterative: Released as open-source so model providers, law firms, and researchers can reproduce results, contribute tasks, and refine evaluation standards over time.
Who It's For and Tradeoffs
Great fit if you want to understand where LLM agents can safely replace, augment, or require human-in-the-loop review for legal work — e.g., law firms, legaltech teams, and researchers building long-horizon agent skills. It helps quantify ROI and identify specific failure modes.
Look elsewhere if you need a plug-and-play production legal assistant today: LAB is a research and benchmarking resource (not a deployed SaaS product), requires setup (the harness and judge configuration), and assumes human oversight for any client-facing use. The benchmark also depends on judged criteria and evolving task sets, so results are only as meaningful as the rubric and the chosen judge configuration.