Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

31 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Harvey LAB

Legal Agent Benchmark (LAB): An open-source benchmark for evaluating agents on real legal work.

Latest version License: MIT Legal practice areas Tasks Test suite

Harvey LAB is an open-source project aimed at benchmarking LLM agents' abilities to perform legal work in realistic environments.

LAB consists of two parts: a dataset of tasks containing agent instructions, documents, and rubrics as well as an execution harness for running and evaluating agents against those tasks.

LAB is an ongoing project and we expect to consistently add to and refine the task set and execution harness.

Read the announcement post: Introducing Harvey's Legal Agent Benchmark

Getting Started

Start with the full walkthrough in docs/tutorial.md — it takes one realistic M&A data-room assignment end to end: setup, task inspection, agent run, scoring, report review, and comparison dashboards.

Additional Documentation

Guide Description
Architecture Task model, harness, tools, adapters, reports, and sweeps
Evaluation Methodology All-pass rubric scoring and LLM judge behavior
Contributing Add tasks, model adapters, evaluation improvements, and docs

Citation

If you use Harvey LAB in your research, please cite it as:

@misc{harveylab2026,
  title   = {Harvey LAB: The Legal Agent Benchmark},
  author  = {{Harvey AI}},
  year    = {2026},
  version = {v1.0},
  url     = {https://github.com/harveyai/harvey-labs/tree/v1.0},
  note    = {Announcement: \url{https://www.harvey.ai/blog/introducing-harveys-legal-agent-benchmark}}
}

About

A benchmark built to evaluate and improve agent capabilities for supporting legal work.

Resources

Contributing

Stars

Watchers

Forks

Packages

Used by

Contributors

Languages