← All projects

BENCHMARK·LEGAL AI·2026

benchmarks.law

An independent legal-AI benchmark: tools tested without vendor consent, results never withdrawn, no LLM judges. The method is published; first results are pending.

benchmarks.law home page: the masthead above the line 'A guide rates the kitchen, invited or not'

/ Overview

Most legal-AI evaluations are run with the vendor’s permission, and some let a vendor withdraw a result it does not like. A restaurant guide does not work that way, and neither should a benchmark that lawyers rely on when they buy.

benchmarks.law is the standard I would want as a buyer, written down before any result exists, so the rules cannot be bent to fit the numbers.

/ The three rules

  • No invitation required. Tools are tested whether or not the vendor agrees to be tested.
  • No quiet exits. A result stands once it is run. A vendor may reply, and the reply is printed beside the finding.
  • No machine judges. A language model never decides whether another model got the law right.

/ The first project

The first project is about hallucination and error-spotting: take a real brief or an executed agreement, introduce one known error, give the tool a single fixed review instruction, and check whether it returns a corrected document. Because the error is planted, the right answer is known in advance and the check is binary.

/ Where it stands

The standard and the method are published. Results are not yet: a first round of runs exists and is being re-run from a frozen dataset so that every published number can be reproduced. Until those are on the site, this is a method and a commitment, not a leaderboard. Contributors are welcome.

/ Lineage

It follows two 2025 reports I co-authored for Legal Benchmarks, which compared AI tools and in-house lawyers on real legal tasks. benchmarks.law is my own later project, built on stricter rules.