BENCHMARK·LEGAL AI·2026
benchmarks.law
An independent legal-AI benchmark: tools tested without vendor consent, results never withdrawn, no LLM judges. The method is published; first results are pending.

/ Overview
Most legal-AI evaluations are run with the vendor’s permission, and some let a vendor withdraw a result it does not like. A restaurant guide does not work that way, and neither should a benchmark that lawyers rely on when they buy.
benchmarks.law is the standard I would want as a buyer, written down before any result exists, so the rules cannot be bent to fit the numbers.
/ The three rules
- No invitation required. Tools are tested whether or not the vendor agrees to be tested.
- No quiet exits. A result stands once it is run. A vendor may reply, and the reply is printed beside the finding.
- No machine judges. A language model never decides whether another model got the law right.
/ The first project
The first project is about hallucination and error-spotting: take a real brief or an executed agreement, introduce one known error, give the tool a single fixed review instruction, and check whether it returns a corrected document. Because the error is planted, the right answer is known in advance and the check is binary.
/ Where it stands
The standard and the method are published. Results are not yet: a first round of runs exists and is being re-run from a frozen dataset so that every published number can be reproduced. Until those are on the site, this is a method and a commitment, not a leaderboard. Contributors are welcome.
/ Lineage
It follows two 2025 reports I co-authored for Legal Benchmarks, which compared AI tools and in-house lawyers on real legal tasks. benchmarks.law is my own later project, built on stricter rules.