The Citation Record, Edition 2026.Q4 (forthcoming)
An independent quarterly benchmark measuring how accurately legal AI systems cite the law.
What this measures
Legal AI systems produce citations. Some of those citations are to cases that do not exist. Others are to real cases with the wrong reporter or page, real cases that say something other than what the system claims, real cases that have been overruled, or real cases from a jurisdiction that does not bind the court in question.
These are six different failures with six different professional consequences, and the standard verification a lawyer performs — confirming the case exists — catches only the first. Every vendor claim we surveyed reports existence, or reports nothing measurable at all. The Citation Record scores all six separately, against a ground-truth corpus owned by none of the systems under test, and publishes the false-positive rate alongside the catch rate.
How it works
Practising attorneys run a standardized query set inside subscriptions they hold and pay for, and score the results. Citation existence and accuracy resolve mechanically against CourtListener. Misattributed holdings, overruled authority, and jurisdictional weight require attorney judgment and receive it.
No AI system scores any item. The methodology, prompt set, and scoring rubric are published in full; the test items are held out and rotated. Each edition is free, carries a permanent identifier, and states the measurement window it describes.
Disclosures
No vendor pays anything
No measured vendor has paid, or will pay, any amount in any form: no participation fees, sponsorship, licensing, consulting, equity, referral commissions, or complimentary access beyond publicly available free tiers. This is a permanent commitment, not a policy for one edition.
The publisher builds one of the systems under test
Citation Firewall is a verification tool built by the publisher. It is scored under the identical rubric, on the same held-out items, by the same panel, with the publisher recused from scoring it. Its results are published whatever they are. Readers should discount them accordingly and are encouraged to replicate them first.
No non-disclosure agreements
The publisher has signed no NDA with any measured vendor and will not. Where a conflict cannot be disclosed, the affected system is withdrawn from testing and the withdrawal is stated.
Vendors see results before publication, and cannot change them
Each measured vendor receives its results 21 days before publication. Responses are printed unedited. Corrections are adopted where warranted; disputes are left standing on the record.
Status
- First edition
- 2026.Q4
- Ground truth
- CourtListener bulk data, generation stated per edition
- Licence
- Creative Commons Attribution 4.0
- Tooling
- github.com/CitationRecord
Taking part
The panel is being assembled. Attorneys who hold paid subscriptions to legal research or drafting systems and do litigation work are welcome to write, as are law librarians willing to review the methodology before it is fixed. The commitment is roughly ninety minutes per quarter. Contributors are credited by name; none are paid.