OpenCase Logo

OpenCase Benchmarks

How often AI cites the law correctly

We asked leading AI models and OpenCase 120 hard questions about American law, then looked up every case, statute and rule they cited. The questions, citations and verdicts are all public.

Legal Citation Benchmark / All setupsTRY OPENCASE
  1. OpenCase Deep99.7%
  2. Claude Opus 5.5 (no tools)98.9%
  3. Muse Spark 1.3 (web search)98.8%
  4. Claude Opus 5.5 (web search)98.3%
  5. GPT 5.6 Sol (web search)96.9%
  6. Claude Haiku 5.5 (web search)96.7%
  7. Claude Haiku 5.5 (no tools)96.5%
  8. OpenCase Fast96.5%
  9. GPT 6 Luna (web search)93.8%
  10. Gemini 3.8 Flash (web search)93.7%
  11. GPT 5.6 Sol (no tools)91.0%
  12. GPT 6 Luna (no tools)90.0%
  13. Gemini 3.8 Flash (no tools)83.3%
  14. Muse Spark 1.3 (no tools)74.6%

Share of citations that point to a real case, statute or rule with the right page, year and court. Higher is better. 120 questions.

Dataset

120 questions an attorney might research, each about a real decision from a named American court. 41 follow decisions from 2025 and 2026, which a model may not have seen. 43 ask about narrow points of law. 36 ask about older law, from decisions between 1881 and 1970, which is rarely cited today. We wrote them in batches of 20 and made each batch hard enough that at least one setup cited an authority that does not exist.

Evaluation methodology

Each model answered every question twice: once with no tools, and once with its maker's own web search, at its maker's default reasoning level. OpenCase answered in its Fast and Deep modes. A question a setup left unanswered is left out of its rate. A checker with web access looked up every citation in court websites, official reporters, the Caselaw Access Project and other open sources. A second pass re-checked every citation marked hallucinated.

What counts as correct

A correct citation is neither of these. Hallucinated: no case, statute or rule matching the citation exists. Miscited: the authority is real, but the page, year or court is wrong. The few citations the checker could settle neither way are left out.

Testing dates

Run in September and October 2026 with gpt-5.6-sol, gpt-6-luna, claude-opus-5-5, claude-haiku-5-5, gemini-3.8-flash and muse-spark-1.3.

#SETUPCORRECT
1OpenCase Deep99.7%
2Claude Opus 5.5 (no tools)98.9%
3Muse Spark 1.3 (web search)98.8%
4Claude Opus 5.5 (web search)98.3%
5GPT 5.6 Sol (web search)96.9%
6Claude Haiku 5.5 (web search)96.7%
7Claude Haiku 5.5 (no tools)96.5%
7OpenCase Fast96.5%
9GPT 6 Luna (web search)93.8%
10Gemini 3.8 Flash (web search)93.7%
11GPT 5.6 Sol (no tools)91.0%
12GPT 6 Luna (no tools)90.0%
13Gemini 3.8 Flash (no tools)83.3%
14Muse Spark 1.3 (no tools)74.6%
OpenCase designed this benchmark and is one of the setups tested. Every verdict is in the CSV.READ THE REPORTDOWNLOAD CSV