OpenCase Logo

LEGAL CITATION BENCHMARK · OCTOBER 2026

2026 Hallucinations in Legal AI Report

How often does AI cite fake case law? We asked six AI models from OpenAI, Anthropic, Google and Meta 120 hard legal research questions and checked every citation they gave. Here is how often each one hallucinated a case or got a real one wrong, with and without web search.

Without tools · higher is better

  1. OpenCaseDeep
    99.7%
  2. Claude Opus 5.5no tools
    98.9%
  3. Claude Haiku 5.5no tools
    96.5%
  4. OpenCaseFast
    96.5%
  5. GPT 5.6 Solno tools
    91.0%
  6. GPT 6 Lunano tools
    90.0%
  7. Gemini 3.8 Flashno tools
    83.3%
  8. Muse Spark 1.3no tools
    74.6%

With web search · higher is better

  1. OpenCaseDeep
    99.7%
  2. Muse Spark 1.3web search
    98.8%
  3. Claude Opus 5.5web search
    98.3%
  4. GPT 5.6 Solweb search
    96.9%
  5. Claude Haiku 5.5web search
    96.7%
  6. OpenCaseFast
    96.5%
  7. GPT 6 Lunaweb search
    93.8%
  8. Gemini 3.8 Flashweb search
    93.7%

Key findings

  • Without tools, Muse Spark 1.3 got 25.4% of its citations wrong and Gemini 3.8 Flash 16.7%. Claude Opus 5.5 got 1.1% wrong, the fewest of any model.
  • Without tools, Muse invented at least one authority in 45% of its answers, and Gemini in 41%.
  • Web search helped the two weakest models most. Muse fell to 1.2% wrong and Gemini to 6.3%. Claude Opus moved from 1.1% to 1.7%.
  • OpenCase Deep was the most accurate setup we tested: 99.7% of its citations were correct. Neither OpenCase mode hallucinated an authority. Every OpenCase error was a real authority with a wrong detail.
#SetupCorrectWrong
1OpenCase Deep99.7%0.3%
2Claude Opus 5.5 (no tools)98.9%1.1%
3Muse Spark 1.3 (web search)98.8%1.2%
4Claude Opus 5.5 (web search)98.3%1.7%
5GPT 5.6 Sol (web search)96.9%3.1%
6Claude Haiku 5.5 (web search)96.7%3.3%
7Claude Haiku 5.5 (no tools)96.5%3.5%
7OpenCase Fast96.5%3.5%
9GPT 6 Luna (web search)93.8%6.2%
10Gemini 3.8 Flash (web search)93.7%6.3%
11GPT 5.6 Sol (no tools)91.0%9.0%
12GPT 6 Luna (no tools)90.0%10.0%
13Gemini 3.8 Flash (no tools)83.3%16.7%
14Muse Spark 1.3 (no tools)74.6%25.4%

What a hallucination looks like

A hallucinated citation reads like a real one. A fake case has a plausible name, a real reporter and a page number. A fake statute has a section number that fits the code around it. The only way to catch either is to look it up.

Gemini 3.8 Flash, no tools · Ninth Circuit sentencingUnited States v. Peraza, 644 F. App'x 788 (9th Cir. 2016)That page is Henry v. Regents. No Ninth Circuit decision named Peraza matches.
Muse Spark 1.3, no tools · Kansas farm foreclosureK.S.A. 60-2422Kansas has no such section. The article of its code on executions and sales ends at 60-2420.
GPT 5.6 Sol, no tools · Michigan employment lawLaborde v City of Grosse Pointe Farms, 199 Mich. App. 590, 502 N.W.2d 336 (1993)Those pages hold Mills v. White Castle and Price v. Long Realty. No case by this name exists.
Gemini 3.8 Flash, web search · Fifth Circuit maintenance ratesBorders v. Abdon Callais Offshore, LLC, 831 F. Supp. 2d 946 (E.D. La. 2011)That page is Myers v. Simpson, a Virginia case. No suit named Borders against Abdon Callais turns up.

What we measured

Hallucinated
No case, statute or rule matching the citation exists.
Miscited
The authority is real, but a detail is wrong, such as the page, the year or the court.
Correct
Neither hallucinated nor miscited, as a share of the citations we could check. The few we could settle neither way are left out.

How we ran it

We wrote 120 questions of the kind a practicing attorney asks, each built around a real decision from a named American court. Every model answered every question twice: once with no tools, and once with its maker's own web search, at its maker's default reasoning level. OpenCase answered twice too, in its Fast and Deep modes.

A checker with web access looked up each citation in court websites, official reporters, the Caselaw Access Project and other open sources. A second, separate pass re-checked every citation the first pass marked as hallucinated.

Limitations

  • OpenCase designed this benchmark and is one of the setups tested. We publish every question, citation and verdict so anyone can re-check our work.
  • A few questions got no answer, and each is left out of that setup's rate. Muse refused 4 questions in both modes under its content filter, and its web search failed on 2 more. Gemini with web search returned no answer to 9, even after repeated retries.
  • Each model ran at its default reasoning level. A higher setting may cite more accurately, at more cost and delay.
  • The rate is per citation. It does not measure whether an answer cited enough authority, or whether a real case supports the point it was cited for. The few citations the checker could settle neither way are left out.