Key findings
- Without tools, Muse Spark 1.3 got 25.4% of its citations wrong and Gemini 3.8 Flash 16.7%. Claude Opus 5.5 got 1.1% wrong, the fewest of any model.
- Without tools, Muse invented at least one authority in 45% of its answers, and Gemini in 41%.
- Web search helped the two weakest models most. Muse fell to 1.2% wrong and Gemini to 6.3%. Claude Opus moved from 1.1% to 1.7%.
- OpenCase Deep was the most accurate setup we tested: 99.7% of its citations were correct. Neither OpenCase mode hallucinated an authority. Every OpenCase error was a real authority with a wrong detail.
| # | Setup | Correct | Wrong |
|---|---|---|---|
| 1 | OpenCase Deep | 99.7% | 0.3% |
| 2 | Claude Opus 5.5 (no tools) | 98.9% | 1.1% |
| 3 | Muse Spark 1.3 (web search) | 98.8% | 1.2% |
| 4 | Claude Opus 5.5 (web search) | 98.3% | 1.7% |
| 5 | GPT 5.6 Sol (web search) | 96.9% | 3.1% |
| 6 | Claude Haiku 5.5 (web search) | 96.7% | 3.3% |
| 7 | Claude Haiku 5.5 (no tools) | 96.5% | 3.5% |
| 7 | OpenCase Fast | 96.5% | 3.5% |
| 9 | GPT 6 Luna (web search) | 93.8% | 6.2% |
| 10 | Gemini 3.8 Flash (web search) | 93.7% | 6.3% |
| 11 | GPT 5.6 Sol (no tools) | 91.0% | 9.0% |
| 12 | GPT 6 Luna (no tools) | 90.0% | 10.0% |
| 13 | Gemini 3.8 Flash (no tools) | 83.3% | 16.7% |
| 14 | Muse Spark 1.3 (no tools) | 74.6% | 25.4% |
What a hallucination looks like
A hallucinated citation reads like a real one. A fake case has a plausible name, a real reporter and a page number. A fake statute has a section number that fits the code around it. The only way to catch either is to look it up.
What we measured
- Hallucinated
- No case, statute or rule matching the citation exists.
- Miscited
- The authority is real, but a detail is wrong, such as the page, the year or the court.
- Correct
- Neither hallucinated nor miscited, as a share of the citations we could check. The few we could settle neither way are left out.
How we ran it
We wrote 120 questions of the kind a practicing attorney asks, each built around a real decision from a named American court. Every model answered every question twice: once with no tools, and once with its maker's own web search, at its maker's default reasoning level. OpenCase answered twice too, in its Fast and Deep modes.
A checker with web access looked up each citation in court websites, official reporters, the Caselaw Access Project and other open sources. A second, separate pass re-checked every citation the first pass marked as hallucinated.
Limitations
- OpenCase designed this benchmark and is one of the setups tested. We publish every question, citation and verdict so anyone can re-check our work.
- A few questions got no answer, and each is left out of that setup's rate. Muse refused 4 questions in both modes under its content filter, and its web search failed on 2 more. Gemini with web search returned no answer to 9, even after repeated retries.
- Each model ran at its default reasoning level. A higher setting may cite more accurately, at more cost and delay.
- The rate is per citation. It does not measure whether an answer cited enough authority, or whether a real case supports the point it was cited for. The few citations the checker could settle neither way are left out.