OpenCase Benchmarks
How often AI cites the law correctly
We asked leading AI models and OpenCase 120 hard questions about American law, then looked up every case, statute and rule they cited. The questions, citations and verdicts are all public.
- OpenCase Deep99.7%
- Claude Opus 5.5 (no tools)98.9%
- Muse Spark 1.3 (web search)98.8%
- Claude Opus 5.5 (web search)98.3%
- GPT 5.6 Sol (web search)96.9%
- Claude Haiku 5.5 (web search)96.7%
- Claude Haiku 5.5 (no tools)96.5%
- OpenCase Fast96.5%
- GPT 6 Luna (web search)93.8%
- Gemini 3.8 Flash (web search)93.7%
- GPT 5.6 Sol (no tools)91.0%
- GPT 6 Luna (no tools)90.0%
- Gemini 3.8 Flash (no tools)83.3%
- Muse Spark 1.3 (no tools)74.6%
Share of citations that point to a real case, statute or rule with the right page, year and court. Higher is better. 120 questions.
Dataset
120 questions an attorney might research, each about a real decision from a named American court. 41 follow decisions from 2025 and 2026, which a model may not have seen. 43 ask about narrow points of law. 36 ask about older law, from decisions between 1881 and 1970, which is rarely cited today. We wrote them in batches of 20 and made each batch hard enough that at least one setup cited an authority that does not exist.
Evaluation methodology
Each model answered every question twice: once with no tools, and once with its maker's own web search, at its maker's default reasoning level. OpenCase answered in its Fast and Deep modes. A question a setup left unanswered is left out of its rate. A checker with web access looked up every citation in court websites, official reporters, the Caselaw Access Project and other open sources. A second pass re-checked every citation marked hallucinated.
What counts as correct
A correct citation is neither of these. Hallucinated: no case, statute or rule matching the citation exists. Miscited: the authority is real, but the page, year or court is wrong. The few citations the checker could settle neither way are left out.
Testing dates
Run in September and October 2026 with gpt-5.6-sol, gpt-6-luna, claude-opus-5-5, claude-haiku-5-5, gemini-3.8-flash and muse-spark-1.3.
| # | SETUP | CORRECT |
|---|---|---|
| 1 | OpenCase Deep | 99.7% |
| 2 | Claude Opus 5.5 (no tools) | 98.9% |
| 3 | Muse Spark 1.3 (web search) | 98.8% |
| 4 | Claude Opus 5.5 (web search) | 98.3% |
| 5 | GPT 5.6 Sol (web search) | 96.9% |
| 6 | Claude Haiku 5.5 (web search) | 96.7% |
| 7 | Claude Haiku 5.5 (no tools) | 96.5% |
| 7 | OpenCase Fast | 96.5% |
| 9 | GPT 6 Luna (web search) | 93.8% |
| 10 | Gemini 3.8 Flash (web search) | 93.7% |
| 11 | GPT 5.6 Sol (no tools) | 91.0% |
| 12 | GPT 6 Luna (no tools) | 90.0% |
| 13 | Gemini 3.8 Flash (no tools) | 83.3% |
| 14 | Muse Spark 1.3 (no tools) | 74.6% |