Official scores are live. The public evaluator, development questions, protocol, and aggregate six-model snapshot are available now.
EvidenceBench measures whether leading AI models can identify the best evidence answer and cite the rules and cases that support it. It is a public educational benchmark—not legal advice, legal research, or an endorsement of any model.
Official holdout
151
sealed questions
Caselaw module
55
case questions
Data.gov expansion
33
published 2020–2024 cases
Current-law tranche
6
published 2026 decisions
Public development set
24
reviewable questions
Tools during runs
Disabled
closed-book protocol
Snapshot charts
These charts use the same checked-in aggregate results as the tables below. Scores are percentages; higher is better except for hallucination rate.
Overall benchmark score
70% answer accuracy · 30% citation F1
- 1. Gemini 3.5 Flash
- 2. Claude Fable 5
- 3. Grok 4.5
- 4. GPT-5.6 Sol
- 5. Kimi K3
- 6. Mistral Large 3
- 7. DeepSeek V4 Pro
Caselaw by source tranche
- Grok 4.5
- Claude Fable 5
- Gemini 3.5 Flash
- DeepSeek V4 Pro
- GPT-5.6 Sol
- Kimi K3
- Mistral Large 3
Caselaw dimensions
The original 16-question matrix is balanced across federal and state authority, civil and criminal proceedings, and widely cited and less familiar decisions. The expanded source tranches reflect the decisions available under their disclosed selection rules, so the complete 55-case module is intentionally not described as balanced. Scores below isolate performance on each axis.
| Model | FRE | Caselaw | Federal | State | Civil | Criminal | Popular | Obscure |
|---|---|---|---|---|---|---|---|---|
| Gemini 3.5 Flash | 96.6% | 65.3% | 71.3% | 60.7% | 80.0% | 59.8% | 98.8% | 59.6% |
| Claude Fable 5 | 94.3% | 67.5% | 74.8% | 61.9% | 77.9% | 63.6% | 96.8% | 62.5% |
| Grok 4.5 | 93.5% | 68.4% | 77.1% | 61.6% | 76.7% | 65.3% | 100.0% | 63.0% |
| GPT-5.6 Sol | 98.3% | 48.9% | 60.0% | 40.3% | 64.0% | 43.3% | 100.0% | 40.2% |
| Kimi K3 | 95.9% | 45.1% | 56.3% | 36.4% | 72.0% | 35.0% | 98.8% | 36.0% |
| Mistral Large 3 | 96.0% | 43.6% | 47.9% | 40.3% | 67.3% | 34.8% | 95.0% | 34.9% |
| DeepSeek V4 Pro | 90.8% | 49.1% | 58.8% | 41.6% | 56.0% | 46.5% | 100.0% | 40.4% |
Recent-source performance
EvidenceBench v3 adds every unique, citable 2020–2024 decision marked published in the public-domain DOJ/NIJ Post-PCAST dataset cataloged by Data.gov, plus six decisions issued in 2026 and verified against official federal court, GovInfo, or New York Official Reports sources. Source snapshots and hashes are checked into the public repository.
| Model | All caselaw | Data.gov 2020–2024 | Published 2026 | 2022–2024 |
|---|---|---|---|---|
| Gemini 3.5 Flash | 65.3% | 48.8% | 70.0% | 50.7% |
| Claude Fable 5 | 67.5% | 53.0% | 70.0% | 44.3% |
| Grok 4.5 | 68.4% | 53.0% | 70.0% | 49.3% |
| GPT-5.6 Sol | 48.9% | 27.0% | 35.0% | 19.3% |
| Kimi K3 | 45.1% | 14.5% | 70.0% | 7.1% |
| Mistral Large 3 | 43.6% | 14.8% | 70.0% | 15.0% |
| DeepSeek V4 Pro | 49.1% | 30.0% | 35.0% | 32.1% |
Data.gov source hash: 90990eb660d45be39515f9f08278abd4b080be3ce758eb1cf9e8854cc5ab95f1. Current 2026 metadata hash: ea97ec705a56e141df2a772b0644a7c33731ace07b13812314ecbabb8dde4467.
Leaderboard
This is the current approved, closed-book snapshot. Model routes, dates, costs, and aggregate metrics are recorded in the public manifest.
| Model | Family | Overall | Accuracy | Citation F1 | Hallucination |
|---|---|---|---|---|---|
| Gemini 3.5 FlashOfficial snapshot | 85.2% | 90.1% | 73.7% | 15.2% | |
| Claude Fable 5Official snapshot | Anthropic | 84.5% | 89.4% | 73.2% | 10.4% |
| Grok 4.5Official snapshot | xAI | 84.3% | 91.4% | 67.9% | 6.9% |
| GPT-5.6 SolOfficial snapshot | OpenAI | 80.3% | 82.1% | 76.2% | 3.6% |
| Kimi K3Official snapshot | Moonshot AI | 77.4% | 81.5% | 67.9% | 7.9% |
| Mistral Large 3Official snapshot | Mistral | 76.9% | 82.8% | 63.2% | 17.3% |
| DeepSeek V4 ProOfficial snapshot | DeepSeek | 75.6% | 80.1% | 65.1% | 6.9% |
Model details
Gemini 3.5 Flash
Google · official
Claude Fable 5
Anthropic · official
Grok 4.5
xAI · official
GPT-5.6 Sol
OpenAI · official
Kimi K3
Moonshot AI · official
Mistral Large 3
Mistral · official
DeepSeek V4 Pro
DeepSeek · official
Category coverage
The sealed holdout covers core evidence topics, the original eight balanced caselaw intersections, a Data.gov-backed forensic-evidence tranche, and published 2026 decisions.
- relevance
- hearsay
- witnesses
- experts
- authentication
- best evidence
- character
- impeachment
- privilege
- caselaw
Public development-question explorer
EB-DEV-004 · Hearsay
A witness testifies that a bystander said the light was red, offered to prove the light was red. What is the best objection?
Gold answer
Hearsay
Citations
FRE 801(c) · FRE 802
Attorney rationale
The statement is an out-of-court assertion offered for its truth and no exception is identified.
The public development set and deterministic evaluator are available in the repository; sealed item-level outputs are never published.
How scoring works
Each response must return an option, a concise explanation, and normalized FRE or case-reporter citations. Overall score is 0.70 * answer_accuracy + 0.30 * citation_f1. Citation F1 rewards required support while penalizing missing, unsupported, and nonexistent citations.
Every official run is closed-book: no browsing, retrieval, or tools. We publish aggregate results and a dataset commitment, but not the sealed questions or item-level outputs.
The authority corpus is frozen at the Federal Rules effective December 1, 2025. The Rule 801(d)(1)(A) amendment remains pending in July 2026 and is excluded until its scheduled December 1, 2026 effective date.