Objection Academy public good

EvidenceBench

Attorney-crafted evidence questions. Transparent model scores. Citation reliability that is easy to inspect.

v3.0.0

Official scores are live. The public evaluator, development questions, protocol, and aggregate six-model snapshot are available now.

EvidenceBench measures whether leading AI models can identify the best evidence answer and cite the rules and cases that support it. It is a public educational benchmark—not legal advice, legal research, or an endorsement of any model.

Official holdout

151

sealed questions

Caselaw module

55

case questions

Data.gov expansion

33

published 2020–2024 cases

Current-law tranche

6

published 2026 decisions

Public development set

24

reviewable questions

Tools during runs

Disabled

closed-book protocol

Snapshot charts

These charts use the same checked-in aggregate results as the tables below. Scores are percentages; higher is better except for hallucination rate.

Overall benchmark score

70% answer accuracy · 30% citation F1

  1. 1. Gemini 3.5 Flash
  2. 2. Claude Fable 5
  3. 3. Grok 4.5
  4. 4. GPT-5.6 Sol
  5. 5. Kimi K3
  6. 6. Mistral Large 3
  7. 7. DeepSeek V4 Pro

Caselaw by source tranche

All caselaw Data.gov 2020–2024Published 2026
  1. Grok 4.5
  2. Claude Fable 5
  3. Gemini 3.5 Flash
  4. DeepSeek V4 Pro
  5. GPT-5.6 Sol
  6. Kimi K3
  7. Mistral Large 3

Caselaw dimensions

The original 16-question matrix is balanced across federal and state authority, civil and criminal proceedings, and widely cited and less familiar decisions. The expanded source tranches reflect the decisions available under their disclosed selection rules, so the complete 55-case module is intentionally not described as balanced. Scores below isolate performance on each axis.

ModelFRECaselawFederalStateCivilCriminalPopularObscure
Gemini 3.5 Flash96.6%65.3%71.3%60.7%80.0%59.8%98.8%59.6%
Claude Fable 594.3%67.5%74.8%61.9%77.9%63.6%96.8%62.5%
Grok 4.593.5%68.4%77.1%61.6%76.7%65.3%100.0%63.0%
GPT-5.6 Sol98.3%48.9%60.0%40.3%64.0%43.3%100.0%40.2%
Kimi K395.9%45.1%56.3%36.4%72.0%35.0%98.8%36.0%
Mistral Large 396.0%43.6%47.9%40.3%67.3%34.8%95.0%34.9%
DeepSeek V4 Pro90.8%49.1%58.8%41.6%56.0%46.5%100.0%40.4%

Recent-source performance

EvidenceBench v3 adds every unique, citable 2020–2024 decision marked published in the public-domain DOJ/NIJ Post-PCAST dataset cataloged by Data.gov, plus six decisions issued in 2026 and verified against official federal court, GovInfo, or New York Official Reports sources. Source snapshots and hashes are checked into the public repository.

ModelAll caselawData.gov 2020–2024Published 20262022–2024
Gemini 3.5 Flash65.3%48.8%70.0%50.7%
Claude Fable 567.5%53.0%70.0%44.3%
Grok 4.568.4%53.0%70.0%49.3%
GPT-5.6 Sol48.9%27.0%35.0%19.3%
Kimi K345.1%14.5%70.0%7.1%
Mistral Large 343.6%14.8%70.0%15.0%
DeepSeek V4 Pro49.1%30.0%35.0%32.1%

Data.gov source hash: 90990eb660d45be39515f9f08278abd4b080be3ce758eb1cf9e8854cc5ab95f1. Current 2026 metadata hash: ea97ec705a56e141df2a772b0644a7c33731ace07b13812314ecbabb8dde4467.

Leaderboard

This is the current approved, closed-book snapshot. Model routes, dates, costs, and aggregate metrics are recorded in the public manifest.

ModelFamilyOverallAccuracyCitation F1Hallucination
Gemini 3.5 FlashOfficial snapshotGoogle85.2%90.1%73.7%15.2%
Claude Fable 5Official snapshotAnthropic84.5%89.4%73.2%10.4%
Grok 4.5Official snapshotxAI84.3%91.4%67.9%6.9%
GPT-5.6 SolOfficial snapshotOpenAI80.3%82.1%76.2%3.6%
Kimi K3Official snapshotMoonshot AI77.4%81.5%67.9%7.9%
Mistral Large 3Official snapshotMistral76.9%82.8%63.2%17.3%
DeepSeek V4 ProOfficial snapshotDeepSeek75.6%80.1%65.1%6.9%

Model details

Gemini 3.5 Flash

Google · official

Claude Fable 5

Anthropic · official

Grok 4.5

xAI · official

GPT-5.6 Sol

OpenAI · official

Kimi K3

Moonshot AI · official

Mistral Large 3

Mistral · official

DeepSeek V4 Pro

DeepSeek · official

Category coverage

The sealed holdout covers core evidence topics, the original eight balanced caselaw intersections, a Data.gov-backed forensic-evidence tranche, and published 2026 decisions.

  • relevance
  • hearsay
  • witnesses
  • experts
  • authentication
  • best evidence
  • character
  • impeachment
  • privilege
  • caselaw

Public development-question explorer

EB-DEV-004 · Hearsay

A witness testifies that a bystander said the light was red, offered to prove the light was red. What is the best objection?

Gold answer

Hearsay

Citations

FRE 801(c) · FRE 802

Attorney rationale

The statement is an out-of-court assertion offered for its truth and no exception is identified.

The public development set and deterministic evaluator are available in the repository; sealed item-level outputs are never published.

How scoring works

Each response must return an option, a concise explanation, and normalized FRE or case-reporter citations. Overall score is 0.70 * answer_accuracy + 0.30 * citation_f1. Citation F1 rewards required support while penalizing missing, unsupported, and nonexistent citations.

Every official run is closed-book: no browsing, retrieval, or tools. We publish aggregate results and a dataset commitment, but not the sealed questions or item-level outputs.

The authority corpus is frozen at the Federal Rules effective December 1, 2025. The Rule 801(d)(1)(A) amendment remains pending in July 2026 and is excluded until its scheduled December 1, 2026 effective date.