docketrouter
Models / Anthropic

Claude Opus 4

by Anthropic · anthropic/claude-opus-4

Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in...

reasoningtool-usevisionreleased 2025-05-22
Legal score · raw
-
not yet benchmarked
Context
200K
max output 32K
Input
$15
per 1M tokens
Output
$75
per 1M tokens
Suite cost
-
run the suite to see

Benchmark results

TaskCategoryRawJuicedCorrectLatencyCostRan
Hearsay IdentificationEvidence------
Bluebook Citation FormatResearch & Writing------
Federal Civil ProcedureProcedure------
Limitations ArithmeticProcedure------
Contract Clause ClassificationContracts------
Citation Hallucination ResistanceReliability------

Measured by DocketBuster

These numbers come from DocketBuster's own legal battery, not from DocketRouter's suite. Latest run per metric, with n and a 95% Wilson interval where the source reports one. See docketbuster.com/benchmarks.

MetricValuenIntervalMeasured
Statute pinpoint, exact section (no retrieval)4.6% (14/307)30795% CI 2.7% to 7.5%2026-08-23
Statute pinpoint, exact section (with DocketBuster retrieval)0.0% (0/307)30795% CI 0.0% to 1.2%2026-08-25
Say-nothing rate (declines to bluff when the answer is not in the record)100.0%23995% CI 98.4% to 100.0%2026-08-21
Abstained on statute pinpoint100.0% (307/307)307count, no interval reported2026-08-25
Coaching quality (GW-14x, 0 to 8)7.20 / 8-rubric mean, no interval reported2026-08-21

Source files: hard-llm-opus.json, hard-llm-opus-fullrag.json, hard-llm-ortier-expensive-say.json, gw14x-ortier-supercharged.json, hard-llm-ortier-expensive-statute.json, hard-llm-level-opus48-statute_rag.json. Raw model name in source: claude-opus-4-8.