docketrouter
Models / Nous

Hermes 4 405B

by Nous · nousresearch/hermes-4-405b

Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mode, where the model can choose to deliberate internally with...

reasoningreleased 2025-08-26
Legal score · raw
-
not yet benchmarked
Context
131K
max output 118K
Input
$1
per 1M tokens
Output
$3
per 1M tokens
Suite cost
-
run the suite to see

Benchmark results

TaskCategoryRawJuicedCorrectLatencyCostRan
Hearsay IdentificationEvidence------
Bluebook Citation FormatResearch & Writing------
Federal Civil ProcedureProcedure------
Limitations ArithmeticProcedure------
Contract Clause ClassificationContracts------
Citation Hallucination ResistanceReliability------

Measured by DocketBuster

These numbers come from DocketBuster's own legal battery, not from DocketRouter's suite. Latest run per metric, with n and a 95% Wilson interval where the source reports one. See docketbuster.com/benchmarks.

MetricValuenIntervalMeasured
Statute pinpoint, exact section (no retrieval)7.7% (23/300)30095% CI 5.2% to 11.2%2026-08-22
Statute pinpoint, exact section (with DocketBuster retrieval)75.6% (232/307)30795% CI 70.5% to 80.0%2026-08-24
Say-nothing rate (declines to bluff when the answer is not in the record)99.6%24495% CI 97.7% to 99.9%2026-08-22
Abstained on statute pinpoint16.3% (50/307)307count, no interval reported2026-08-24

Source files: hard-llm-ortier-hermes4-405b.json, hard-llm-level-hermes4-405b-statute_rag.json. Raw model name in source: nousresearch/hermes-4-405b.