docketrouter
Models / Nous

Hermes 4 70B

by Nous · nousresearch/hermes-4-70b

Hermes 4 70B is a hybrid reasoning model from Nous Research, built on Meta-Llama-3.1-70B. It introduces the same hybrid mode as the larger 405B release, allowing the model to either...

reasoningreleased 2025-08-26
Legal score · raw
-
not yet benchmarked
Context
131K
max output 118K
Input
$0.13
per 1M tokens
Output
$0.40
per 1M tokens
Suite cost
-
run the suite to see

Benchmark results

TaskCategoryRawJuicedCorrectLatencyCostRan
Hearsay IdentificationEvidence------
Bluebook Citation FormatResearch & Writing------
Federal Civil ProcedureProcedure------
Limitations ArithmeticProcedure------
Contract Clause ClassificationContracts------
Citation Hallucination ResistanceReliability------

Measured by DocketBuster

These numbers come from DocketBuster's own legal battery, not from DocketRouter's suite. Latest run per metric, with n and a 95% Wilson interval where the source reports one. See docketbuster.com/benchmarks.

MetricValuenIntervalMeasured
Statute pinpoint, exact section (no retrieval)12.4% (37/299)29995% CI 9.1% to 16.6%2026-08-23
Statute pinpoint, exact section (with DocketBuster retrieval)76.2% (234/307)30795% CI 71.2% to 80.6%2026-08-24
Say-nothing rate (declines to bluff when the answer is not in the record)96.3%30095% CI 93.6% to 97.9%2026-08-23
Abstained on statute pinpoint9.8% (30/307)307count, no interval reported2026-08-24
Coaching quality (GW-14x, 0 to 8)6.60 / 8-rubric mean, no interval reported2026-08-24

Source files: hard-llm-ortier-hermes4-70b.json, hard-llm-level-hermes4-70b-statute_rag.json, gw14x-ortier-hermes4-70b.json. Raw model name in source: nousresearch/hermes-4-70b.