docketrouter
Models / xAI

Grok 4.6

by xAI · x-ai/grok-4.6

Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

reasoningtool-usevisionreleased 2026-08-12
Legal score · raw
-
not yet benchmarked
Context
500K
max output 450K
Input
$2
per 1M tokens
Output
$6
per 1M tokens
Suite cost
-
run the suite to see

Benchmark results

TaskCategoryRawJuicedCorrectLatencyCostRan
Hearsay IdentificationEvidence------
Bluebook Citation FormatResearch & Writing------
Federal Civil ProcedureProcedure------
Limitations ArithmeticProcedure------
Contract Clause ClassificationContracts------
Citation Hallucination ResistanceReliability------

Measured by DocketBuster

These numbers come from DocketBuster's own legal battery, not from DocketRouter's suite. Latest run per metric, with n and a 95% Wilson interval where the source reports one. See docketbuster.com/benchmarks.

MetricValuenIntervalMeasured
Statute pinpoint, exact section (no retrieval)36.7% (110/300)30095% CI 31.4% to 42.3%2026-08-22
Statute pinpoint, exact section (with DocketBuster retrieval)0.0% (0/306)30695% CI 0.0% to 1.2%2026-08-25
Say-nothing rate (declines to bluff when the answer is not in the record)100.0%7995% CI 95.4% to 100.0%2026-08-22
Abstained on statute pinpoint100.0% (306/306)306count, no interval reported2026-08-25
Coaching quality (GW-14x, 0 to 8)7.32 / 8-rubric mean, no interval reported2026-08-23

Source files: hard-llm-ortier-grok46.json, gw14x-ortier-grok46.json, hard-llm-level-grok46-statute_rag.json. Raw model name in source: x-ai/grok-4.6.