docketrouter

LLM Rankings

Six models we track closely. Raw = the model alone. With DocketRouter = same model grounded and verified by our layer.

62 exam-style items (hearsay, Bluebook, FRCP deadlines, limitations math, clause classification, citation traps). Frontier models pass this suite; it is a floor, not a ranking. Kept public so you can audit every answer.

Bluebook Citation Format. Four variants of a citation to a well-known case or statute; only one follows Bluebook form (reporter abbreviation, volume/page order, parenthetical year, section symbol). Tests fine-grained formatting discipline that matters in filed briefs. Items →

#ModelProviderRawWith DocketRouterΔLatencyTask costInput $/M
1Grok 4.3
x-ai/grok-4.3
xAI100%100%+04032ms$0.0368$1.25
2Claude Sonnet 4.5
anthropic/claude-sonnet-4.5
Anthropic0%100%+1007461ms$0.1130$3
3DeepSeek V3.1
deepseek/deepseek-chat-v3.1
DeepSeek100%88%-121542ms$0.0117$0.55
4Gemini 2.5 Pro
google/gemini-2.5-pro
Google100%88%-127538ms$0.0686$1.25
5GPT-5
openai/gpt-5
OpenAI100%75%-257796ms$0.0579$1.25
6Qwen3 32B
qwen/qwen3-32b
Qwen75%63%-1212114ms$0.0029$0.08