docketrouter

LLM Rankings

Six models we track closely. Raw = the model alone. With DocketRouter = same model grounded and verified by our layer.

62 exam-style items (hearsay, Bluebook, FRCP deadlines, limitations math, clause classification, citation traps). Frontier models pass this suite; it is a floor, not a ranking. Kept public so you can audit every answer.

Federal Civil Procedure. Short-answer questions on the Federal Rules of Civil and Appellate Procedure: which rule governs a motion, how many days a deadline is, numeric limits on discovery. These are the kinds of facts a paralegal must never get wrong. Items →

#ModelProviderRawWith DocketRouterΔLatencyTask costInput $/M
1DeepSeek V3.1
deepseek/deepseek-chat-v3.1
DeepSeek100%100%+01453ms$0.0153$0.55
2Grok 4.3
x-ai/grok-4.3
xAI100%100%+02247ms$0.0432$1.25
3GPT-5
openai/gpt-5
OpenAI100%100%+02911ms$0.0516$1.25
4Gemini 2.5 Pro
google/gemini-2.5-pro
Google100%100%+05245ms$0.0833$1.25
5Claude Sonnet 4.5
anthropic/claude-sonnet-4.5
Anthropic100%100%+02981ms$0.0932$3
6Qwen3 32B
qwen/qwen3-32b
Qwen83%92%+810529ms$0.0032$0.08