docketrouter

LLM Rankings

Six models we track closely. Raw = the model alone. With DocketRouter = same model grounded and verified by our layer.

62 exam-style items (hearsay, Bluebook, FRCP deadlines, limitations math, clause classification, citation traps). Frontier models pass this suite; it is a floor, not a ranking. Kept public so you can audit every answer.

Contract Clause Classification. Classify a contract clause into one of eight categories drawn from the CUAD taxonomy: Governing Law, Non-Compete, Indemnification, Limitation of Liability, Confidentiality, Termination, Assignment, Force Majeure. The core skill behind contract-review products. Items →

#ModelProviderRawWith DocketRouterΔLatencyTask costInput $/M
1Qwen3 32B
qwen/qwen3-32b
Qwen100%100%+08997ms$0.0031$0.08
2DeepSeek V3.1
deepseek/deepseek-chat-v3.1
DeepSeek100%100%+03941ms$0.0144$0.55
3Grok 4.3
x-ai/grok-4.3
xAI100%100%+02593ms$0.0421$1.25
4GPT-5
openai/gpt-5
OpenAI100%100%+03173ms$0.0490$1.25
5Gemini 2.5 Pro
google/gemini-2.5-pro
Google100%100%+06025ms$0.0931$1.25
6Claude Sonnet 4.5
anthropic/claude-sonnet-4.5
Anthropic100%100%+03842ms$0.1021$3