docketrouter

LLM Rankings

Six models we track closely. Raw = the model alone. Juiced = same model through the DocketRouter legal model.

62 exam-style items (hearsay, Bluebook, FRCP deadlines, limitations math, clause classification, citation traps). Frontier models pass this suite; it is a floor, not a ranking. Kept public so you can audit every answer.

#ModelProviderRawJuicedΔLatencySuite costInput $/M
1Claude Sonnet 4.5
anthropic/claude-sonnet-4.5
Anthropic67%98%+314487ms$0.6010$3
2Grok 4.3
x-ai/grok-4.3
xAI99%97%-13412ms$0.2496$1.25
3Gemini 2.5 Pro
google/gemini-2.5-pro
Google96%97%+16216ms$0.4913$1.25
4DeepSeek V3.1
deepseek/deepseek-chat-v3.1
DeepSeek86%94%+81941ms$0.0820$0.55
5GPT-5
openai/gpt-5
OpenAI74%78%+36222ms$0.3540$1.25
6Qwen3 32B
qwen/qwen3-32b
Qwen70%74%+512278ms$0.0190$0.08