docketrouter

LLM Rankings

Six models we track closely. Raw = the model alone. With DocketRouter = same model grounded and verified by our layer.

62 exam-style items (hearsay, Bluebook, FRCP deadlines, limitations math, clause classification, citation traps). Frontier models pass this suite; it is a floor, not a ranking. Kept public so you can audit every answer.

Hearsay Identification. Given a short fact pattern, decide whether the out-of-court statement is being offered for the truth of the matter asserted (hearsay) or for a non-hearsay purpose (effect on listener, verbal act, state of mind, impeachment). Modeled on the LegalBench hearsay task. Items →

#ModelProviderRawWith DocketRouterΔLatencyTask costInput $/M
1DeepSeek V3.1
deepseek/deepseek-chat-v3.1
DeepSeek80%100%+201422ms$0.0110$0.55
2Grok 4.3
x-ai/grok-4.3
xAI100%100%+03281ms$0.0357$1.25
3Gemini 2.5 Pro
google/gemini-2.5-pro
Google100%100%+06106ms$0.0748$1.25
4Claude Sonnet 4.5
anthropic/claude-sonnet-4.5
Anthropic100%100%+04487ms$0.0847$3
5GPT-5
openai/gpt-5
OpenAI70%90%+206637ms$0.0621$1.25
6Qwen3 32B
qwen/qwen3-32b
Qwen80%50%-3017815ms$0.0029$0.08