Benchmarks
Two suites, kept separate. Humanity's Last Lawsuit is the one that ranks models. Fundamentals is a floor every frontier model clears. Every number on this page comes from a run we executed; nothing is projected.
HLL 0.1-TX
Humanity's Last Lawsuit
Blinded real Texas appellate cases. Score = issue, standard, authority, application, outcome, procedure; fabricated citation caps at 25%.
363
audited items
147
source opinions
34
graded answers
Score by model
raw juiced
saturated
Fundamentals
62 exam-style items across six tasks. Exact-match grading, temperature 0.
Overall by model
raw juiced
Claude Sonnet 4.567% → 98%
Grok 4.399% → 97%
Gemini 2.5 Pro96% → 97%
DeepSeek V3.186% → 94%
GPT-574% → 78%
Qwen3 32B70% → 74%
Mean by task (featured models)
Hearsay Identification88% → 90%
Bluebook Citation Format79% → 85%
Federal Civil Procedure97% → 99%
Limitations Arithmetic63% → 83%
Contract Clause Classification100% → 100%
Citation Hallucination Resistance64% → 81%
Fundamentals tasks
Public items with gold labels, so every answer is auditable.
Procedure