LeXi AI AIBE 20 Evaluation Report.
An independent evaluation of AI performance across AIBE 20 questions — measuring accuracy, interpretive logic, contextual reasoning, and response quality across five major models.
We conducted this evaluation to establish a transparent benchmark for legal AI — one that goes beyond accuracy to test how well AI systems reason through ambiguity, uphold ethical standards, and perform under real-world legal conditions.
The order of merit
Ranked by accuracy on 95 scored questions. Five withdrawn items excluded from all totals.
LeXi AI
GPT 5.5
Gemini 3.1 Pro
Deepseek v3.2
Claude Opus 4.8
A measurable step-change.
LeXi AI, measured against the same exam format one cycle earlier — sharper on every axis that matters for legal reasoning.
Nine more correct. Nearly half the latency. Accuracy within a point of the ceiling.
More correct, higher accuracy.
Every point is a model. The horizontal axis is number of correct answers; the vertical axis is accuracy. Hover a point to inspect it.
Explore Question by Question Analysis
Dive deep into individual AIBE 20 questions to see how each model answered. View detailed answer comparisons, incorrect answer analysis, and performance patterns across all 95 scored questions.
View all 95 Questions