Benchmark Report - AIBE 20 - AI-Powered Legal Platform
Skip to main content

LeXi AI AIBE 20 Evaluation Report.

An independent evaluation of AI performance across AIBE 20 questions — measuring accuracy, interpretive logic, contextual reasoning, and response quality across five major models.

We conducted this evaluation to establish a transparent benchmark for legal AI — one that goes beyond accuracy to test how well AI systems reason through ambiguity, uphold ethical standards, and perform under real-world legal conditions.

I · LEADERBOARD

The order of merit

Ranked by accuracy on 95 scored questions. Five withdrawn items excluded from all totals.

LeXi AI

ACCURACY
98.95%
CORRECT94/95
WRONG1
AVG / Q10.56s

GPT 5.5

ACCURACY
96.84%
CORRECT92/95
WRONG3
AVG / Q7.96s

Gemini 3.1 Pro

ACCURACY
95.79%
CORRECT91/95
WRONG4
AVG / Q7.08s

Deepseek v3.2

ACCURACY
89.47%
CORRECT85/95
WRONG10
AVG / Q3.96s

Claude Opus 4.8

ACCURACY
88.42%
CORRECT84/95
WRONG11
AVG / Q3.81s
II · THE LEAP

A measurable step-change.

LeXi AI, measured against the same exam format one cycle earlier — sharper on every axis that matters for legal reasoning.

METRIC
PREVIOUS
CURRENT
Δ
01Correct answers
85/93
94/95
+9
02Accuracy
91.4%
98.95%
+7.55 pp
03Avg response time
20s +
11s
45% faster
IN SUMMARY

Nine more correct. Nearly half the latency. Accuracy within a point of the ceiling.

+9ANSWERS
45%FASTER
+7.55PP ACCURACY
III · ACCURACY VS CORRECT ANSWERS

More correct, higher accuracy.

Every point is a model. The horizontal axis is number of correct answers; the vertical axis is accuracy. Hover a point to inspect it.

98%94%90%86%8084889296100QUESTIONS CORRECTACCURACYLeXi AIGPT 5.5Gemini 3.1 ProDeepseek v3.2Claude Opus 4.8

Explore Question by Question Analysis

Dive deep into individual AIBE 20 questions to see how each model answered. View detailed answer comparisons, incorrect answer analysis, and performance patterns across all 95 scored questions.

View all 95 Questions