Benchmarks measure what AI can do.
We measure what it actually does.

Error rates, costs, and failure modes from AI systems running in production — ours. Every figure here comes from a run we can point to.

Latest report — August 13, 2026

5.8%

of claims got flatly opposite verdicts: one model said the sources support it, the other said they do not. Same claim, same sources, same prompt.

Agreed 87.2%Contradicted 5.8%One hedged 7.0%

n = 399 claims, each judged twice (798 model calls) · window claims drawn from 2026-05-03 → 2026-08-09, run 2026-08-13

All reports