Benchmarks measure what AI can do.
We measure what it actually does.
Error rates, costs, and failure modes from AI systems running in production — ours. Every figure here comes from a run we can point to.
of claims got flatly opposite verdicts: one model said the sources support it, the other said they do not. Same claim, same sources, same prompt.
We gave the same claims to two models. They contradict each other on 1 in 17.
All reports
We gave the same claims to two models. They contradict each other on 1 in 17.
Our fact checker is a model, and we had never measured its error rate. So we put 399 claims to Claude and GPT-4o under identical conditions: same sources, same prompt, same output contract. They agree 87% of the time, and where they disagree is the part that matters.
The same model fabricates twice as often writing about politics as about technology.
Three months of automated fact-checking split by subject. Same models, same prompts, same pipeline: the failure rate depends on what the article is about, and the gap runs opposite to the usual assumption.
We fact-checked 5,500 claims from AI-written news. 30% weren't supported by the sources.
Three months of automated fact-checking across 1,140 AI-generated news articles, verified claim by claim against the source material the system itself used to write them.