Methodology

Everything published here is measured on systems we operate. This page states how, so any figure can be argued with.

What we measure

We run autonomous AI pipelines in production — systems that research, write, verify, and publish without a human approving each step. Those systems generate operational data as a by-product: error rates, costs per task, failure modes, and the outcome of every automated check.

We publish that data. It is not produced for a benchmark and not collected in a lab: it comes from systems doing real work under real conditions, including their bad days.

Fact-check status vocabulary

Several reports use an automated fact-checker that examines each factual claim in a generated article and tries to trace it to the source material the writing agent was given. Its verdicts:

StatusDefinition
VERIFIEDThe claim is supported by the provided sources.
FLAGGEDNo support found in the sources. The claim may still be true in the world — it was not traceable at write time.
RECENTEvent too recent to confirm against available material.
CORRECTEDThe claim was rewritten to match what the source actually says.

An article is held from publication if any claim is FLAGGED. That is a deliberately strict rule, and it makes article-level pass rates look much worse than claim-level ones. Both are reported.

Rules we follow

  • Sample size and window on every figure. A number without an n and a date range is not a measurement.
  • Raw counts, not just percentages. Percentages are derived from published counts so the two cannot drift.
  • Limitations stated in the report itself, not in a footnote nobody reads.
  • Corrections are visible. If a figure turns out to be wrong, the report is updated and the change noted, not quietly edited.
  • No projections. We publish what happened, not what we expect will happen.

Known limitations

  • Single operator. All measurements come from pipelines we run. They reflect our architecture, prompts, and model choices — not the behaviour of AI systems in general.
  • The checker is itself a model. Our fact-checker has its own error rate, which we have not yet quantified. Measuring it is planned.
  • Observational, not controlled. These are production systems, not experiments. Variables change between periods when we fix things.

Corrections

If you find an error, write to us and we will check it. Reports carry an updated date when figures change.