Research that helps you see what the claims leave out.
Our approach to benchmarks, product intelligence, evidence, and explainers, and an honest account of its limits.
What we do
Research areas
Capability benchmarks
Repeatable tests with published tasks, conditions, sample sizes, and limitations. None have been published yet.
Benchmark requirementsProduct intelligence
Tracking version changes, pricing, deployment options, and documentation, each with a dated source.
Evaluation queueEvidence library
Source links, verification dates, and claim scope for every important claim.
Evidence standardTechnical explainers
Plain-language explainers that separate current software from research on, and speculation about, advanced machine intelligence.
Read an explainerPublished explainers
Published research articles
Superintelligence and today's tools: keeping the terms apart
Why the word 'superintelligence' in a brand name, a headline, or a policy document does not make a current product superintelligent.
Reading benchmark claims critically
A checklist of questions that determine whether a benchmark number tells you anything about the tool you plan to use.
Limitations of our research
Limitations
- We have not run any benchmark or product test yet. Everything published so far is methodology and explanation.
- Results from tests of AI products can vary between runs and change with product versions. We will record version and date, and results will expire.
- Where we summarise third-party research we have not reproduced it, and we label it that way.