A score is only useful when you can inspect how it was made.
How products are evaluated, what counts as evidence, and how commercial relationships are kept out of the result.
Six dimensions, proposed weights
Weights are whole percentages that total exactly 100. Categories may define their own weights; each change is recorded in the changelog below.
Dimensions
- Capability & output quality25%
- Reliability & consistency20%
- Workflow fit & integrations15%
- Pricing & value15%
- Privacy & security evidence15%
- Support & documentation10%
| Dimension | Weight |
|---|---|
| Capability & output quality | 25% |
| Reliability & consistency | 20% |
| Workflow fit & integrations | 15% |
| Pricing & value | 15% |
| Privacy & security evidence | 15% |
| Support & documentation | 10% |
| Total | 100% |
What each dimension means
Dimension definitions
- Capability & output quality
How well the product performs the tasks it is meant for, judged against a defined task set.
- Task success against a published rubric
- Quality of outputs on representative work
- Handling of hard or ambiguous inputs
- Reliability & consistency
Whether results hold up when the same task is repeated and over time.
- Variation between repeated runs
- Failure modes and how visible they are
- Stability across product versions
- Workflow fit & integrations
How well the product fits the tools, environments, and habits of its intended users.
- Supported environments and integrations
- Administrative and review controls
- Effort needed to adopt
- Pricing & value
What it costs for the usage a typical buyer in the category needs, and what limits apply.
- Verified, dated pricing
- Usage limits and overage rules
- Cost relative to the alternatives tested
- Privacy & security evidence
What documented, checkable evidence exists about data handling and security for the specific product and deployment.
- Data-use and retention terms
- Scoped certifications or authorizations from authoritative sources
- Administrative data controls
- Support & documentation
The quality of documentation and the support a customer can expect.
- Documentation accuracy and coverage
- Published support channels and terms
- Change communication
How a score is calculated
Calculation
- Each dimension is scored from 0 to 10, in steps of 0.1, from recorded evidence that is independently measured or independently observed.
- The final score is the sum of each score multiplied by its weight, divided by 100. It is calculated with integer arithmetic and rounded half up to one decimal, so it is exactly reproducible.
- Every score links to its evidence records, and is marked as a measured result or an editorial judgment.
- Each evaluation records its test protocol, test date, and the product version or model tested when known.
- Every published score comes with a plain-language explanation of how each dimension contributed.
Minimum evidence rules
Minimum evidence
- A final score is calculated only when every weighted dimension is backed by independently measured or independently observed evidence with at least one linked evidence record.
- Vendor-reported and unknown information is shown, labelled, and excluded from scoring. It is never replaced by an average or a neutral value.
- A ranking is published only when at least two published, fully evidenced evaluations exist for the same category under the same methodology version.
- Each evaluation records its test protocol, test date, and, where available, the product version or model tested.
Evidence classes
Evidence classes
- Vendor-reported claim
A statement made by the company that sells the product. Useful as a lead, not as a finding.
- Independently observed
Behaviour a reviewer saw directly but did not measure under a repeatable protocol.
- Independently measured
A result produced under a documented, repeatable protocol run by someone independent of the vendor.
- Third-party research
Work published by researchers or institutions other than us or the vendor. We link it; we have not reproduced it.
- Editorial interpretation
Our reading of what the evidence means. Judgment, clearly separated from the evidence itself.
- Unverified
A claim we could not confirm, or have not yet checked. It should not be relied on.
Evaluation statuses
Evaluation statuses
- Planned
- Planned
Selected for evaluation. No testing has started and no evidence is published.
- Research in progress
- Research in progress
Testing or evidence gathering has started. No conclusions are published.
- Published
- Published
A completed evaluation with its evidence and protocol is available.
- Needs review
- Needs review
A published evaluation that may be out of date or is under re-check.
- Archived
- Archived
Retained for the record. No longer maintained.
Commercial relationships
Commercial relationships
- Affiliate commissions, sponsorships, and vendor payments are never inputs to a score, a ranking, or a conclusion.
- The scoring code has no access to commercial fields; a test enforces this.
- Affiliate links are labelled where they appear and carry rel="sponsored".
- Sponsored content, if ever offered, will be labelled as such and kept out of rankings.
Changelog
Every change to dimensions, weights, or evidence rules is recorded here.
Methodology changelog
- v0.1.0
- — Initial proposed baseline: six dimensions with weights 25/20/15/15/15/10. No evaluation has been published under this version. Category-specific weights are not yet defined.