Methodology & transparency

A score is only useful when you can inspect how it was made.

How products are evaluated, what counts as evidence, and how commercial relationships are kept out of the result.

Methodology version 0.1.0

Six dimensions, proposed weights

Weights are whole percentages that total exactly 100. Categories may define their own weights; each change is recorded in the changelog below.

Dimensions

  • Capability & output quality25%
  • Reliability & consistency20%
  • Workflow fit & integrations15%
  • Pricing & value15%
  • Privacy & security evidence15%
  • Support & documentation10%
Dimensions and weights
DimensionWeight
Capability & output quality25%
Reliability & consistency20%
Workflow fit & integrations15%
Pricing & value15%
Privacy & security evidence15%
Support & documentation10%
Total100%

What each dimension means

Dimension definitions

Capability & output quality

How well the product performs the tasks it is meant for, judged against a defined task set.

  • Task success against a published rubric
  • Quality of outputs on representative work
  • Handling of hard or ambiguous inputs
Reliability & consistency

Whether results hold up when the same task is repeated and over time.

  • Variation between repeated runs
  • Failure modes and how visible they are
  • Stability across product versions
Workflow fit & integrations

How well the product fits the tools, environments, and habits of its intended users.

  • Supported environments and integrations
  • Administrative and review controls
  • Effort needed to adopt
Pricing & value

What it costs for the usage a typical buyer in the category needs, and what limits apply.

  • Verified, dated pricing
  • Usage limits and overage rules
  • Cost relative to the alternatives tested
Privacy & security evidence

What documented, checkable evidence exists about data handling and security for the specific product and deployment.

  • Data-use and retention terms
  • Scoped certifications or authorizations from authoritative sources
  • Administrative data controls
Support & documentation

The quality of documentation and the support a customer can expect.

  • Documentation accuracy and coverage
  • Published support channels and terms
  • Change communication

How a score is calculated

Calculation

  1. Each dimension is scored from 0 to 10, in steps of 0.1, from recorded evidence that is independently measured or independently observed.
  2. The final score is the sum of each score multiplied by its weight, divided by 100. It is calculated with integer arithmetic and rounded half up to one decimal, so it is exactly reproducible.
  3. Every score links to its evidence records, and is marked as a measured result or an editorial judgment.
  4. Each evaluation records its test protocol, test date, and the product version or model tested when known.
  5. Every published score comes with a plain-language explanation of how each dimension contributed.

Minimum evidence rules

Minimum evidence

  • A final score is calculated only when every weighted dimension is backed by independently measured or independently observed evidence with at least one linked evidence record.
  • Vendor-reported and unknown information is shown, labelled, and excluded from scoring. It is never replaced by an average or a neutral value.
  • A ranking is published only when at least two published, fully evidenced evaluations exist for the same category under the same methodology version.
  • Each evaluation records its test protocol, test date, and, where available, the product version or model tested.

Evidence classes

Evidence classes

  • Vendor-reported claim

    A statement made by the company that sells the product. Useful as a lead, not as a finding.

  • Independently observed

    Behaviour a reviewer saw directly but did not measure under a repeatable protocol.

  • Independently measured

    A result produced under a documented, repeatable protocol run by someone independent of the vendor.

  • Third-party research

    Work published by researchers or institutions other than us or the vendor. We link it; we have not reproduced it.

  • Editorial interpretation

    Our reading of what the evidence means. Judgment, clearly separated from the evidence itself.

  • Unverified

    A claim we could not confirm, or have not yet checked. It should not be relied on.

Evaluation statuses

Evaluation statuses

Planned
Planned

Selected for evaluation. No testing has started and no evidence is published.

Research in progress
Research in progress

Testing or evidence gathering has started. No conclusions are published.

Published
Published

A completed evaluation with its evidence and protocol is available.

Needs review
Needs review

A published evaluation that may be out of date or is under re-check.

Archived
Archived

Retained for the record. No longer maintained.

Commercial relationships

Commercial relationships

  • Affiliate commissions, sponsorships, and vendor payments are never inputs to a score, a ranking, or a conclusion.
  • The scoring code has no access to commercial fields; a test enforces this.
  • Affiliate links are labelled where they appear and carry rel="sponsored".
  • Sponsored content, if ever offered, will be labelled as such and kept out of rankings.

Changelog

Every change to dimensions, weights, or evidence rules is recorded here.

Methodology changelog

v0.1.0
— Initial proposed baseline: six dimensions with weights 25/20/15/15/15/10. No evaluation has been published under this version. Category-specific weights are not yet defined.