Category

Agents & automation

Software that plans and carries out multi-step work with some autonomy, calling tools and acting on a person's or organisation's behalf.

Who this is for

  • Teams automating multi-step work
  • Developers building on agent tools
  • Risk owners who must bound what an agent may do

How we will evaluate it

Agent evaluation is harder to make repeatable than single-turn tasks. Protocols will state how variation between runs is handled before any score is published.

This category uses the proposed baseline weights. See the methodology.

Category criteria

The questions we will ask of every product in this category. These are criteria, not results.

Completion and correctness

Does the agent finish multi-step tasks, and how often does it report success incorrectly?

Control and permissions

What can it access, what needs approval, and can its actions be audited?

Failure behaviour

How does it fail, and is a failure visible to the person responsible?

Cost predictability

Can you predict and cap usage-based cost for a given job?

Weights used for this category
  • Capability & output quality: 25%
  • Reliability & consistency: 20%
  • Workflow fit & integrations: 15%
  • Pricing & value: 15%
  • Privacy & security evidence: 15%
  • Support & documentation: 10%

Products in the queue

Alphabetical, not ranked.

No ranking is published for this category. A ranking needs at least 2 published, fully evidenced evaluations under the same methodology version; 0 exist.

Planned comparisons

Guides