Agents & automation
Software that plans and carries out multi-step work with some autonomy, calling tools and acting on a person's or organisation's behalf.
Who this is for
- Teams automating multi-step work
- Developers building on agent tools
- Risk owners who must bound what an agent may do
How we will evaluate it
Agent evaluation is harder to make repeatable than single-turn tasks. Protocols will state how variation between runs is handled before any score is published.
This category uses the proposed baseline weights. See the methodology.
Category criteria
The questions we will ask of every product in this category. These are criteria, not results.
Completion and correctness
Does the agent finish multi-step tasks, and how often does it report success incorrectly?
Control and permissions
What can it access, what needs approval, and can its actions be audited?
Failure behaviour
How does it fail, and is a failure visible to the person responsible?
Cost predictability
Can you predict and cap usage-based cost for a given job?
Weights used for this category
- Capability & output quality: 25%
- Reliability & consistency: 20%
- Workflow fit & integrations: 15%
- Pricing & value: 15%
- Privacy & security evidence: 15%
- Support & documentation: 10%
Products in the queue
Alphabetical, not ranked.
- Anthropic · Coding & developmentPlannedNo evidence record publishedCriteria for Claude Code
- OpenAI · Coding & developmentPlannedNo evidence record publishedCriteria for OpenAI Codex
No ranking is published for this category. A ranking needs at least 2 published, fully evidenced evaluations under the same methodology version; 0 exist.
Planned comparisons
Guides
- How to read an intelligent-software evaluation — What a trustworthy evaluation shows, what to check before relying on a score, and the warning signs of one that is not.