Coding & development
Tools that help write, explain, review, test, and maintain code, from in-editor completion to assistants that work across a whole repository.
Who this is for
- Individual developers
- Engineering teams
- Engineering and security leads who set policy for code and data
How we will evaluate it
Coding evaluations will use the baseline weights until a task set and protocol are published. Capability will be measured on tasks we publish in advance, with the number of runs and the pass criteria stated.
This category uses the proposed baseline weights. See the methodology.
Category criteria
The questions we will ask of every product in this category. These are criteria, not results.
Task completion on real repositories
Does the tool finish defined tasks correctly in a codebase similar to yours, including passing the project's own tests?
Reviewability
Can a human see, understand, and reverse what the tool changed?
Context handling
How well does it use code outside the open file, and how does it behave when context runs out?
Environment fit
Which editors, languages, and CI or review systems does it support?
Code and data handling
What happens to proprietary code: retention, training use, and administrative controls?
Weights used for this category
- Capability & output quality: 25%
- Reliability & consistency: 20%
- Workflow fit & integrations: 15%
- Pricing & value: 15%
- Privacy & security evidence: 15%
- Support & documentation: 10%
Products in the queue
Alphabetical, not ranked.
- Anthropic · Coding & developmentPlannedNo evidence record publishedCriteria for Claude Code
- Anysphere · Coding & developmentPlannedNo evidence record publishedCriteria for Cursor
- Google · Coding & developmentPlannedNo evidence record publishedCriteria for Gemini Code Assist
- GitHub · Coding & developmentPlannedNo evidence record publishedCriteria for GitHub Copilot
- OpenAI · Coding & developmentPlannedNo evidence record publishedCriteria for OpenAI Codex
No ranking is published for this category. A ranking needs at least 2 published, fully evidenced evaluations under the same methodology version; 0 exist.
Planned comparisons
Guides
- Choosing a coding assistant: what to test in your own code — A practical way to trial AI coding tools on your own repositories before you rely on published comparisons.