Coding agents and harnesses, compared

An agent is a model plus a harness; six reach a pull request, but nothing here merges or deploys unsupervised.

Providers
17
Fields compared
26
Source confidence
35%
Last verified
2026-01-15 (8mo ago)
Re-verified
every 14 days
Data confidence0% of 365 figures verified in the last 14d
80sourced of 365149stale, oldest 8mo ago
17 tools · verified 8mo ago
opencode logoopencodeProvider-agnostic open-source TUI with a client/server split and LSP awareness.84.8

Features

Surface
CLIweb
Default model
None, you choose a provider at first run
Model breadth
any provider / router
Repo indexing
AST / LSP repo map
Plan mode
Yes
Checkpoint & undo
Partial
Headless / CI
Yes
Local models
Yes

Agentic

Where it stops
4
MCP
Yes
Subagents
Yes
Hooks
Partial
Skills / commands
Yes
Background agents
Partial
Opens PRs
Yes

Governance & control

Permissions
granular allowlist
BYOK
Yes
Licence
MIT

Pricing

Pricing model
BYOK tokens only
Entry price (/seat/mo)
$0/seat/mo
Token markup
No

Traction

GitHub stars
25,000
Terminal-Bench (%)
·
Bench model
·

Positioning

Positioning
The open-source AI coding agent built for the terminal.

Timeline

First release
2025-06-01
Claude Code logoClaude CodeAnthropic's terminal-first agent, with IDE, web and GitHub Action surfaces.80.0

Features

Surface
CLIIDE extension+2
Default model
Latest Claude Opus/Sonnet, switchable with /model
Model breadth
one vendor + custom endpoint
Repo indexing
agentic grep
Plan mode
Yes
Checkpoint & undo
Yes
Headless / CI
Yes
Local models
No

Agentic

Where it stops
4
MCP
Yes
Subagents
Yes
Hooks
Yes
Skills / commands
Yes
Background agents
Yes
Opens PRs
Yes

Governance & control

Permissions
allowlist + OS sandbox
BYOK
Yes
Licence
proprietary

Pricing

Pricing model
subscription + token overage
Entry price (/seat/mo)
$20/seat/mo
Token markup
No

Traction

GitHub stars
Terminal-Bench (%)
Bench model
·

Positioning

Positioning
Anthropic's coding agent, in your terminal and wherever else you work.

Timeline

First release
2025-02-24
Codex CLI logoCodex CLIOpenAI's open-source Rust agent, with the strongest local sandboxing on this page.76.0

Features

Surface
CLIIDE extension+2
Default model
Latest GPT-5 Codex model
Model breadth
one vendor + custom endpoint
Repo indexing
agentic grep
Plan mode
Partial
Checkpoint & undo
Partial
Headless / CI
Yes
Local models
Yes

Agentic

Where it stops
4
MCP
Yes
Subagents
No
Hooks
Partial
Skills / commands
Partial
Background agents
Yes
Opens PRs
Yes

Governance & control

Permissions
allowlist + OS sandbox
BYOK
Yes
Licence
Apache-2.0

Pricing

Pricing model
subscription + token overage
Entry price (/seat/mo)
$20/seat/mo
Token markup
No

Traction

GitHub stars
42,000
Terminal-Bench (%)
Bench model
·

Positioning

Positioning
A coding agent that runs locally, in your terminal.

Timeline

First release
2025-04-16
Goose logoGooseBlock's Apache-2.0 agent, MCP-native from the start, as both a CLI and a desktop app.75.7

Features

Surface
CLIstandalone editor
Default model
None, you configure a provider on first run
Model breadth
any provider / router
Repo indexing
agentic grep
Plan mode
Yes
Checkpoint & undo
No
Headless / CI
Yes
Local models
Yes

Agentic

Where it stops
3
MCP
Yes
Subagents
Yes
Hooks
Partial
Skills / commands
Yes
Background agents
Partial
Opens PRs
No

Governance & control

Permissions
allowlist + OS sandbox
BYOK
Yes
Licence
Apache-2.0

Pricing

Pricing model
BYOK tokens only
Entry price (/seat/mo)
$0/seat/mo
Token markup
No

Traction

GitHub stars
20,000
Terminal-Bench (%)
·
Bench model
·

Positioning

Positioning
An open-source, extensible AI agent that goes beyond code suggestions.

Timeline

First release
2025-01-28
Cursor logoCursorVS Code fork with a real codebase index, background cloud agents and a CLI.75.5

Features

Surface
standalone editorCLI+2
Default model
Cursor's in-house Composer model, with frontier models selectable
Model breadth
multi-vendor
Repo indexing
embeddings index
Plan mode
Yes
Checkpoint & undo
Yes
Headless / CI
Yes
Local models
No

Agentic

Where it stops
4
MCP
Yes
Subagents
Partial
Hooks
Yes
Skills / commands
Yes
Background agents
Yes
Opens PRs
Yes

Governance & control

Permissions
allowlist + OS sandbox
BYOK
Partial
Licence
proprietary

Pricing

Pricing model
subscription + token overage
Entry price (/seat/mo)
$20/seat/mo
Token markup
Partial

Traction

GitHub stars
Terminal-Bench (%)
·
Bench model
·

Positioning

Positioning
The AI code editor, built to make you extraordinarily productive.

Timeline

First release
2023-03-01
GitHub Copilot (Agent HQ) logoGitHub Copilot (Agent HQ)Copilot's coding agent plus a control plane for running third-party agents on your repos.73.7

Features

Surface
GitHub-nativeIDE extension+3
Default model
Selectable across OpenAI, Anthropic and Google models
Model breadth
multi-vendor
Repo indexing
hybrid index + agentic
Plan mode
Partial
Checkpoint & undo
Partial
Headless / CI
Yes
Local models
Partial

Agentic

Where it stops
4
MCP
Yes
Subagents
Partial
Hooks
Unknown
Skills / commands
Yes
Background agents
Yes
Opens PRs
Yes

Governance & control

Permissions
vendor-hosted sandbox
BYOK
Partial
Licence
proprietary

Pricing

Pricing model
subscription + token overage
Entry price (/seat/mo)
$10/seat/mo
Token markup
Partial

Traction

GitHub stars
Terminal-Bench (%)
·
Bench model
·

Positioning

Positioning
Mission control for every agent working on your repositories.

Timeline

First release
2021-06-29
Kilo Code logoKilo CodeOpen-source VS Code agent in the Cline/Roo lineage, with an orchestrator mode and a CLI.72.8

Features

Surface
IDE extensionCLI
Default model
·
Model breadth
any provider / router
Repo indexing
embeddings index
Plan mode
Yes
Checkpoint & undo
Yes
Headless / CI
Partial
Local models
Yes

Agentic

Where it stops
3
MCP
Yes
Subagents
Yes
Hooks
Unknown
Skills / commands
Yes
Background agents
No
Opens PRs
No

Governance & control

Permissions
granular allowlist
BYOK
Yes
Licence
Apache-2.0

Pricing

Pricing model
BYOK tokens only
Entry price (/seat/mo)
$0/seat/mo
Token markup
No

Traction

GitHub stars
9,000
Terminal-Bench (%)
·
Bench model
·

Positioning

Positioning
Open-source AI coding agent for VS Code, with everything included.

Timeline

First release
2025-03-01
Gemini CLI logoGemini CLIApache-2.0 terminal agent from Google with an unusually generous free tier.72.6

Features

Surface
CLIIDE extension+1
Default model
Latest Gemini Pro
Model breadth
one vendor
Repo indexing
agentic grep
Plan mode
Partial
Checkpoint & undo
Yes
Headless / CI
Yes
Local models
No

Agentic

Where it stops
3
MCP
Yes
Subagents
Partial
Hooks
Unknown
Skills / commands
Yes
Background agents
Partial
Opens PRs
Partial

Governance & control

Permissions
allowlist + OS sandbox
BYOK
Yes
Licence
Apache-2.0

Pricing

Pricing model
subscription + token overage
Entry price (/seat/mo)
$0/seat/mo
Token markup
No

Traction

GitHub stars
75,000
Terminal-Bench (%)
·
Bench model
·

Positioning

Positioning
An open-source AI agent that brings Gemini into your terminal.

Timeline

First release
2025-06-25
Qwen Code logoQwen CodeApache-2.0 Gemini CLI fork tuned for Qwen3-Coder, with a large free daily quota.66.6

Features

Surface
CLI
Default model
Latest Qwen3-Coder model
Model breadth
one vendor + custom endpoint
Repo indexing
agentic grep
Plan mode
Partial
Checkpoint & undo
Yes
Headless / CI
Yes
Local models
Yes

Agentic

Where it stops
3
MCP
Yes
Subagents
Unknown
Hooks
Unknown
Skills / commands
Yes
Background agents
No
Opens PRs
No

Governance & control

Permissions
granular allowlist
BYOK
Yes
Licence
Apache-2.0

Pricing

Pricing model
BYOK tokens only
Entry price (/seat/mo)
$0/seat/mo
Token markup
No

Traction

GitHub stars
13,000
Terminal-Bench (%)
·
Bench model
·

Positioning

Positioning
A command-line coding agent tuned for the Qwen3-Coder models.

Timeline

First release
2025-07-01
Grok Build logoGrok BuildxAI's coding harness, launched July 2026, on this page but not yet verified.66.5

Features

Surface
·
Default model
·
Model breadth
·
Repo indexing
·
Plan mode
Unknown
Checkpoint & undo
Unknown
Headless / CI
Unknown
Local models
Unknown

Agentic

Where it stops
MCP
Unknown
Subagents
Unknown
Hooks
Unknown
Skills / commands
Unknown
Background agents
Unknown
Opens PRs
Unknown

Governance & control

Permissions
·
BYOK
Unknown
Licence
·

Pricing

Pricing model
·
Entry price (/seat/mo)
·
Token markup
Unknown

Traction

GitHub stars
·
Terminal-Bench (%)
·
Bench model
·

Positioning

Positioning
·

Timeline

First release
2026-07-01
Cline logoClineOpen-source VS Code and JetBrains agent built around an explicit plan/act split.64.8

Features

Surface
IDE extensionCLI
Default model
None, you pick a provider at setup
Model breadth
any provider / router
Repo indexing
agentic grep
Plan mode
Yes
Checkpoint & undo
Yes
Headless / CI
Partial
Local models
Yes

Agentic

Where it stops
3
MCP
Yes
Subagents
No
Hooks
No
Skills / commands
Yes
Background agents
No
Opens PRs
No

Governance & control

Permissions
granular allowlist
BYOK
Yes
Licence
Apache-2.0

Pricing

Pricing model
BYOK tokens only
Entry price (/seat/mo)
$0/seat/mo
Token markup
No

Traction

GitHub stars
50,000
Terminal-Bench (%)
·
Bench model
·

Positioning

Positioning
Autonomous coding agent in your IDE, with your API key.

Timeline

First release
2024-07-01
DIY harness logoDIY harnessYour own loop over a model API with an agent SDK, the build-it-yourself reference point.59.8

Features

Surface
CLI
Default model
Whichever you wire up
Model breadth
any provider / router
Repo indexing
manual context only
Plan mode
No
Checkpoint & undo
No
Headless / CI
Yes
Local models
Yes

Agentic

Where it stops
MCP
Partial
Subagents
Partial
Hooks
Yes
Skills / commands
No
Background agents
Partial
Opens PRs
Partial

Governance & control

Permissions
applies everything
BYOK
Yes
Licence

Pricing

Pricing model
BYOK tokens only
Entry price (/seat/mo)
$0/seat/mo
Token markup
No

Traction

GitHub stars
·
Terminal-Bench (%)
·
Bench model
·

Positioning

Positioning
The reference point: what you get for a few hundred lines and an API key.

Timeline

First release
·
Zed logoZedGPL-licensed Rust editor whose agent panel can host Claude Code, Codex and Gemini CLI.58.0

Features

Surface
standalone editor
Default model
None, configure a provider, or sign in for hosted models
Model breadth
any provider / router
Repo indexing
AST / LSP repo map
Plan mode
Partial
Checkpoint & undo
Yes
Headless / CI
No
Local models
Yes

Agentic

Where it stops
3
MCP
Yes
Subagents
No
Hooks
No
Skills / commands
Partial
Background agents
No
Opens PRs
No

Governance & control

Permissions
granular allowlist
BYOK
Yes
Licence
GPL-3.0

Pricing

Pricing model
BYOK tokens only
Entry price (/seat/mo)
$0/seat/mo
Token markup
No

Traction

GitHub stars
62,000
Terminal-Bench (%)
·
Bench model
·

Positioning

Positioning
The editor for humans and AI, written in Rust.

Timeline

First release
2023-03-22
Aider logoAiderMinimal terminal pair-programmer with a tree-sitter repo map and automatic git commits.57.0

Features

Surface
CLI
Default model
None, set by --model or environment keys
Model breadth
any provider / router
Repo indexing
AST / LSP repo map
Plan mode
Partial
Checkpoint & undo
Yes
Headless / CI
Yes
Local models
Yes

Agentic

Where it stops
3
MCP
No
Subagents
No
Hooks
No
Skills / commands
Partial
Background agents
No
Opens PRs
No

Governance & control

Permissions
prompts every action
BYOK
Yes
Licence
Apache-2.0

Pricing

Pricing model
BYOK tokens only
Entry price (/seat/mo)
$0/seat/mo
Token markup
No

Traction

GitHub stars
37,000
Terminal-Bench (%)
·
Bench model
·

Positioning

Positioning
AI pair programming in your terminal.

Timeline

First release
2023-05-01
Windsurf logoWindsurfAgentic IDE built around Cascade, now owned by Cognition after a turbulent 2025.53.5

Features

Surface
standalone editorIDE extension
Default model
Windsurf's in-house SWE model, with frontier models selectable
Model breadth
multi-vendor
Repo indexing
embeddings index
Plan mode
Yes
Checkpoint & undo
Yes
Headless / CI
No
Local models
No

Agentic

Where it stops
3
MCP
Yes
Subagents
Unknown
Hooks
Unknown
Skills / commands
Yes
Background agents
Partial
Opens PRs
Partial

Governance & control

Permissions
granular allowlist
BYOK
Partial
Licence
proprietary

Pricing

Pricing model
opaque credits
Entry price (/seat/mo)
$15/seat/mo
Token markup
Yes

Traction

GitHub stars
Terminal-Bench (%)
·
Bench model
·

Positioning

Positioning
The agentic IDE where you and the agent stay in flow together.

Timeline

First release
2024-11-13
Devin logoDevinFully delegated cloud engineer: file a task in Slack or the web app, get a PR back.48.9

Features

Surface
webcloud delegate
Default model
Undisclosed, Cognition-hosted, including in-house SWE models
Model breadth
no model choice
Repo indexing
hybrid index + agentic
Plan mode
Yes
Checkpoint & undo
No
Headless / CI
Partial
Local models
No

Agentic

Where it stops
4
MCP
Unknown
Subagents
No
Hooks
No
Skills / commands
Partial
Background agents
Yes
Opens PRs
Yes

Governance & control

Permissions
vendor-hosted sandbox
BYOK
No
Licence
proprietary

Pricing

Pricing model
opaque credits
Entry price (/seat/mo)
$20/seat/mo
Token markup
Yes

Traction

GitHub stars
Terminal-Bench (%)
·
Bench model
·

Positioning

Positioning
An AI software engineer you delegate tickets to.

Timeline

First release
2024-03-12
Amp logoAmpSourcegraph's opinionated agent with no model picker and an ad-supported free tier.44.2

Features

Surface
CLIIDE extension+1
Default model
Chosen by Amp, there is no model picker
Model breadth
no model choice
Repo indexing
agentic grep
Plan mode
Partial
Checkpoint & undo
Partial
Headless / CI
Yes
Local models
No

Agentic

Where it stops
3
MCP
Yes
Subagents
Yes
Hooks
Unknown
Skills / commands
Partial
Background agents
Partial
Opens PRs
No

Governance & control

Permissions
granular allowlist
BYOK
No
Licence
proprietary

Pricing

Pricing model
opaque credits
Entry price (/seat/mo)
$0/seat/mo
Token markup
Yes

Traction

GitHub stars
Terminal-Bench (%)
·
Bench model
·

Positioning

Positioning
An agentic coding tool that always runs the best model, chosen for you.

Timeline

First release
2025-05-01
yespartialnounknown
Sources shown beside each value · Learn how sourcing works
measured vendor-claimed community inferredExpand any row for the source, verification date, and caveat behind every cell.

Where each one wins

  • Teams switching model providers often, or comparing them head-to-head
  • Running the agent on a big remote machine and driving it from a laptop
  • Anyone who wants an MIT-licensed harness with subagents and a PR path

The strongest provider-agnostic harness, the only MIT one reaching a PR, but with no OS sandbox.

  • Terminal-native teams who want subagents and hooks without building them
  • Repos where an OS sandbox is a hard requirement for unattended runs
  • Anyone already paying for Max who wants the subscription to cover agent use

The most complete extensibility surface here, at the cost of lock-in to one model vendor.

  • Unattended local runs where a real OS sandbox is non-negotiable
  • Teams already on ChatGPT business plans wanting agent use covered
  • Anyone who wants an Apache-2.0 harness they can audit and fork

The best local containment story here, but weaker than Claude Code on orchestration primitives.

Goose

75.7
  • Organisations with an existing fleet of internal MCP servers
  • Recipes: parameterised, shareable tasks run non-interactively
  • Teams wanting an Apache-2.0 agent with a desktop app for non-terminal users

The most MCP-committed harness here, a safer open-source bet than most, but it stops at your working tree.

Cursor

75.5
  • Large monorepos where agentic grep runs out of context before it finds anything
  • Teams who want to review agentic diffs in an editor rather than a terminal
  • Shops wanting background PR-opening agents without wiring up CI

A real persistent codebase index inside the editor you review in, paid for in lock-in.

  • Enterprises needing audit trails and existing branch protections respected
  • Running several vendors' agents on one repo without separate integrations
  • Teams whose review process already lives entirely in GitHub

The only vendor to properly solve agent identity, with opaque premium-request billing.

  • VS Code teams who want mode-based orchestration without leaving the IDE
  • People who liked Roo Code and want a more actively developed descendant

A fast Cline and Roo Code descendant with orchestrator mode, but the least-documented commercial side here.

  • Solo developers and students who want daily agent use at zero cost
  • Teams already on Vertex who want one billing relationship
  • Anyone wanting an Apache-2.0 base to fork, as Qwen Code did

The cheapest credible daily agent, a step behind on orchestration, with its PR path in a separate Action.

  • Running open-weight coding models without building a harness
  • Regions or budgets where frontier API pricing is the blocker
  • Air-gapped or self-hosted inference behind an OpenAI-compatible endpoint

The clearest route to strong open-weight models in a mature loop, with a viable free quota.

  • Nothing we can responsibly recommend yet

Listed for completeness, scored nowhere because we have not run it yet.

Cline

64.8
  • VS Code and JetBrains teams who want BYOK without leaving the IDE
  • Reviewers who want to approve the plan before any file changes
  • MCP-heavy setups that benefit from the server marketplace

The clearest plan/act interaction and safest BYOK default, but it stays inside the editor.

  • One repeated task with a known shape and a tight tool surface
  • Environments where no third-party harness can be installed
  • Understanding what the commercial harnesses are actually doing for you

Right for narrow, repeated jobs, wrong for open-ended engineering.

Zed

58.0
  • Driving Claude Code or Gemini CLI from an editor instead of a terminal
  • Teams who want a genuinely open-source editor with local-model support

The hedge: a fast GPL editor hosting other agents over ACP, but weak as an agent itself.

Aider

57.0
  • Scripted, repeatable edits where you want a git commit per change
  • Local or unusual model endpoints other harnesses will not talk to
  • Anyone who wants to read the whole harness in an afternoon

The most auditable harness here, but visibly slowed while the category sprinted past.

  • Developers who want the agent tracking what they are editing, not just what they typed
  • Teams already invested in Codeium's retrieval and enterprise deployment

Strong editor-aware Cascade, but 2025 ownership churn is a continuity risk and credit pricing clouds cost.

Devin

48.9
  • Backlogs of small, well-specified tickets nobody wants to pick up
  • Teams who want work dispatched from Slack rather than a terminal

The clearest full delegation, great for well-specified tickets, but ACU pricing bills you afterwards.

Amp

44.2
  • Teams who would rather not run a model-selection debate every quarter
  • Sourcegraph shops wanting shared, linkable agent threads

Removing the model picker is defensible, but you control neither cost nor provider.

Which AI coding agent should I use in 2026?

  • Agent = model + harness: the model gives capability, the harness turns it into finished work.
  • The harness owns the tool loop, context strategy, permission gate and retry-on-failing-test behaviour.
  • The same frontier model in two harnesses gives wildly different completion rates on one repo.
  • A benchmark counts only when harness, model and benchmark version are all named.
  • The Terminal-Bench column ships empty rather than reprinting model scores as harness scores.
  • The killer field is where does it stop? scored 1-5: suggests → edits → tests → PR → merge/deploy.
  • Every scored harness sits at rung 3 or 4, so the column really separates working tree from pull request.
  • Claude CodeClaude Code logo, Codex CLICodex CLI logo, opencodeopencode logo, CursorCursor logo, DevinDevin logo and GitHub Copilot reach rung 4.
  • Nothing on this page ships rung 5; Copilot cannot approve its own PR, DevinDevin logo stops at review.
  • Any "autonomous deployment" claim describes a CI pipeline you wired downstream of a rung-4 agent.
  • Vendor-coupled harnesses (Claude CodeClaude Code logo, Codex CLICodex CLI logo, Gemini CLIGemini CLI logo, DevinDevin logo, Grok BuildGrok Build logo) ship agentic features first.
  • Amp sits beside them: no model picker, opaque about which of several vendors it uses.
  • Provider-agnostic harnesses (opencodeopencode logo, ClineCline logo, Goose, AiderAider logo, ZedZed logo, Qwen CodeQwen Code logo, Kilo CodeKilo Code logo) trail but survive pricing changes.
  • The agnostic camp runs local models and never leaves you renegotiating retiered subscriptions.
  • opencodeopencode logo now reaches a pull request through its own GitHub App, narrowing the gap.
  • Three pricing models: BYOK tokens (opencodeopencode logo, AiderAider logo, Goose, ZedZed logo, ClineCline logo, Qwen CodeQwen Code logo).
  • Flat subscription with soft limits: Claude CodeClaude Code logo on Pro or Max, Codex on ChatGPT plans.
  • Credits (WindsurfWindsurf logo, DevinDevin logo, Amp) are the trap: conversion rate is vendor-controlled and changeable.
  • Copilot's premium requests are the same unforecastable problem in different clothing.
  • For a budget number, a subscription or your own API key are the only honest answers.
  • Expect this page to age fast; intended cadence is 14 days and we are behind.
  • Every dated cell was last checked 15 January 2026, roughly six months old, so treat cells as leads.
  • WindsurfWindsurf logo changed hands: OpenAIOpenAI logo deal collapsed, Google licensed the tech, Cognition bought what remained.
  • AiderAider logo's release cadence slowed; GitHub rebuilt Copilot around a multi-agent control plane.
  • Grok BuildGrok Build logo appeared July 2026 with nothing verifiable published.
  • pi.dev arrived after the pass and stays in prose as one to watch, not on the roster.
  • Pick tools with low exit cost: a harness that reads AGENTS.md and takes your key is easy to replace.

The autonomy ladder, rung by rung

  • Rung 1 suggests: it proposes a diff and you apply it.
  • Rung 2 edits files: it writes directly to your working tree.
  • Rung 3 runs commands and tests: it iterates on the failures it sees.
  • Rung 4 opens a pull request: it commits, pushes a branch and raises a PR through a first-class path.
  • Rung 5 merges and deploys: it lands on the default branch or triggers a release with no human.
  • The interesting gap is between 3 and 4, and it is organisational rather than technical.
  • Rung 4 requires the vendor to solve identity: whose token pushes, what CI runs, how review comments are handled.
  • GitHub Copilot runs in Actions under a bot identity and is barred from approving its own work.
  • DevinDevin logo and CursorCursor logo solve it with a hosted VM and their own identity.
  • Claude CodeClaude Code logo and opencodeopencode logo solve it with published GitHub Apps, so licence has nothing to do with the rung.
  • Nothing here scores 1 or 2; tools that only suggest diffs are autocomplete in a different category.
  • Nothing scores 5, so the killer field is really a two-value sort.
  • Rung 3 hides a wide band: AiderAider logo fixing a failing test does more unattended work than an approval-gated IDE agent.
  • Read this column with the permissions and headless columns, not alone.
  • The gap between 4 and 5 lives in branch protection rules, not in a harness.
  • Treat any rung-5 claim as a claim to have been given credentials, and audit it accordingly.

Pricing traps

  • The flat subscription is the most honest deal here, but read the limits.
  • Every subscription harness enforces a rolling usage window in vendor-specific units, not tokens.
  • Heavy agentic use burns those windows far faster than chat; then you wait or move to metered billing.
  • Credits are worse for planning: WindsurfWindsurf logo, DevinDevin logo and Amp price in units the vendor defines.
  • The vendor can re-price the credit-to-token conversion when model prices change, and you find out afterwards.
  • Copilot's premium-request multipliers have the same shape.
  • For a defensible quarterly number, use a per-seat subscription or your own provider key.
  • BYOK looks cheapest but moves the variance onto your API bill rather than removing it.
  • A harness that re-reads a large file every turn costs multiples of one that caches aggressively.
  • At team scale, meter BYOK per repo for a fortnight and check for prompt caching before extrapolating.

How to actually choose

  • For well-specified tickets on a repo with good tests, buy autonomy.
  • Use Claude CodeClaude Code logo or Codex CLICodex CLI logo headless behind a PR, or Copilot if you already live in GitHub.
  • Optimise for rung 4, background execution and a sandbox you trust.
  • For exploratory or architectural work, buy ergonomics instead.
  • CursorCursor logo and WindsurfWindsurf logo let you review a large agentic diff in an editor with real navigation.
  • Both maintain a persistent embeddings index that grep-based harnesses cannot fake on a monorepo.
  • Kilo CodeKilo Code logo indexes too, and Copilot and DevinDevin logo read an index alongside live search.
  • What CursorCursor logo has is that index sitting inside the editor you review in.
  • ZedZed logo is the third option: a fast editor hosting other agents over the Agent Client Protocol.
  • If constrained by procurement, data residency or lock-in, take a provider-agnostic open-source harness.
  • opencodeopencode logo for a modern TUI and broad providers, Goose if MCP is central, ClineCline logo for VS Code, AiderAider logo for minimalism.
  • You will be a quarter behind on features but can repoint any of them at a new provider anytime.
  • Whatever you pick, write AGENTS.md, pin a test command, and keep the sandbox on.

What is coding agents compared: claude code vs codex cli vs cursor vs the rest?

Coding agents compared: Claude Code vs Codex CLI vs Cursor vs the rest

On toolweight, Coding agents compared: Claude CodeClaude Code logo vs Codex CLICodex CLI logo vs CursorCursor logo vs the rest means the 17 tools benchmarked on this page, Claude CodeClaude Code logo, Codex CLICodex CLI logo, Gemini CLIGemini CLI logo, opencodeopencode logo, Kilo CodeKilo Code logo, ClineCline logo, CursorCursor logo, WindsurfWindsurf logo, AiderAider logo, DevinDevin logo, Amp, ZedZed logo, Goose, GitHub Copilot (Agent HQ)GitHub Copilot (Agent HQ) logo, Grok BuildGrok Build logo, Qwen CodeQwen Code logo, DIY harness, judged on the same 26 fields, from the same sources, on the same date. The question it exists to answer: Which AI coding agent should I use in 2026?

How does toolweight compare these?

  • Every cell comes from vendor docs, the tool's repo, or hands-on use, each carrying its provenance.
  • Pricing-page facts are marked vendor-claimed with the URL.
  • Inferences from adjacent behaviour, and judgement calls like scores, are marked inferred.
  • Widely reported but undocumented facts are marked community.
  • Unknown facts are null and render as a dash rather than a guess.
  • Closed-source harnesses have null GitHub stars because no comparable repo exists.
  • Verify cadence is 14 days; the last full pass was 15 January 2026, well outside it.
  • We publish dates rather than refresh them; a moved date without a redone check is a lie.
  • Read every date as the last time a human actually looked.
  • The killer field is scored on a fixed five-rung ladder, at the highest rung reached out of the box.
  • Rung 4 needs a first-class PR path: built-in command, vendor-published GitHub App, or hosted runner.
  • Invoking gh because you asked it to is not rung 4; any shell can be scripted up a rung.
  • A vendor Action with no App identity behind it is graded partial, for first- and third-party alike.
  • The Terminal-Bench column ships empty because no traceable figure met the bar.
  • Traceable figures named a model and version without a harness, or came from the benchmark's own reference agent.
  • Terminal-Bench 1.0 results are superseded by 2.0 and not comparable.
  • Gaps between top entries were smaller than one task on an eighty-task benchmark, so not a ranking.
  • An empty column is the honest state until we run the benchmark against pinned harness versions.
Full methodology and sourcing policy →
Cite this comparisonCC-BY-4.0 · verified 2026-01-15

An agent is a model plus a harness; six reach a pull request, but nothing here merges or deploys unsupervised. — toolweight, https://toolweight.com/compare/coding-agents, verified 2026-01-15. Data from toolweight (https://toolweight.com), licensed CC-BY-4.0.

Frequently asked questions

Does the harness matter more than the model?

  • Below the frontier, no: a weak model fails regardless of harness.
  • At the frontier, yes: context strategy, tool design and failure feedback decide completion rate.
  • The same model in Claude CodeClaude Code logo and in a bare chat window are not the same product.

Which coding agents can open a pull request on their own?

Can any of them merge and deploy without a human?

  • Not out of the box; treat any claim otherwise as a description of your own CI.
  • Copilot's agent is explicitly forbidden from approving its own pull request, and DevinDevin logo stops at review.
  • Claude CodeClaude Code logo and Codex will merge only with a privileged token and approvals disabled, which is you removing the gate.

What is the cheapest way to run a coding agent all day?

  • A flat subscription with soft limits: Claude CodeClaude Code logo on a Max plan, Codex on a ChatGPT Pro plan.
  • Or a free harness with your own key: opencodeopencode logo, AiderAider logo, Goose, ClineCline logo, ZedZed logo.
  • Gemini CLIGemini CLI logo and Qwen CodeQwen Code logo both have usable free tiers tied to a personal account.
  • Avoid credit-based pricing if you need to forecast; the credit-to-token rate is the vendor's to change.

Are Terminal-Bench scores comparable between harnesses?

  • Only when harness version, model and benchmark version are all stated; the column is empty because nothing traceable met that bar.
  • Terminal-Bench 2.0 results are not comparable to 1.0, and a harness can move several points between releases.
  • A figure from the benchmark's own reference agent is a model result, not a harness result.
  • Half a point on an eighty-task benchmark is a fraction of one task, not a ranking.

Is it safe to let an agent run shell commands unattended?

  • Only inside a sandbox with a real filesystem and network boundary.
  • Codex CLICodex CLI logo (Seatbelt on macOS, Landlock on Linux) and Claude CodeClaude Code logo ship OS-level sandboxing.
  • Gemini CLIGemini CLI logo and Goose ship one too, but off by default, so check the cell note.
  • DevinDevin logo and Copilot run in vendor-hosted VMs, which we rank level with a local kernel sandbox.
  • Most IDE extensions offer only an auto-approve allowlist, which is convenience, not a security boundary.

Should my team standardise on one harness?

  • Standardise the repo-level configuration, not the harness.
  • AGENTS.md, MCP server definitions and a documented test command are portable across nearly everything here.
  • Skills, hooks and rules files mostly are not portable.
  • Then swapping harnesses next quarter costs a config change rather than a migration.