A coding agent is a model plus a harness: the prompts, tools and loop around it. To compare harnesses, hold the model still. We gave graff and the Devin CLI the same model, gpt-6.1-sol at medium effort, the same seven coding tasks and the same graders, and ran them side by side. On these tasks the harness alone made a 2.8x difference in time and a 5.7x difference in cost.

The short version
On gpt-6.1-sol, graff v0.0.302.21 took 46.0 seconds per task on average (median 38.1) and $0.025 of list price. The Devin CLI took 128.2 seconds (median 104.7) and $0.142. graff was faster and cheaper on every one of the seven tasks.
Devin passed 19 of 21 runs and graff 18 of 21. The difference is a single run of json-stream, a task no harness we have tested had passed until now. Both passed every run of the other six tasks.
- 46.0sper task for graff on gpt-6.1-sol, against 128.2s for the Devin CLI on the same model
- $0.025of list price per task for graff, against $0.142 for Devin: 5.7 times less
- 5.2model calls per task for graff, against 15.1 for Devin
- 18 / 19runs passed out of 21, graff and Devin; the difference is one run of one task
How to read it. Each dot is one harness on the same model; down and to the left is faster and cheaper. The line is the Pareto frontier, the setups nothing else beats on both at once. graff is the frontier on its own: Devin sits in the region graff beats on both time and cost.
gpt-6.1-sol: time and cost per task
Sources: graff-evals: tasks, checks and runner · Codegraff v0.0.302.21 release notes
How we kept it fair
Same model and effort: gpt-6.1-sol at medium effort on both. Devin runs it as gpt-6-1-sol-medium; graff runs it at its default medium effort, logged in as for a user.
Same tasks and graders: the seven tasks of graff-evals' swe suite, each a small repository with a request and a hidden check the agent never sees. Three runs per task per harness, interleaved task by task in one session, four at a time, so both harnesses saw the same network and service conditions.
Same pricing: cost is computed from each run's own token counts at gpt-6.1-sol's list price ($2 per million input tokens, $0.10 per million cached, $10 per million output), the price both Devin and OpenAI list. Neither harness is charged for the other's discounts or plans.
Each harness as it ships: graff is the released v0.0.302.21 binary running a scripted session. Devin is the Devin CLI 3000.11.3 driven over the Agent Client Protocol, the same way the Harness desktop app drives it, with tools auto-approved. We read Devin's token usage from the per-call usage updates it sends over ACP.
Sources: graff-evals: tasks, checks and runner · Agent Client Protocol · Devin CLI
Where the difference comes from
Devin made 15.1 model calls per task against graff's 5.2. Each call re-reads the conversation, so input grew faster still: 221,000 input tokens per task against 50,700.
Devin also wrote far more: 5,742 output tokens per task against graff's 877. Output tokens are the slow and expensive part of a call. A call's time follows the tokens it writes far more than the context it reads, which is what our last round of work on graff was about.
Prompt caching is not the gap. Devin read 85.2% of its input from the cache and graff 88.6%, close enough that the difference in cost comes almost entirely from volume: more calls, more input per call, and more output.
Sources: Codegraff v0.0.302.21 release notes
Task by task
graff was faster on all seven tasks. The closest was validated, at 80 seconds against 85, where one of graff's three runs took 167 seconds. The widest was cookie-store, 52 seconds against 228.
Cost followed the same pattern: between $0.013 and $0.042 per task for graff, and between $0.11 and $0.21 for Devin.
The one Devin got and graff didn't
json-stream asks for a streaming JSON-sequence reader. Its hidden check has an edge case every harness we had tested missed, so we used to call 18 of 21 a perfect score on this suite. Devin passed it once in three runs. graff failed all three.
One pass in three is a small edge, but a real one. We are reading Devin's passing run to see what it did differently.
What this doesn't show
This is one suite of seven tasks, 21 runs per harness. It measures the harness on short, self-contained coding tasks with automatic checks, not long sessions, large repositories or Devin's cloud features.
The two harnesses reached the model through different services: Devin through its own, graff through a ChatGPT plan. Some of the time difference may come from the route rather than the harness. The token counts, and so the cost, do not depend on the route.
Costs are list-price estimates from each run's token counts, not bills. One run per harness lost its usage report, and those two runs are left out of the cost and token figures but kept in the pass and time figures.
Sources: graff-evals: tasks, checks and runner · Devin CLI
The numbers
gpt-6.1-sol at medium effort, seven tasks from graff-evals' swe suite, three runs each, both harnesses interleaved in one session. Cost is the list price of each run's input, cached input and output tokens.
See the numbersShow ▾Hide ▴
| Harness | Pass | Wall | Median | List cost | Model calls | Input tokens | Cache hit | Output tokens |
|---|---|---|---|---|---|---|---|---|
| graff v0.0.302.21 | 18/21 | 46.0s | 38.1s | $0.0249 | 5.2 | 50,679 | 88.6% | 877 |
| Devin CLI 3000.11.3 | 19/21 | 128.2s | 104.7s | $0.1416 | 15.1 | 221,342 | 85.2% | 5,742 |
Sources: graff-evals: tasks, checks and runner · Codegraff v0.0.302.21 release notes
Try it
graff v0.0.302.21 installs with one line on macOS and Linux, and now natively on Windows with irm https://codegraff.com/install.ps1 | iex. The tasks, checks and runner are in graff-evals, so you can run the same comparison with your own model and harnesses.
Sources: Codegraff v0.0.302.21 release notes · graff-evals: tasks, checks and runner
Sources and method
All numbers come from one interleaved round on October 4, 2026: graff v0.0.302.21 (the released binary) and the Devin CLI 3000.11.3, both on gpt-6.1-sol at medium effort, seven tasks three times each. Devin was driven over the Agent Client Protocol as the Harness desktop app drives it, with token usage read from the per-call usage updates Devin sends; graff's usage comes from its own per-call accounting. Cost is computed from tokens at the same list price for both.