# graff and Devin on the same model: 2.8 times faster, 5.7 times cheaper

We ran graff v0.0.302.21 and the Devin CLI on the same model, gpt-6.1-sol at medium effort, over seven coding tasks with hidden checks, three runs each. graff took 46 seconds and $0.025 of list price per task; Devin took 128 seconds and $0.142. Devin passed one more run, on the one task graff failed every time.

Published: 2026-10-04
Author: Rach Pradhan
Canonical: https://codegraff.com/blog/graff-vs-devin

A coding agent is a model plus a harness: the prompts, tools and loop around it. To compare harnesses, hold the model still. We gave graff and the Devin CLI the same model, gpt-6.1-sol at medium effort, the same seven coding tasks and the same graders, and ran them side by side. On these tasks the harness alone made a 2.8x difference in time and a 5.7x difference in cost.

![Two mouse engineers drive different brass machines: one feeds a long punched program strip into a tall loom, the other sets a short row of levers; both pass task cards into one checking jig.](https://codegraff.com/blog/code-mode/two-machines.webp)

Two machines, one checking jig: the same tasks and graders for both. Conceptual illustration.

## The short version

On gpt-6.1-sol, graff v0.0.302.21 took 46.0 seconds per task on average (median 38.1) and $0.025 of list price. The Devin CLI took 128.2 seconds (median 104.7) and $0.142. graff was faster and cheaper on every one of the seven tasks.

Devin passed 19 of 21 runs and graff 18 of 21. The difference is a single run of json-stream, a task no harness we have tested had passed until now. Both passed every run of the other six tasks.

Sources: [tasks: graff-evals: tasks, checks and runner](https://github.com/justrach/codegraff/tree/main/graff-evals); [notes: Codegraff v0.0.302.21 release notes](https://github.com/justrach/codegraff/blob/v0.0.302.21/docs/releases/v0.0.302.21.md)

## How we kept it fair

Same model and effort: gpt-6.1-sol at medium effort on both. Devin runs it as gpt-6-1-sol-medium; graff runs it at its default medium effort, logged in as for a user.

Same tasks and graders: the seven tasks of graff-evals' swe suite, each a small repository with a request and a hidden check the agent never sees. Three runs per task per harness, interleaved task by task in one session, four at a time, so both harnesses saw the same network and service conditions.

Same pricing: cost is computed from each run's own token counts at gpt-6.1-sol's list price ($2 per million input tokens, $0.10 per million cached, $10 per million output), the price both Devin and OpenAI list. Neither harness is charged for the other's discounts or plans.

Each harness as it ships: graff is the released v0.0.302.21 binary running a scripted session. Devin is the Devin CLI 3000.11.3 driven over the Agent Client Protocol, the same way the Harness desktop app drives it, with tools auto-approved. We read Devin's token usage from the per-call usage updates it sends over ACP.

Sources: [tasks: graff-evals: tasks, checks and runner](https://github.com/justrach/codegraff/tree/main/graff-evals); [acp: Agent Client Protocol](https://agentclientprotocol.com); [devin: Devin CLI](https://cli.devin.ai)

## Where the difference comes from

Devin made 15.1 model calls per task against graff's 5.2. Each call re-reads the conversation, so input grew faster still: 221,000 input tokens per task against 50,700.

Devin also wrote far more: 5,742 output tokens per task against graff's 877. Output tokens are the slow and expensive part of a call. A call's time follows the tokens it writes far more than the context it reads, which is what our last round of work on graff was about.

Prompt caching is not the gap. Devin read 85.2% of its input from the cache and graff 88.6%, close enough that the difference in cost comes almost entirely from volume: more calls, more input per call, and more output.

Sources: [notes: Codegraff v0.0.302.21 release notes](https://github.com/justrach/codegraff/blob/v0.0.302.21/docs/releases/v0.0.302.21.md)

## Task by task

graff was faster on all seven tasks. The closest was validated, at 80 seconds against 85, where one of graff's three runs took 167 seconds. The widest was cookie-store, 52 seconds against 228.

Cost followed the same pattern: between $0.013 and $0.042 per task for graff, and between $0.11 and $0.21 for Devin.

Sources: [tasks: graff-evals: tasks, checks and runner](https://github.com/justrach/codegraff/tree/main/graff-evals)

## The one Devin got and graff didn't

json-stream asks for a streaming JSON-sequence reader. Its hidden check has an edge case every harness we had tested missed, so we used to call 18 of 21 a perfect score on this suite. Devin passed it once in three runs. graff failed all three.

One pass in three is a small edge, but a real one. We are reading Devin's passing run to see what it did differently.

Sources: [tasks: graff-evals: tasks, checks and runner](https://github.com/justrach/codegraff/tree/main/graff-evals)

## What this doesn't show

This is one suite of seven tasks, 21 runs per harness. It measures the harness on short, self-contained coding tasks with automatic checks, not long sessions, large repositories or Devin's cloud features.

The two harnesses reached the model through different services: Devin through its own, graff through a ChatGPT plan. Some of the time difference may come from the route rather than the harness. The token counts, and so the cost, do not depend on the route.

Costs are list-price estimates from each run's token counts, not bills. One run per harness lost its usage report, and those two runs are left out of the cost and token figures but kept in the pass and time figures.

Sources: [tasks: graff-evals: tasks, checks and runner](https://github.com/justrach/codegraff/tree/main/graff-evals); [devin: Devin CLI](https://cli.devin.ai)

## The numbers

gpt-6.1-sol at medium effort, seven tasks from graff-evals' swe suite, three runs each, both harnesses interleaved in one session. Cost is the list price of each run's input, cached input and output tokens.

Per task, mean of 21 runs per harness (cost and tokens over 20 runs with usage)

| Harness | Pass | Wall | Median | List cost | Model calls | Input tokens | Cache hit | Output tokens |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| graff v0.0.302.21 | 18/21 | 46.0s | 38.1s | $0.0249 | 5.2 | 50,679 | 88.6% | 877 |
| Devin CLI 3000.11.3 | 19/21 | 128.2s | 104.7s | $0.1416 | 15.1 | 221,342 | 85.2% | 5,742 |

Sources: [tasks: graff-evals: tasks, checks and runner](https://github.com/justrach/codegraff/tree/main/graff-evals); [notes: Codegraff v0.0.302.21 release notes](https://github.com/justrach/codegraff/blob/v0.0.302.21/docs/releases/v0.0.302.21.md)

## Try it

graff v0.0.302.21 installs with one line on macOS and Linux, and now natively on Windows with irm https://codegraff.com/install.ps1 | iex. The tasks, checks and runner are in graff-evals, so you can run the same comparison with your own model and harnesses.

Sources: [notes: Codegraff v0.0.302.21 release notes](https://github.com/justrach/codegraff/blob/v0.0.302.21/docs/releases/v0.0.302.21.md); [tasks: graff-evals: tasks, checks and runner](https://github.com/justrach/codegraff/tree/main/graff-evals)

## Sources

All numbers come from one interleaved round on October 4, 2026: graff v0.0.302.21 (the released binary) and the Devin CLI 3000.11.3, both on gpt-6.1-sol at medium effort, seven tasks three times each. Devin was driven over the Agent Client Protocol as the Harness desktop app drives it, with token usage read from the per-call usage updates Devin sends; graff's usage comes from its own per-call accounting. Cost is computed from tokens at the same list price for both.

- [tasks: graff-evals: tasks, checks and runner](https://github.com/justrach/codegraff/tree/main/graff-evals)

- [notes: Codegraff v0.0.302.21 release notes](https://github.com/justrach/codegraff/blob/v0.0.302.21/docs/releases/v0.0.302.21.md)

- [acp: Agent Client Protocol](https://agentclientprotocol.com)

- [devin: Devin CLI](https://cli.devin.ai)
