# Same model, fewer round trips: graff against the Codex app server and Claude Code

On the same model, graff now finishes our 21 coding and MCP tasks 22% faster than the Codex app server and 34% faster than Claude Code. Claude Code and pi are still quicker on small coding tasks.

Published: 2026-10-01
Author: Rach Pradhan
Canonical: https://codegraff.com/blog/graff-vs-codex-and-claude-code

On the same model and ChatGPT account, graff now finishes our 21 coding and MCP tasks in 22.2 seconds per task against 28.6 for OpenAI's Codex app server: 22% less time, at 69% lower list cost. On one model served through the Codegraff gateway, it takes 8.7 seconds per task against 13.2 for Claude Code and 14.1 for OpenCode 2, and half their time on the MCP tasks. Claude Code and pi are still quicker on the small coding tasks. Every harness passed every run.

![Mouse engineers compare differently shaped brass coding machines that receive matching task cards.](https://codegraff.com/blog/graff-frontier-harness-evals/harness-exhibition.webp)

Different machines, the same task cards. Conceptual illustration.

## The short version

A harness is the program around the model. It decides what the model sees, which tools it can call, and how many round trips a task takes. Give two harnesses the same model and the same task, and the difference you measure is the harness.

We ran 21 tasks: 11 small coding jobs (fix a bug, rename a function, write tests, transform JSON) and 10 jobs against a Linear-shaped MCP server (fetch the issues and their comments, then write a report). Each harness ran each task three times, interleaved with the others, and every run had to pass the task's own check. Every harness passed every run.

On the same model, graff now takes 22% less time per task than the Codex app server and 34% less than Claude Code. Most of the lead is on the MCP tasks, where graff finishes in half the time Claude Code takes.

Sources: [codex-results: graff vs the Codex app server and pi (results and every run)](https://github.com/justrach/codegraff/blob/main/evals/graff-vs-codex-app-server/RESULTS.md); [cc-results: graff vs Claude Code, OpenCode 2 and pi (results and every run)](https://github.com/justrach/codegraff/blob/main/evals/graff-vs-claude-code/RESULTS.md)

## Against the Codex app server

OpenAI's Codex app server is the engine behind the Codex apps. We drove it through its JSON-RPC interface on gpt-6.1-sol, the same model and ChatGPT account graff used, each at the model's default effort.

graff took 22.2 seconds per task against 28.6: 16.0 against 20.4 on the coding tasks, and 29.0 against 37.8 on the MCP tasks. It made 3.1 model calls per task to the app server's 4.4, and at list price it cost $0.013 per task against $0.043.

That is a reversal. In the first version of this comparison, earlier the same day, graff was 13% slower than the app server, and the v0.0.302.12 release tied it in the afternoon. The changes below put graff 18% under its own release.

pi, a small open-source harness, ran the coding tasks on the same account; it has no MCP client. It took 17.0 seconds per task against graff's 16.0, and costs less there because its prompt is about a tenth the size of graff's.

gpt-6.1-sol, one ChatGPT account. Mean per task over 3 runs; list cost at the model's public rates.

| Harness | All tasks | Coding | MCP | Calls per task | List cost per task |
| --- | --- | --- | --- | --- | --- |
| graff | 22.2s | 16.0s | 29.0s | 3.1 | $0.013 |
| Codex app server | 28.6s | 20.4s | 37.8s | 4.4 | $0.043 |
| pi | — | 17.0s | no MCP client | 3.2 (coding) | $0.007 (coding) |

Sources: [codex-results: graff vs the Codex app server and pi (results and every run)](https://github.com/justrach/codegraff/blob/main/evals/graff-vs-codex-app-server/RESULTS.md)

## Against Claude Code, OpenCode 2 and pi

For Anthropic's Claude Code we used one model served through the Codegraff gateway, with the same model id and key for every harness. The gateway takes OpenAI-style chat completions and Claude Code speaks only Anthropic's Messages API, so Claude Code reached it through a local LiteLLM proxy. The proxy added no measurable time, but it does cost Claude Code some prompt caching; more on that in the method.

graff took 8.7 seconds per task, Claude Code 13.2 and OpenCode 2 14.1. On the MCP tasks graff took 10.6 seconds against 21.2 and 22.1. Claude Code keeps every MCP result whole in its context: 179k prompt tokens per MCP task against graff's 44k, and nearly twice the output.

On the coding tasks the order flips. Claude Code and pi both took 6.0 seconds per task, OpenCode 2 6.8, and graff 7.0, about 0.3 seconds of which is one slow gateway response. Small tasks on this model are where graff has the least to give.

One gateway model, same id and key for every harness. Mean per task over 3 runs.

| Harness | All tasks | Coding | MCP | Calls per task | Prompt tokens per MCP task |
| --- | --- | --- | --- | --- | --- |
| graff | 8.7s | 7.0s | 10.6s | 3.5 | 44k |
| Claude Code | 13.2s | 6.0s | 21.2s | 3.5 | 179k |
| OpenCode 2 | 14.1s | 6.8s | 22.1s | 4.3 | 63k |
| pi | — | 6.0s | no MCP client | 2.8 (coding) | — |

Sources: [cc-results: graff vs Claude Code, OpenCode 2 and pi (results and every run)](https://github.com/justrach/codegraff/blob/main/evals/graff-vs-claude-code/RESULTS.md)

## What changed in graff

Each of these removes a model call that did no work the task asked for, and every call costs seconds.

Servers you name load up front. When your message names an MCP server ("the linear server"), graff loads its tools with the first request instead of spending a round trip on a load call, and asking for subagents does the same for the subagent tools. The note that comes with the tools explains how large results are trimmed, and that one rlm script can fetch, process and write a file in a single step. On half of the MCP tasks graff now finishes in two model calls, where the release needed three to five.

write_file is safer, and says what it did. It will not replace a file the session has not read unless the call says replace: true, so the model no longer checks whether a path is free before writing. Its result says whether it created or replaced the file and whether a .json file parses, so the model stops reading back what it just wrote. Close to half of the runs that wrote a file used to spend a call on one of those checks.

Subagents can use their parent's MCP tools at once. They used to start with a tool search. The task that fans out to two subagents went from about 65 seconds to about 50.

Unattended runs skip what no one reads. Under -p or in a piped session, graff no longer offers a question tool nobody can answer, and at low effort it skips reasoning summaries that only a screen would show. A piped graff repl also reads a pasted multi-line block as one prompt.

Two ideas did not survive measurement: letting the final write and the "done" signal share one response (the model never did it), and telling the model that a batch of tool calls runs in order (it started splitting edits apart instead).

Sources: [adr-0231: ADR 0231: no round trip the model does not need](https://github.com/justrach/codegraff/blob/main/docs/adr/0231-no-round-trip-the-model-does-not-need.md); [release: graff v0.0.302.13](https://github.com/justrach/codegraff/releases/tag/v0.0.302.13)

## How we ran it

Every task runs in a fresh git repository with a fresh home directory, so no harness sees user-level settings or MCP servers. Harnesses run interleaved, task by task, a few at a time, each after one unscored warm-up request. Wall time is the runner's clock around each process, from start to exit.

Costs are list prices of the tokens each harness reports: $2 per million input tokens, $0.10 per million cached and $10 per million output for gpt-6.1-sol. A ChatGPT plan is not billed per token; the dollars make the harnesses comparable, not a bill. The Codex app server sends one warm-up request per thread that its usage does not report, so its cost here is a floor.

Through the proxy, Claude Code's own prompt-cache markers do not survive the translation, and only 49% of its MCP input was read from cache, against graff's 96%. Against an endpoint that takes Anthropic's API directly, Claude Code's MCP times would likely be better, so treat its MCP gap here as an upper bound.

OpenCode 2 could not be signed into a ChatGPT plan without an interactive login, so it appears only in the gateway comparison. Three runs per task on one machine means a second or two on any single task is noise; the suite-level gaps held across every run that day.

Sources: [codex-results: graff vs the Codex app server and pi (results and every run)](https://github.com/justrach/codegraff/blob/main/evals/graff-vs-codex-app-server/RESULTS.md); [cc-results: graff vs Claude Code, OpenCode 2 and pi (results and every run)](https://github.com/justrach/codegraff/blob/main/evals/graff-vs-claude-code/RESULTS.md); [graff-evals: graff-evals: tasks, harness drivers and runner](https://github.com/justrach/codegraff/tree/main/graff-evals)

## Try it

These changes ship in graff v0.0.302.13. The tasks, the harness drivers and every run's data are in the codegraff repo, so you can rerun the comparison with graff-evals on your own tasks and models.

Sources: [release: graff v0.0.302.13](https://github.com/justrach/codegraff/releases/tag/v0.0.302.13); [graff-evals: graff-evals: tasks, harness drivers and runner](https://github.com/justrach/codegraff/tree/main/graff-evals)

## Sources

Every number comes from the public eval write-ups in the codegraff repo: 21 tasks, 3 runs of each per harness, arms interleaved task by task, each harness at its own defaults. The Codex comparison ran gpt-6.1-sol on one ChatGPT account. The gateway comparison ran one model through the Codegraff gateway, with the same model id and key for every harness; Claude Code reached it through a local proxy because the gateway takes chat completions only. List costs use gpt-6.1-sol's public rates; the gateway comparison reports tokens, not dollars.

- [codex-results: graff vs the Codex app server and pi (results and every run)](https://github.com/justrach/codegraff/blob/main/evals/graff-vs-codex-app-server/RESULTS.md)

- [cc-results: graff vs Claude Code, OpenCode 2 and pi (results and every run)](https://github.com/justrach/codegraff/blob/main/evals/graff-vs-claude-code/RESULTS.md)

- [adr-0231: ADR 0231: no round trip the model does not need](https://github.com/justrach/codegraff/blob/main/docs/adr/0231-no-round-trip-the-model-does-not-need.md)

- [release: graff v0.0.302.13](https://github.com/justrach/codegraff/releases/tag/v0.0.302.13)

- [graff-evals: graff-evals: tasks, harness drivers and runner](https://github.com/justrach/codegraff/tree/main/graff-evals)
