Graff / Evaluations

Shorter turns: how graff got ahead of Pi's codemode

We hill-climbed graff's default setup against Pi 1.0 with codemode on gpt-6.1-sol and gpt-6-astra. A model call's time follows the tokens it writes, not the context it reads, so the gains came from writing less. v0.0.302.18 takes 19% less time than v0.0.302.17 on gpt-6.1-sol and 14% less on gpt-6-astra. On gpt-6.1-sol it beats Pi with codemode on time, cost and cache hits; on gpt-6-astra the two are level on time and graff is 21% cheaper.

Read as MarkdownStructured article dataRelease notes

Yesterday's post ended with Pi 1.0 and its codemode ahead of graff on gpt-6-astra. We spent the next day hill-climbing graff against it, on gpt-6.1-sol as well, with one rule: the same models, plan and effort for both, and graff as it ships, with no special lean prompt. Five changes made it into v0.0.302.18. None of them is a new feature. Each one removes something graff's turns were paying for and did not need.

Codegraff mice route repeated blue-backed context cards through a fast cache path while a changed card returns to a slower press.
A stable prefix lets later passes reuse cached work; change it and the shortcut disappears. Conceptual illustration.

The short version

On gpt-6.1-sol, v0.0.302.18 took 33.0 seconds per task, against 34.2 for Pi with codemode and 40.8 for v0.0.302.17. It cost $0.019 of list price per task against Pi's $0.028, and read 89.8% of its input from the prompt cache against Pi's 48.3%. Nothing beats it on any of the three.

On gpt-6-astra it took 41.4 seconds per task against Pi's 40.7 and v0.0.302.17's 48.2. Its median run was faster than Pi's, 31.7 seconds against 32.9, so call it level on time. It cost $0.106 per task against Pi's $0.135, with 91.2% of its input cached against 56.1%.

Every setup passed all 33 of its runs on both models: 11 tasks, 3 runs each, interleaved task by task, with graff logged in as it is for its users.

The gains were not where we first looked. The size of graff's prompt did not matter. What it wrote did.

  • 33.0sper task for graff v0.0.302.18 on gpt-6.1-sol, against 34.2s for Pi 1.0 with codemode
  • 19%less time per task than v0.0.302.17 on gpt-6.1-sol, and 14% on gpt-6-astra
  • 32%lower list cost per task than Pi with codemode on gpt-6.1-sol, 21% on gpt-6-astra
  • 0.98the correlation between a call's duration and the tokens it writes; context size had none

How to read it. Each dot is one setup (Pi is Pi 1.0 with codemode); down and to the left is faster and cheaper. The line is the Pareto frontier, the setups nothing else beats on both at once. On gpt-6.1-sol, v0.0.302.18 is the frontier on its own: faster and cheaper than Pi with codemode and than v0.0.302.17.

gpt-6.1-sol: time and cost per task

gpt-6.1-sol on a ChatGPT plan, graff logged in with Jev offered: 33 runs per setup, 11 tasks 3 times each, interleaved in one session. Every run passed.

How to read it. On gpt-6-astra Pi with codemode is 0.7 seconds faster on the mean, though its median run is slower, and v0.0.302.18 is 21% cheaper because most of its input is read from the cache. Both sit on the frontier; v0.0.302.17 is beaten on both.

gpt-6-astra: time and cost per task

gpt-6-astra on a ChatGPT plan, graff logged in with Jev offered: 33 runs per setup, 11 tasks 3 times each, interleaved in one session. Every run passed.

Sources: Codegraff v0.0.302.18 release notes

Apples to apples

Both harnesses ran on the same ChatGPT plan, with the same model and the same reasoning effort: graff's default for each model, medium on gpt-6-astra and low on gpt-6.1-sol, which Pi was given with --thinking. Pi ran with codemode in its default tools, which also routes its MCP calls through scripts.

graff ran as it ships. A leaner prompt and tool catalog was faster, but it is not what people run, so it is not in these numbers. Every change below is a change to graff's default behaviour, the same in the line REPL, the fullscreen client and ACP.

graff ran logged in to Codegraff, as a user's copy does, so its tool catalog included jev_effort: the model can ask Jev, a hosted selector, to choose the reasoning effort for its next request. It never asked. Across 132 graff runs on the two models jev_effort was offered every time and called none, so these numbers are graff at its default effort. Our earlier rounds ran without a login, and their numbers were within a second of these.

The tasks: five that pull issues from an MCP server and join them with files or git history, a commit digest, and five multi-file coding tasks (fix three bugs, summarize logs, rename an API, check documentation coverage, build a sidecar). Each round interleaved the setups task by task, 8 runs at a time, so they shared the same network and the same hour.

Sources: graff-evals: tasks, checks and the Pi 1.0 wrapper · Pi: codemode documentation

A call's time is the tokens it writes

Across 251 model calls graff made on gpt-6.1-sol, a call's duration fell close to a straight line in its output tokens: about 3.2 seconds before the first token, then about 30 milliseconds for each token written, a correlation of 0.98. The size of the context it read had no relation to the time to the first token (a correlation of -0.09). Most of that context is cached anyway.

graff made fewer calls per task than Pi, 4.1 against 4.7 on gpt-6.1-sol with v0.0.302.17, but each call wrote more: 874 output tokens per task against Pi's 612. At 30 milliseconds a token, that gap alone is about eight seconds a task.

So the question became: what was graff writing that the task did not need?

Mean output tokens per task over 33 runs (11 tasks, 3 runs each), reasoning included. A call's time grows by about 30 milliseconds per output token on gpt-6.1-sol and 39 on gpt-6-astra. Shorter is faster.

Sources: Codegraff v0.0.302.18 release notes

What graff stopped writing

Narration. graff's prompt asked the model to say what it had found and what it would do before each step, and to give a one- or two-sentence heads-up before larger chunks of work. Models wrote a short paragraph ahead of most tool calls; Pi asks for none and its models wrote almost none. The heads-up now asks for one short line. Prose ahead of tool calls fell from about 330 characters per task to about 130 on gpt-6.1-sol and about 70 on gpt-6-astra. The transcript still says what is happening, in a line.

Second derivations. "Make the requested thing work and prove it" read as license to compute an answer twice: one log-census script counted every file, then counted them again a different way and asserted the two agreed. A script that computes an answer is now its own proof. In the round that measured it, runs with self-checking code fell from 10 of 22 to 2 of 22.

Reads of files that did not exist yet. read_file asked to be called before editing any text file, so models read paths they were about to create. It now asks only for existing files.

rlm round trips. graff's code mode shows a slim view of large MCP results when a script prints them. That view now names the fields the whole value has, so the model does not print a sample row to learn them. The slim rule says to compute inside the same script rather than fetch in one call and compute in the next, and a call that binds no name prints its result instead of nothing.

Version errors. On a Mac whose python3 is the system 3.9, models wrote 3.11 code, such as datetime.fromisoformat with a trailing Z, which failed and cost a call to rewrite. The shell tool now says which python3 the shell runs, read from where python3 on PATH resolves. graff never runs it to find out: on a Mac without the developer tools, that opens an install dialog.

Narration the model wrote in the same responses as its tool calls, not counting the final answer. Mean per task over 33 runs.

Sources: ADR 0239: heads-ups are one line, a computed answer is its own proof · ADR 0240: a slim view names its fields, one script fetches and computes · ADR 0242: the shell tool names python3's version

A stall that was the model thinking

graff watches a streaming response for silence. Once visible prose arrives, it gives the stream a quarter of the usual budget, on the reasoning that bytes stopping mid-sentence mean a dead socket. But a heads-up is its own output item, and once it closes the model can think in silence while it composes several tool calls. On gpt-6-astra that silence crossed the quarter budget, graff reconnected, and the whole generation was paid for again, about 45 seconds each time.

The quarter budget now holds only while the prose item is open. When the item closes, the next one gets the full budget. A connection that dies mid-sentence is still caught as quickly as before.

Sources: ADR 0241: prose tightens the stall budget only until its item closes

The instructions cache as one unit

graff keys the ChatGPT plan's prompt cache by account, so a new repo's first call can reuse the system prompt another repo warmed. Request dumps showed how much it really reused. The backend renders the tool definitions first and graff's instructions after them, and caches the instructions as one unit: two repos whose instructions differed only in the project layout shared about 6,400 tokens of tool definitions and none of the roughly 4,000 tokens of instructions.

The instructions carried two blocks that change from repo to repo, the project's instruction file and its layout, with static text after them. Those two blocks now ride as their own message ahead of the conversation, and the instructions are identical in every repo. A second repo's first call read 10,112 of its 10,350 input tokens from the cache.

One thing to know if you run evals like these: two prompt variants on one account's cache key evict each other's entries, which made the new prompt look worse at caching than it was. Measured on separate keys, it cached as well as the old one.

Cached input tokens over all input tokens, over 33 runs per setup. Cached input is billed at a tenth (gpt-6-astra) to a twentieth (gpt-6.1-sol) of the uncached price. Longer is better.

Sources: ADR 0243: repo context rides as an input item

What did not help

Dropping the publishing, commit-authoring and issue-filing rules from the prompt saved nothing measurable, which fits the finding that the size of the context does not drive a call's time. Those rules stay; they are safety rules.

Asking for no narration at all was a little faster still than one line, but it leaves the transcript with tool rows and nothing in between. One line is the better trade.

The leaner prompt and catalog mentioned above was faster again, but it is not what people run, so we did not count it.

Sources: Codegraff v0.0.302.18 release notes

The numbers

Both models on a ChatGPT plan, 11 tasks, 3 runs each, the three setups interleaved task by task, graff logged in with Jev offered. Cost is the list price of each run's input, cached input and output tokens.

See the numbersShow ▾
Per task, mean of 33 runs per setup
SetupPassWallMedianList costCache hitOutput tokensModel calls
gpt-6.1-sol: v0.0.302.1733/3340.8s37.5s$0.025987.0%8744.1
gpt-6.1-sol: v0.0.302.1833/3333.0s26.8s$0.019289.8%7013.5
gpt-6.1-sol: Pi 1.0 + codemode33/3334.2s29.0s$0.028148.3%6124.7
gpt-6-astra: v0.0.302.1733/3348.2s38.2s$0.12588.3%7673.8
gpt-6-astra: v0.0.302.1833/3341.4s31.7s$0.10691.2%6933.5
gpt-6-astra: Pi 1.0 + codemode33/3340.7s32.9s$0.13556.1%6964.8

Sources: Codegraff v0.0.302.18 release notes

Try it

v0.0.302.18 also moves the fullscreen UI into its own client: graff tui launches an installed graff-tui, and graff in a terminal starts the line REPL. Update with graff update, or install it from the docs.

Sources: Codegraff v0.0.302.18 release notes

Sources and method

The headline numbers come from one round with the shipped v0.0.302.18 binary, logged in as for a user, against v0.0.302.17 and Pi 1.0 with codemode, interleaved on each model. The per-call timing and the narration, self-checking and cache measurements come from the hill-climbing rounds before it, which measured each change on its own. Pi's codemode is described from its own documentation.

  1. Codegraff v0.0.302.18 release notes
  2. graff-evals: tasks, checks and the Pi 1.0 wrapper
  3. Pi: codemode documentation
  4. ADR 0239: heads-ups are one line, a computed answer is its own proof
  5. ADR 0240: a slim view names its fields, one script fetches and computes
  6. ADR 0241: prose tightens the stall budget only until its item closes
  7. ADR 0242: the shell tool names python3's version
  8. ADR 0243: repo context rides as an input item