{"schemaVersion":1,"canonical":"https://codegraff.com/blog/graff-shorter-turns","markdown":"https://codegraff.com/blog/graff-shorter-turns/markdown","slug":"graff-shorter-turns","title":"Shorter turns: how graff got ahead of Pi's codemode","seoTitle":"graff v0.0.302.18 vs Pi 1.0 codemode on gpt-6.1-sol and gpt-6-astra: time, cost and cache hits","description":"We hill-climbed graff's default setup against Pi 1.0 with codemode on gpt-6.1-sol and gpt-6-astra. A model call's time follows the tokens it writes, not the context it reads, so the gains came from writing less. v0.0.302.18 takes 19% less time than v0.0.302.17 on gpt-6.1-sol and 14% less on gpt-6-astra. On gpt-6.1-sol it beats Pi with codemode on time, cost and cache hits; on gpt-6-astra the two are level on time and graff is 21% cheaper.","publishedAt":"2026-10-03","keywords":["graff","Pi 1.0","codemode","code mode","gpt-6.1-sol","gpt-6-astra","ChatGPT plan","prompt caching","coding agent benchmark","Pareto frontier","output tokens","latency"],"intro":"Yesterday's post ended with Pi 1.0 and its codemode ahead of graff on gpt-6-astra. We spent the next day hill-climbing graff against it, on gpt-6.1-sol as well, with one rule: the same models, plan and effort for both, and graff as it ships, with no special lean prompt. Five changes made it into v0.0.302.18. None of them is a new feature. Each one removes something graff's turns were paying for and did not need.","hero":{"src":"/blog/graff-0-0-277/cache-discipline-workshop.webp","alt":"Codegraff mice route repeated blue-backed context cards through a fast cache path while a changed card returns to a slower press.","caption":"A stable prefix lets later passes reuse cached work; change it and the shortcut disappears. Conceptual illustration."},"sections":[{"id":"short","title":"The short version","paragraphs":["On gpt-6.1-sol, v0.0.302.18 took 33.0 seconds per task, against 34.2 for Pi with codemode and 40.8 for v0.0.302.17. It cost $0.019 of list price per task against Pi's $0.028, and read 89.8% of its input from the prompt cache against Pi's 48.3%. Nothing beats it on any of the three.","On gpt-6-astra it took 41.4 seconds per task against Pi's 40.7 and v0.0.302.17's 48.2. Its median run was faster than Pi's, 31.7 seconds against 32.9, so call it level on time. It cost $0.106 per task against Pi's $0.135, with 91.2% of its input cached against 56.1%.","Every setup passed all 33 of its runs on both models: 11 tasks, 3 runs each, interleaved task by task, with graff logged in as it is for its users.","The gains were not where we first looked. The size of graff's prompt did not matter. What it wrote did."],"sources":["notes"]},{"id":"fair","title":"Apples to apples","paragraphs":["Both harnesses ran on the same ChatGPT plan, with the same model and the same reasoning effort: graff's default for each model, medium on gpt-6-astra and low on gpt-6.1-sol, which Pi was given with --thinking. Pi ran with codemode in its default tools, which also routes its MCP calls through scripts.","graff ran as it ships. A leaner prompt and tool catalog was faster, but it is not what people run, so it is not in these numbers. Every change below is a change to graff's default behaviour, the same in the line REPL, the fullscreen client and ACP.","graff ran logged in to Codegraff, as a user's copy does, so its tool catalog included jev_effort: the model can ask Jev, a hosted selector, to choose the reasoning effort for its next request. It never asked. Across 132 graff runs on the two models jev_effort was offered every time and called none, so these numbers are graff at its default effort. Our earlier rounds ran without a login, and their numbers were within a second of these.","The tasks: five that pull issues from an MCP server and join them with files or git history, a commit digest, and five multi-file coding tasks (fix three bugs, summarize logs, rename an API, check documentation coverage, build a sidecar). Each round interleaved the setups task by task, 8 runs at a time, so they shared the same network and the same hour."],"sources":["tasks","pi-codemode"]},{"id":"output","title":"A call's time is the tokens it writes","paragraphs":["Across 251 model calls graff made on gpt-6.1-sol, a call's duration fell close to a straight line in its output tokens: about 3.2 seconds before the first token, then about 30 milliseconds for each token written, a correlation of 0.98. The size of the context it read had no relation to the time to the first token (a correlation of -0.09). Most of that context is cached anyway.","graff made fewer calls per task than Pi, 4.1 against 4.7 on gpt-6.1-sol with v0.0.302.17, but each call wrote more: 874 output tokens per task against Pi's 612. At 30 milliseconds a token, that gap alone is about eight seconds a task.","So the question became: what was graff writing that the task did not need?"],"sources":["notes"]},{"id":"less","title":"What graff stopped writing","paragraphs":["Narration. graff's prompt asked the model to say what it had found and what it would do before each step, and to give a one- or two-sentence heads-up before larger chunks of work. Models wrote a short paragraph ahead of most tool calls; Pi asks for none and its models wrote almost none. The heads-up now asks for one short line. Prose ahead of tool calls fell from about 330 characters per task to about 130 on gpt-6.1-sol and about 70 on gpt-6-astra. The transcript still says what is happening, in a line.","Second derivations. \"Make the requested thing work and prove it\" read as license to compute an answer twice: one log-census script counted every file, then counted them again a different way and asserted the two agreed. A script that computes an answer is now its own proof. In the round that measured it, runs with self-checking code fell from 10 of 22 to 2 of 22.","Reads of files that did not exist yet. read_file asked to be called before editing any text file, so models read paths they were about to create. It now asks only for existing files.","rlm round trips. graff's code mode shows a slim view of large MCP results when a script prints them. That view now names the fields the whole value has, so the model does not print a sample row to learn them. The slim rule says to compute inside the same script rather than fetch in one call and compute in the next, and a call that binds no name prints its result instead of nothing.","Version errors. On a Mac whose python3 is the system 3.9, models wrote 3.11 code, such as datetime.fromisoformat with a trailing Z, which failed and cost a call to rewrite. The shell tool now says which python3 the shell runs, read from where python3 on PATH resolves. graff never runs it to find out: on a Mac without the developer tools, that opens an install dialog."],"sources":["adr-0239","adr-0240","adr-0242"]},{"id":"stalls","title":"A stall that was the model thinking","paragraphs":["graff watches a streaming response for silence. Once visible prose arrives, it gives the stream a quarter of the usual budget, on the reasoning that bytes stopping mid-sentence mean a dead socket. But a heads-up is its own output item, and once it closes the model can think in silence while it composes several tool calls. On gpt-6-astra that silence crossed the quarter budget, graff reconnected, and the whole generation was paid for again, about 45 seconds each time.","The quarter budget now holds only while the prose item is open. When the item closes, the next one gets the full budget. A connection that dies mid-sentence is still caught as quickly as before."],"sources":["adr-0241"]},{"id":"cache","title":"The instructions cache as one unit","paragraphs":["graff keys the ChatGPT plan's prompt cache by account, so a new repo's first call can reuse the system prompt another repo warmed. Request dumps showed how much it really reused. The backend renders the tool definitions first and graff's instructions after them, and caches the instructions as one unit: two repos whose instructions differed only in the project layout shared about 6,400 tokens of tool definitions and none of the roughly 4,000 tokens of instructions.","The instructions carried two blocks that change from repo to repo, the project's instruction file and its layout, with static text after them. Those two blocks now ride as their own message ahead of the conversation, and the instructions are identical in every repo. A second repo's first call read 10,112 of its 10,350 input tokens from the cache.","One thing to know if you run evals like these: two prompt variants on one account's cache key evict each other's entries, which made the new prompt look worse at caching than it was. Measured on separate keys, it cached as well as the old one."],"sources":["adr-0243"]},{"id":"not","title":"What did not help","paragraphs":["Dropping the publishing, commit-authoring and issue-filing rules from the prompt saved nothing measurable, which fits the finding that the size of the context does not drive a call's time. Those rules stay; they are safety rules.","Asking for no narration at all was a little faster still than one line, but it leaves the transcript with tool rows and nothing in between. One line is the better trade.","The leaner prompt and catalog mentioned above was faster again, but it is not what people run, so we did not count it."],"sources":["notes"]},{"id":"results","title":"The numbers","paragraphs":["Both models on a ChatGPT plan, 11 tasks, 3 runs each, the three setups interleaved task by task, graff logged in with Jev offered. Cost is the list price of each run's input, cached input and output tokens."],"table":{"caption":"Per task, mean of 33 runs per setup","headers":["Setup","Pass","Wall","Median","List cost","Cache hit","Output tokens","Model calls"],"rows":[["gpt-6.1-sol: v0.0.302.17","33/33","40.8s","37.5s","$0.0259","87.0%","874","4.1"],["gpt-6.1-sol: v0.0.302.18","33/33","33.0s","26.8s","$0.0192","89.8%","701","3.5"],["gpt-6.1-sol: Pi 1.0 + codemode","33/33","34.2s","29.0s","$0.0281","48.3%","612","4.7"],["gpt-6-astra: v0.0.302.17","33/33","48.2s","38.2s","$0.125","88.3%","767","3.8"],["gpt-6-astra: v0.0.302.18","33/33","41.4s","31.7s","$0.106","91.2%","693","3.5"],["gpt-6-astra: Pi 1.0 + codemode","33/33","40.7s","32.9s","$0.135","56.1%","696","4.8"]]},"sources":["notes"]},{"id":"try","title":"Try it","paragraphs":["v0.0.302.18 also moves the fullscreen UI into its own client: graff tui launches an installed graff-tui, and graff in a terminal starts the line REPL. Update with graff update, or install it from the docs."],"sources":["notes"]}],"sourceNote":"The headline numbers come from one round with the shipped v0.0.302.18 binary, logged in as for a user, against v0.0.302.17 and Pi 1.0 with codemode, interleaved on each model. The per-call timing and the narration, self-checking and cache measurements come from the hill-climbing rounds before it, which measured each change on its own. Pi's codemode is described from its own documentation.","sources":[{"id":"notes","title":"Codegraff v0.0.302.18 release notes","url":"https://github.com/justrach/codegraff/blob/v0.0.302.18/docs/releases/v0.0.302.18.md"},{"id":"tasks","title":"graff-evals: tasks, checks and the Pi 1.0 wrapper","url":"https://github.com/justrach/codegraff/tree/main/graff-evals"},{"id":"pi-codemode","title":"Pi: codemode documentation","url":"https://github.com/earendil-works/pi/blob/main/packages/coding-agent/docs/codemode.md"},{"id":"adr-0239","title":"ADR 0239: heads-ups are one line, a computed answer is its own proof","url":"https://github.com/justrach/codegraff/blob/v0.0.302.18/docs/adr/0239-heads-ups-are-one-line-and-a-computed-answer-is-its-own-proof.md"},{"id":"adr-0240","title":"ADR 0240: a slim view names its fields, one script fetches and computes","url":"https://github.com/justrach/codegraff/blob/v0.0.302.18/docs/adr/0240-a-slim-view-names-its-fields-and-one-script-fetches-and-computes.md"},{"id":"adr-0241","title":"ADR 0241: prose tightens the stall budget only until its item closes","url":"https://github.com/justrach/codegraff/blob/v0.0.302.18/docs/adr/0241-prose-tightens-the-stall-budget-only-until-its-item-closes.md"},{"id":"adr-0242","title":"ADR 0242: the shell tool names python3's version","url":"https://github.com/justrach/codegraff/blob/v0.0.302.18/docs/adr/0242-the-shell-tool-names-python3s-version.md"},{"id":"adr-0243","title":"ADR 0243: repo context rides as an input item","url":"https://github.com/justrach/codegraff/blob/v0.0.302.18/docs/adr/0243-repo-context-rides-as-an-input-item-on-responses.md"}]}