{"schemaVersion":1,"canonical":"https://codegraff.com/blog/graff-rlm-vs-pi-codemode","markdown":"https://codegraff.com/blog/graff-rlm-vs-pi-codemode/markdown","slug":"graff-rlm-vs-pi-codemode","title":"What code mode is worth: graff's rlm and Pi 1.0's codemode on two models","seoTitle":"graff rlm vs Pi 1.0 codemode: a code-mode benchmark on MiMo v2.6 Pro and gpt-6-astra","description":"Two harnesses, two models, eleven harder tasks, each code mode on and off. On MiMo v2.6 Pro neither code mode saved time; on gpt-6-astra both did, Pi's the most. graff's rlm was held back because its scripts never saw whole results; graff v0.0.302.17 changes that, and its time on the tasks that join data fell by 27% on gpt-6-astra.","publishedAt":"2026-10-02","keywords":["graff","Pi coding agent","codemode","code mode","rlm","MCP","MiMo v2.6 Pro","gpt-6-astra","coding agent benchmark","codegraff"],"intro":"Pi 1.0 shipped codemode: instead of calling tools one at a time, the model writes a JavaScript script that calls them, and only what the script prints comes back. graff has its own code mode, rlm, a much smaller script language. To see what each is worth, we ran both harnesses with their code modes on and off, on eleven tasks harder than our usual suite and on two models: MiMo v2.6 Pro through one endpoint and key, and gpt-6-astra on a ChatGPT plan.","hero":{"src":"/blog/code-mode/two-machines.webp","alt":"Two mouse engineers drive different brass machines: one feeds a long punched program strip into a tall loom, the other sets a short row of levers; both pass task cards into one checking jig.","caption":"One machine reads a long program, the other a few levers, and both feed the same checking jig. Conceptual illustration."},"sections":[{"id":"short","title":"The short version","paragraphs":["Code mode is not free. What it is worth depends on the model and on the task: it pays when a task makes many calls and needs little of each result, and it costs when the model has to read the data anyway.","On MiMo v2.6 Pro neither code mode saved time. graff was quickest with rlm switched off, 55.1 seconds per task against 81.5 with it on, as v0.0.302.16 ships. Pi was quickest on the MCP tasks with its MCP tools declared as ordinary tools, 30.5 seconds per task against 34.3 and 37.8 through codemode, though codemode read about a third fewer input tokens.","On gpt-6-astra both code modes paid off, and every setup passed all 33 runs. Pi with codemode in its default tools was the fastest setup at 41.7 seconds per task, against 48.1 for Pi as installed; astra also used codemode for local work, which MiMo never did. graff took 54.1 seconds with rlm and 57.3 without.","graff's rlm was held back by a choice we made to save tokens: an MCP result was trimmed to ids and titles before an rlm script saw it, so a task that needed a priority or an estimate had to leave the script. graff v0.0.302.17 changes that: a script now keeps the whole result, and only what it prints is trimmed. On the three tasks that join data, graff's time fell from 51.4 to 37.7 seconds on astra, close to Pi with codemode's 34.8, and from 39.9 to 34.1 on MiMo, ahead of Pi's 40.6."],"sources":["results","pi-codemode","adr-0238"]},{"id":"designs","title":"Two ways to let the model write the glue","paragraphs":["Pi's codemode is one tool that takes JavaScript. The script runs in a QuickJS sandbox, every tool is an async function on a tools object, and only what the script logs reaches the model. A script can await tools.mcp__linear__list_issues({}), keep the fields it needs and print a summary.","In Pi 1.0 codemode is also how MCP tools are reached by default: they can be called from a script but are not declared to the model, nor listed in codemode's description, so a script finds them with searchTools() or describeTool() first, and Pi switches codemode on whenever an MCP server connects. Adding codemode to Pi's default tools makes it available on every task, for reading files and running commands too.","graff's rlm is a call-only script: issues = mcp__linear__list_issues(), then comments = each(issues, mcp__linear__list_comments, \"id\"), then print(len(issues), project(comments, \"latest_author\")). It has no expressions, loops or indexing; anything that needs computing goes to write_file and a shell script. graff starts the calls while the model is still writing the script, and names persist from one script to the next.","The bets differ. rlm bets that a small language the model cannot get badly wrong, plus trimming big results before anyone reads them, keeps the context small. Codemode bets that models already write good JavaScript, so give them the real thing and keep the data inside the sandbox."],"sources":["pi-codemode","pi-mcp","adr-0029"]},{"id":"tasks","title":"Harder tasks","paragraphs":["Our usual MCP tasks ask for issue ids, titles, comment counts and each issue's latest commenter, which is exactly what graff's trimming keeps. Three new tasks need more. Cross-reference: for each issue key, count the files under src/ that mention it, where ENG-1010 is not ENG-101, a lowercase key does not count and a file counts once. Roll-up: group the issues by priority and add up their estimates and comments. Stale triage: find the priority 1 and 2 issues that no commit authored since a given date mentions, in the subject or the body.","The other eight: the two report tasks from our usual suite; a commit digest that counts, per top-level directory, the commits touching it in a pinned slice of this repository's history, in one script so the raw log never reaches the model; and five multi-package tasks from our sub-agent suite, asked without delegation: fix three broken packages, count log lines across plain and gzipped files, rename an API across three packages, compare the routes in code with the API docs, and fix a config parser while documenting its format.","Every task is graded by its own script against the workspace and the fixture, never by a model."],"sources":["results","tasks"]},{"id":"mimo","title":"MiMo v2.6 Pro: plain tool calls won","paragraphs":["With 33 runs per setup, graff with rlm off was the fastest: 55.1 seconds per task, 30 runs passed. Pi with codemode in its default tools took 72.0 seconds and passed 31, Pi as installed 92.8 and 32, and graff v0.0.302.16 with rlm on 81.5 and 27.","On the report tasks, graff with rlm off asked for all eight issues' comments in one turn of parallel calls, and the trimmed results were small: 27.4 seconds per task. On the join tasks the two Pi setups and graff with rlm off all took 47 to 49 seconds, and graff with rlm on 87.5. On the multi-package tasks graff with rlm off took 72.7 seconds and Pi with codemode 101.5, but Pi passed 14 of 15 and graff 13 and 11.","MiMo never called codemode on the six tasks without an MCP server, in 17 runs, so Pi's two setups differ mainly in noise there; on the MCP tasks both go through codemode. To measure codemode itself we ran a second Pi session on the five MCP tasks with the MCP tools declared directly. Without codemode Pi was quicker: 30.5 seconds per task against 34.3 and 37.8, on the report tasks (31.6 against 33.9 and 38.3) and on the join tasks (29.8 against 34.5 and 37.4). What codemode saved was input: on the report tasks it read 34.8k and 41.4k tokens per task against 77.0k, because the script kept the bulky comment lists out of the conversation."],"table":{"caption":"MiMo v2.6 Pro, one model id, endpoint and key. Mean per task over 33 runs: 11 tasks, 3 runs each, interleaved in one session.","headers":["Setup","Runs passed","Seconds per task","Model calls per task","Input tokens per task","Output tokens per task"],"rows":[["graff v0.0.302.16 (rlm on)","27/33","81.5s","8.3","75.9k","1,505"],["graff, rlm off","30/33","55.1s","8.3","79.1k","1,807"],["Pi 1.0 + codemode","31/33","72.0s","10.3","91.1k","2,158"],["Pi 1.0","32/33","92.8s","12.4","110.7k","2,753"]]},"sources":["results"]},{"id":"astra","title":"gpt-6-astra: code mode paid off","paragraphs":["On gpt-6-astra, through a ChatGPT plan, every setup passed all 33 runs, so the differences are time and tokens. Pi with codemode in its default tools took 41.7 seconds per task, Pi as installed 48.1, graff v0.0.302.17, which carries the rlm change, 48.3, graff v0.0.302.16 54.1 and graff with rlm off 57.3.","rlm helped graff here, a little: 54.1 against 57.3 seconds. astra wrote scripts graff accepted every time, where MiMo hit a refusal in 5 of the 13 runs that tried.","Codemode helped Pi most on the multi-package tasks, 52.1 against 65.3 seconds: astra used codemode in 8 of the 12 runs without an MCP server, reading files and running commands from one script.","A second astra session separated codemode from Pi's other defaults on the five MCP tasks. On the report tasks codemode was clearly better: 32.3 and 34.2 seconds per task against 55.8 with the MCP tools declared directly, and less than half the input tokens. On the join tasks the direct setup was quicker, 30.5 seconds against 36.5 and 47.1, and read fewer tokens too.","graff read about twice the input Pi did, 44.0k to 52.2k tokens per task against 20.9k to 21.9k. Its first request is about 10,500 tokens of prompt and tool definitions, Pi's about 1,600. graff's prompt cache covered 89% of its input against Pi's 54% to 57%, which narrows the gap in cost but not in time."],"table":{"caption":"gpt-6-astra on a ChatGPT plan. Mean per task over 33 runs: 11 tasks, 3 runs each, interleaved in one session. Every setup passed all 33.","headers":["Setup","Seconds per task","Model calls per task","Input tokens per task","Read from cache","Output tokens per task"],"rows":[["graff v0.0.302.16 (rlm on)","54.1s","4.5","49.4k","89%","788"],["graff, rlm off","57.3s","4.6","52.2k","89%","927"],["graff v0.0.302.17","48.3s","3.8","44.0k","89%","835"],["Pi 1.0 + codemode","41.7s","4.8","20.9k","54%","647"],["Pi 1.0","48.1s","5.2","21.9k","57%","704"]]},"sources":["results"]},{"id":"out-of-the-box","title":"Where code mode works out of the box","paragraphs":["One pattern held on both models. Code mode pays when a task makes many calls and needs little of each result, like counting the comments on eight issues: one script fetches them all and returns the counts. When the model has to read the data to decide what to do next, as in the join tasks, calling the tools directly was as fast or faster, because a script costs a round trip of its own and Pi's model first has to find the tools.","Pi's model went through codemode on every MCP run, on both models, and fetched and transformed the data in one script in most of them: 28 of 30 runs on MiMo and 18 of 30 on astra. The cost is finding the tools. They are not listed, so a run spent about two calls on searchTools() and describeTool() before its first real fetch. On astra, codemode also took over local work; on MiMo it never did.","graff's rlm worked out of the box for astra and not for MiMo. MiMo used rlm in 13 of 33 runs and hit a refusal in 5: it wrote Python into the call-only language, with loops, comprehensions and calls with several positional arguments, and each refusal is a round trip. When the trimmed result lacked a field it paged the stored result back in; one run spent its whole five minutes looking for the full data, down to spying on the MCP server's raw output.","Two costs on MiMo were not code mode at all. On the route-coverage task graff's model wrote the answer file by hand in all six runs and got five wrong, unsorted or missing an endpoint; Pi's model computed it with a script in all six and got every one right. And graff's prompt does not name the working directory, so in six runs MiMo invented one, and its first command failed. astra did neither."],"sources":["results"]},{"id":"change","title":"What we changed in rlm","paragraphs":["In v0.0.302.16 an MCP call inside an rlm script bound the trimmed cut: rows with only id, identifier, title and name, and comment lists folded to a count and the latest author. The trim was built for our report tasks and saved tokens there, but a script that needed anything else could not reach it. The model had to leave the script, call the tool directly and page the stored result back in 16 KB slices, each ending in a byte-range marker that broke the JSON when the slices were joined.","Now a bind keeps the tool's whole result, which never reaches the model unless the script prints it. print() still shows the trimmed view, and says it is one. project(x, field) reads any field, so project(issues, \"priority\") works, while n and latest_author still read the comment fold. write_file saves the whole result for a shell script, and inside a script read_tool_result binds a stored result whole.","On gpt-6-astra, in the same session as the rest, v0.0.302.17 took graff from 54.1 to 48.3 seconds per task, and on the join tasks from 51.4 to 37.7, with 3.8 model calls per task against 4.5. No run paged a stored result back in, against 6 of 33 before.","On MiMo we measured it in a second session on the five MCP tasks, with v0.0.302.17's code: 31.0 seconds per task against 37.1 for v0.0.302.16. On the join tasks it was the fastest of the four setups, 34.1 seconds against 40.6 for Pi with codemode and 51.7 for graff with rlm off. On the report tasks rlm off stayed quickest, 18.3 seconds: parallel calls with trimmed results beat writing a script.","The change ships in graff v0.0.302.17. Still open: graff's prompt is several times the size of Pi's, so it reads more input per task, and on MiMo plain tool calls remain faster for simple fan-out work. If you run graff on MiMo today, --no-rlm (or GRAFF_RLM=0) is the setting to try."],"table":{"caption":"Seconds per task on the MCP tasks. gpt-6-astra: the main session, 3 runs of each task. MiMo v2.6 Pro: a second session on the five MCP tasks, 3 runs each, the four setups interleaved.","headers":["Setup","astra, join tasks","astra, report tasks","MiMo, join tasks","MiMo, report tasks"],"rows":[["graff v0.0.302.17","37.7s","31.4s","34.1s","26.2s"],["graff v0.0.302.16","51.4s","31.5s","39.9s","32.9s"],["graff, rlm off","55.8s","38.5s","51.7s","18.3s"],["Pi 1.0 + codemode","34.8s","30.5s","40.6s","37.4s"]]},"sources":["adr-0238","adr-0225","results"]},{"id":"method","title":"How we ran it","paragraphs":["Every task runs in a fresh git repository, and each harness reads only the task's MCP server from a private configuration. Setups run interleaved, task by task, eight runs at a time. Wall time is the runner's clock around each process. Model calls and tokens come from each harness's own usage report, and input tokens include cached ones.","graff ran as the published v0.0.302.16 build, as a scripted graff repl at its default effort: on MiMo that sends thinking off, on astra medium reasoning. rlm off is the same build with GRAFF_RLM=0, and the v0.0.302.17 arm ran that build with the rlm change, the same code v0.0.302.17 ships. Pi ran as 1.0.0, headless, with the same thinking setting: off on MiMo, medium on astra. Its codemode setup adds codemode to the default tools through its settings rather than --tools, which would drop the MCP tools; its no-codemode setup declares the MCP tools directly.","On astra both harnesses used a copy of one ChatGPT sign-in with its refresh token removed, so no run could rotate the real login. Three runs per task is a small sample, and response times swing over a day, so we only compare setups that ran in the same session."],"sources":["results","tasks"]},{"id":"try","title":"Try it","paragraphs":["The tasks, the Pi 1.0 wrapper and every run's numbers are in the codegraff repo, so you can rerun the comparison with graff-evals on your own models. graff update installs the latest release."],"sources":["results","tasks"]}],"sourceNote":"Every number comes from the public eval write-up in the codegraff repo: 11 tasks, 3 runs each per setup, setups interleaved task by task. MiMo v2.6 Pro ran with one model id, endpoint and key; gpt-6-astra ran on a ChatGPT plan. Extra sessions on the five MCP tasks measured Pi without codemode on both models and the rlm change on MiMo. Pi's codemode is described from its own documentation.","sources":[{"id":"results","title":"graff and Pi 1.0 with and without their code modes (results per task)","url":"https://github.com/justrach/codegraff/blob/main/evals/graff-vs-pi-codemode/RESULTS.md"},{"id":"tasks","title":"graff-evals: tasks, checks and the Pi 1.0 wrapper","url":"https://github.com/justrach/codegraff/tree/main/graff-evals"},{"id":"pi-codemode","title":"Pi: codemode documentation","url":"https://github.com/earendil-works/pi/blob/main/packages/coding-agent/docs/codemode.md"},{"id":"pi-mcp","title":"Pi: MCP tool exposure","url":"https://github.com/earendil-works/pi/blob/main/packages/coding-agent/docs/mcp.md"},{"id":"adr-0238","title":"ADR 0238: rlm binds keep the whole MCP result","url":"https://github.com/justrach/codegraff/blob/main/docs/adr/0238-rlm-binds-keep-the-whole-mcp-result.md"},{"id":"adr-0225","title":"ADR 0225: the MCP load result states the slim rule","url":"https://github.com/justrach/codegraff/blob/main/docs/adr/0225-mcp-load-result-states-the-slim-rule.md"},{"id":"adr-0029","title":"ADR 0029: MCP inside rlm and return shapes","url":"https://github.com/justrach/codegraff/blob/main/docs/adr/0029-mcp-inside-rlm-and-return-shapes.md"}]}