{"schemaVersion":1,"canonical":"https://codegraff.com/blog/graff-0-0-302-16","markdown":"https://codegraff.com/blog/graff-0-0-302-16/markdown","slug":"graff-0-0-302-16","title":"DeepSeek gets to work sooner in graff, and funded companies need a license","seoTitle":"graff v0.0.302.16: faster DeepSeek V4 Pro by default, and a commercial license for funded companies","description":"graff now sends DeepSeek its own low reasoning level by default: 26% less time per task on DeepSeek V4 Pro across our coding and MCP tasks, and 39% less on sub-agent tasks. From v0.0.302.16, companies that have raised more than US$500k need a commercial license.","publishedAt":"2026-10-02","keywords":["graff","DeepSeek V4 Pro","reasoning effort","coding agent","MCP","sub-agents","AGPL","commercial license","codegraff"],"intro":"Two releases today. v0.0.302.15 changes what graff's default effort sends to DeepSeek: thinking stays on, at DeepSeek's own low level. On DeepSeek V4 Pro our 21 coding and MCP tasks took 26% less time per task and our sub-agent tasks 39% less, with one missed run in 62. v0.0.302.16 adds one term to graff's license: from this version, a company that has raised more than US$500,000 needs a commercial license to use graff. Individuals and other organizations keep using it under the AGPL.","hero":{"src":"/blog/graff-frontier-harness-evals/cost-accounting.webp","alt":"Two mouse engineers count brass tokens and verify task cards beside a curling receipt and discarded failed attempts.","caption":"Fewer tokens per task, and every run still checked against the task. Conceptual illustration."},"sourceNote":"The DeepSeek numbers come from the public eval write-up in the codegraff repo: 21 coding and MCP tasks and 10 sub-agent tasks, 2 runs of each task per build or setting, interleaved task by task, on DeepSeek V4 Pro through DeepSeek's own API with one key. The license summary follows the LICENSE in v0.0.302.16, and the LICENSE text is what applies.","sections":[{"id":"short","title":"The short version","paragraphs":["graff's default effort is medium, and it used to send DeepSeek exactly that. DeepSeek V4 Pro then spent the first request of an MCP task reasoning for thousands of characters, planning the whole job before its first tool call. From v0.0.302.15 the default sends low with thinking still on, and the model gets to work sooner.","We ran v0.0.302.14 and v0.0.302.15 side by side on DeepSeek's own API with one key: 26% less time per task across 11 coding and 10 MCP tasks, 39% less on 10 sub-agent tasks, and 36% fewer output tokens.","v0.0.302.16 changes the license, not the code. Companies that have raised more than US$500,000 now need a commercial license for any use of graff. Versions up to v0.0.302.15 keep the license they shipped with."],"sources":["deepseek-results","release-15","release-16"]},{"id":"deepseek","title":"What changed for DeepSeek","paragraphs":["Effort is how hard a model thinks before it acts. graff has one effort setting for every provider, and its default is medium. DeepSeek documents three levels, low, high and max, and graff used to pass medium through as it was.","The traces showed where the time went: the first request of each turn. On tasks that read an MCP server, V4 Pro reasoned for 7,800 to 13,200 characters on average before its first tool call, laying out the whole job in advance. deepseek-harness, DeepSeek's own agent harness, sends no effort setting and reasoned for a fraction of that on the same tasks.","graff's default now sends low with thinking enabled. On those tasks V4 Pro's first request reasons for about 3,200 characters, and the rest of the turn works as before. /effort high still asks for more thinking, and /effort low turns thinking off. The picker and status line still call the default Medium: it is still graff's medium, and only what DeepSeek receives changed."],"sources":["adr-0237","deepseek-results"]},{"id":"results","title":"v0.0.302.14 against v0.0.302.15","paragraphs":["Both releases ran DeepSeek V4 Pro on DeepSeek's API with the same key, interleaved task by task, 2 runs of each task. Every run had to pass the task's own check.","v0.0.302.15 finished the coding and MCP tasks in 15.9 seconds per task against 21.6, most of the gain on the MCP tasks: 26.5 seconds against 37.8. On the sub-agent tasks, which ask the model to split work across parallel sub-agents, it took 37.7 seconds against 61.9, and both releases passed all 20 runs.","It missed one run. On a task that counts the log lines starting with the token ERROR, it ran grep -c '^ERROR', which also counts the ERRORS_TOTAL line the prompt rules out, and wrote 5 instead of 4. Less thinking means less checking; for work that needs exact figures, /effort high gives the model room to check.","Speed moves between rounds. An earlier round, with a build that sent the same request, measured 11.9 seconds against 24.2 per task. The numbers above come from this release's own code."],"sources":["deepseek-results"],"table":{"caption":"DeepSeek V4 Pro on DeepSeek's API, one key. Mean time per task over 2 runs of each task.","headers":["Release","Coding and MCP","Coding","MCP","Sub-agent tasks","Runs passed"],"rows":[["v0.0.302.14","21.6s","6.8s","37.8s","61.9s","62/62"],["v0.0.302.15","15.9s","6.2s","26.5s","37.7s","61/62"]]}},{"id":"thinking-off","title":"Why not turn thinking off","paragraphs":["In v0.0.302.14 we turned MiMo's thinking off below high effort, because thinking was most of a MiMo task's time. DeepSeek is different. With thinking off, the coding and MCP tasks took 7.6 seconds per task, but the model missed one coding run and 5 of 20 sub-agent runs, where low with thinking on missed none and 2 in the same rounds.","The misses with thinking off were the kind a moment's checking catches: a count that included a line the prompt excluded, two services' results swapped while merging, a route's method misread. So the default keeps thinking on at the low level, and thinking off stays one command away for anyone who prefers that trade.","deepseek-harness, at its own defaults, took 25.5 seconds per task in the same round and passed every run."],"sources":["deepseek-results","dsh-results"],"table":{"caption":"The round that chose the default: coding and MCP tasks, 2 runs of each. Low with thinking on is what v0.0.302.15 sends.","headers":["Setting","Per task","Coding","MCP","Runs passed","Output tokens per task"],"rows":[["v0.0.302.14 default (medium)","24.2s","5.3s","45.0s","42/42","2,578"],["Low, thinking on","11.9s","5.2s","19.3s","42/42","1,104"],["Thinking off","7.6s","4.7s","10.9s","41/42","560"],["deepseek-harness","25.5s","7.0s","45.8s","42/42","2,501"]]}},{"id":"license","title":"A license term for funded companies","paragraphs":["graff is licensed under a modified AGPL-3.0, held jointly by its two authors. From v0.0.302.16 the license has one more term.","A company that, together with its affiliates, has raised more than US$500,000 from investors or lenders (equity, convertible notes, SAFEs or similar, at any valuation) may use graff only under a commercial license granted jointly by the authors. Use includes running graff on the company's own computers, servers, CI and networks, and making it available to its people or users over a network. A company that crosses the line has 30 days to get a license or stop using graff.","Nothing changes for individuals using graff for themselves, or for organizations that are not funded companies: they use graff under the AGPL as before. Versions up to and including v0.0.302.15 keep the license they shipped with. The Apache-2.0 SDKs keep their license, but a copy of graff they include or download is under graff's license.","This is a summary, and the LICENSE text is what applies. For a commercial license, support or deployment help, write to rach@standardharness.com."],"sources":["license","release-16"]},{"id":"method","title":"How we ran it","paragraphs":["Every task runs in a fresh git repository with a fresh home directory, so no run sees user-level settings or MCP servers. Builds and settings run interleaved, task by task, a few at a time, each after one unscored warm-up request. Wall time is the runner's clock around each process, from start to exit.","The MCP tasks read a Linear-shaped MCP server and write a report. The sub-agent tasks ask the model to split work across parallel sub-agents, and a matching set does the same work without delegating. Every run used DeepSeek V4 Pro through DeepSeek's own API with one key.","Two runs per task on one machine is a small sample, and DeepSeek's response times move through the day: the gap was 2.0x in one round and 1.36x in the next. Low with thinking on beat v0.0.302.14's default in every round we ran."],"sources":["deepseek-results","graff-evals"]},{"id":"try","title":"Try it","paragraphs":["Run graff update to get v0.0.302.16. If you use DeepSeek there is nothing to configure: the default effort now sends low. The tasks, the harness drivers and the per-task results are in the codegraff repo, so you can rerun the comparison with graff-evals on your own tasks."],"sources":["release-15","release-16","graff-evals"]}],"sources":[{"id":"deepseek-results","title":"DeepSeek V4 Pro at graff's default effort (results per task)","url":"https://github.com/justrach/codegraff/blob/main/evals/graff-vs-deepseek-harness/DEFAULT-EFFORT.md"},{"id":"dsh-results","title":"graff vs deepseek-harness on DeepSeek V4 Pro","url":"https://github.com/justrach/codegraff/blob/main/evals/graff-vs-deepseek-harness/RESULTS.md"},{"id":"adr-0237","title":"ADR 0237: DeepSeek thinks at its low level by default","url":"https://github.com/justrach/codegraff/blob/main/docs/adr/0237-deepseek-thinks-at-its-low-level-by-default.md"},{"id":"license","title":"graff's LICENSE in v0.0.302.16","url":"https://github.com/justrach/codegraff/blob/v0.0.302.16/LICENSE"},{"id":"release-15","title":"graff v0.0.302.15","url":"https://github.com/justrach/codegraff/releases/tag/v0.0.302.15"},{"id":"release-16","title":"graff v0.0.302.16","url":"https://github.com/justrach/codegraff/releases/tag/v0.0.302.16"},{"id":"graff-evals","title":"graff-evals: tasks, harness drivers and runner","url":"https://github.com/justrach/codegraff/tree/main/graff-evals"}]}