{"schemaVersion":1,"canonical":"https://codegraff.com/blog/codedb-with-graff","markdown":"https://codegraff.com/blog/codedb-with-graff/markdown","author":{"name":"Rach Pradhan","url":"https://justrach.com"},"slug":"codedb-with-graff","title":"How CodeDB works with Graff and your coding agent","seoTitle":"CodeDB + Graff: Free AI Coding Agent Setup & MCP Guide","description":"Set up free CodeDB with Graff, connect Codex/ChatGPT, Grok or Kimi, and use MCP repository context to complete and evaluate coding tasks.","publishedAt":"2026-09-11","modifiedAt":"2026-09-12","readTime":"11 min","sections":[{"id":"roles","label":"Where CodeDB fits","blocks":[{"kind":"paragraph","children":["Start with Graff and the model account you already use: Kimi Code, Codex/ChatGPT, Grok/SuperGrok, or another supported provider. Connect your account, open a repository, and give Graff a coding task. CodeDB supplies repository context so the agent can find the right implementation, understand its callers, and locate the tests."],"lead":true},{"kind":"paragraph","children":["In the Codegraff stack, each tool has a different job:"]},{"kind":"diagram","caption":"Graff uses CodeDB for context. Pro adds optional local tools; your agent can also use its own editor. Motion illustrates the connections, not live activity."},{"kind":"paragraph","children":["Graff and CodeDB are free and open source. You can use Graff on its own, or add CodeDB to a coding agent you already use. Pro is an optional upgrade. The"," ",{"tag":"a","href":"/docs#stack","children":["product overview"]}," ","explains how the pieces fit together."]}]},{"id":"start","label":"Connect your model","title":"Start with Graff and your preferred model","blocks":[{"kind":"paragraph","children":["Graff supports multiple providers out of the box. Connect a supported subscription through its built-in sign-in flow, or use a provider API key. You choose the model; Graff handles the coding session, tools, and execution around it."]},{"kind":"paragraph","children":["For the macOS app,"," ",{"tag":"a","href":"https://github.com/justrach/codegraff/releases/latest/download/Codegraff-macos-arm64.dmg","children":["download Codegraff for Apple Silicon"]},", open the disk image, and move the app to Applications. The"," ",{"tag":"a","href":"https://github.com/justrach/codegraff#install","children":["Graff installation guide"]}," ","also includes the standalone CLI installer. The commands below use that CLI."]},{"kind":"heading","text":"Connect an existing subscription"},{"kind":"paragraph","children":["Choose the account you want to use. Run its command, then complete the browser authorization. These subscription connections use sign-in rather than requiring a separate API key:"]},{"kind":"copy","text":"graff login codex","label":"Copy Codex / ChatGPT sign-in","code":true,"title":"Codex / ChatGPT"},{"kind":"copy","text":"graff login xai","label":"Copy Grok / SuperGrok sign-in","code":true,"title":"Grok / SuperGrok"},{"kind":"copy","text":"graff login kimi","label":"Copy Kimi Code sign-in","code":true,"title":"Kimi Code"},{"kind":"heading","text":"Use another provider"},{"kind":"paragraph","children":["Graff also supports API-key connections for providers including Anthropic, OpenAI, DeepSeek, Google Gemini, Z.AI, OpenRouter, Vercel AI Gateway, MiniMax, Mistral, Groq, and Cerebras. Subscription sign-in and API-key access are different connection methods; choose the one offered for your provider."]},{"kind":"paragraph","children":["Start Graff inside your project:"]},{"kind":"code","text":"cd /path/to/your/project\ngraff"},{"kind":"paragraph","children":["Enter ",{"tag":"code","children":["/model"]}," to browse providers and models. When a provider needs credentials, Graff offers its supported sign-in or API-key entry. You can also select a provider directly with ",{"tag":"code","children":["/model codex"]},", ",{"tag":"code","children":["/model xai"]},","," ",{"tag":"code","children":["/model kimi"]},", or, for example, ",{"tag":"code","children":["/model deepseek"]},". See the"," ",{"tag":"a","href":"https://github.com/justrach/codegraff","children":["Graff command and provider reference"]}," ","for the available options."]},{"kind":"paragraph","children":["Then describe the work: “Explain how session expiry works in this repository, find the relevant tests, and propose a fix for expired sessions being accepted.” The workflow is the same whichever connected provider you choose."]},{"kind":"paragraph","children":["To use CodeDB for the investigation, follow the"," ",{"tag":"a","href":"#connect","children":["CodeDB connection steps below"]},", check ",{"tag":"code","children":["/mcp"]},", and use the copyable first-task prompt. Your provider connection supplies the model; the MCP connection supplies CodeDB’s repository tools."]}]},{"id":"free","label":"What is free?","title":"What is free?","blocks":[{"kind":"paragraph","children":[{"tag":"strong","children":["CodeDB is free and open source."]}," You can install it, index your repository, and connect it to a supported agent without buying CodeDB Pro. Graff is also free and open source. There is no requirement to switch to Graff to use CodeDB."]},{"kind":"paragraph","children":["Use an existing supported subscription or connect a provider API key. Your subscription’s limits or your API provider’s usage charges still apply. Graff and CodeDB are free software; CodeDB Pro is the optional paid tool layer described"," ",{"tag":"a","href":"#pro","children":["below"]},"."]},{"kind":"paragraph","children":["This walkthrough starts with Graff. CodeDB also works with Claude Code, Codex, Cursor, and other supported clients if you want to use it in another workflow. Both projects are available on GitHub:"," ",{"tag":"a","href":"https://github.com/justrach/codedb","children":["CodeDB"]}," ","and"," ",{"tag":"a","href":"https://github.com/justrach/codegraff","children":["Graff"]},"."]}]},{"id":"capabilities","label":"What you can do","title":"What you can actually do with it","blocks":[{"kind":"paragraph","children":["You keep describing work in ordinary language. The agent uses CodeDB behind the conversation to find evidence. Here are four places to start:"]},{"kind":"list","ordered":false,"items":[[{"tag":"strong","children":["Understand an unfamiliar repository."]}," Ask where a request enters the application, which module owns the behavior, and where its tests live. Use the returned paths and definitions to build a small map before changing anything."],[{"tag":"strong","children":["Investigate a bug."]}," Start from the symptom, locate the implementation, and follow the relevant calls. For an expired-session bug, that means checking both the expiry comparison and the request path that reaches it."],[{"tag":"strong","children":["Plan a refactor."]}," Inspect a function and its callers before changing its signature. Ask the agent to identify affected tests and any references that need a separate text search."],[{"tag":"strong","children":["Review a proposed change."]}," Ask the agent to explain how the modified code connects to its surrounding modules. Review the actual diff and run the project checks before accepting the patch."]]},{"kind":"paragraph","children":["The common sequence is find, understand, change, verify. CodeDB helps with the first two; your agent’s editor, shell, and test runner carry the work through to a checked result."]}]},{"id":"index","label":"From files to context","title":"From files to useful context","blocks":[{"kind":"paragraph","children":["CodeDB scans a project and builds indexes for several kinds of questions. A structural outline identifies declarations and their locations. A word index finds identifiers. A trigram index narrows the files that might contain a text match. A dependency graph records connections between files."]},{"kind":"paragraph","children":["Those indexes let an agent ask for a function or its surrounding code without first reading every file in the directory. File watching keeps the working index updated as the repository changes. See the"," ",{"tag":"a","href":"https://github.com/justrach/codedb/blob/main/docs/architecture.md","children":["architecture guide"]}," ","for the underlying data structures."]},{"kind":"paragraph","children":["The useful output is evidence the agent can inspect: paths, line numbers, definitions, and related code. Language support and static resolution affect what it can find. A caller lookup is a starting point for checking impact; runtime dispatch, generated code, and configuration can still require a closer look."]}]},{"id":"workflow","label":"A task, step by step","title":"A task, step by step","blocks":[{"kind":"paragraph","children":["Suppose your task is “fix expired sessions still being accepted.” First give the agent a concrete success condition: an expired session is rejected, a valid session still works, and a regression test covers the boundary. An agent can use CodeDB to narrow the investigation before editing:"]},{"kind":"list","ordered":true,"items":[[{"tag":"strong","children":["Orient."]}," Ask ",{"tag":"code","children":["codedb_context"]}," for the task with a bounded response budget."],[{"tag":"strong","children":["Inspect."]}," Use ",{"tag":"code","children":["codedb_explain"]}," on the relevant session validator to read its definition and callers."],[{"tag":"strong","children":["Trace."]}," If the connection is unclear, ask"," ",{"tag":"code","children":["codedb_callpath"]}," for the resolved chain between the request handler and validator."],[{"tag":"strong","children":["Change and verify."]}," Use the agent’s file-editing tools, then run the affected tests and inspect the diff."]]},{"kind":"paragraph","children":["Here is an example argument object for the first tool call; the task describes a hypothetical bug:"]},{"kind":"code","text":"{\n  \"task\": \"Find session expiry validation and its tests\",\n  \"semantic\": \"local\",\n  \"max_tokens\": 2400\n}"},{"kind":"paragraph","children":["This example explicitly selects local retrieval. The default hybrid mode can use remote semantic services; choose ",{"tag":"code","children":["semantic: \"local\""]}," when the retrieval call needs to stay on the machine. Your coding agent’s model connection is a separate data boundary. The"," ",{"tag":"a","href":"https://github.com/justrach/codedb#-mcp-tools","children":["tool reference"]}," ","describes the options."]}]},{"id":"connect","label":"Add CodeDB to Graff","title":"Add CodeDB to Graff","blocks":[{"kind":"paragraph","children":["MCP is the connection that lets your coding client call CodeDB’s tools. The client launches the local CodeDB process, sends a tool request, and gives its response to the agent. You do not need to upload a repository to the Codegraff website to use the local index."]},{"kind":"heading","text":"1. Install CodeDB"},{"kind":"paragraph","children":["On macOS or Linux, run this in your terminal:"]},{"kind":"copy","text":"curl -fsSL https://codedb.codegraff.com/install.sh | bash","label":"Copy CodeDB install command","code":true},{"kind":"paragraph","children":["The installer registers CodeDB with supported clients it detects, including Claude Code, Codex, and Cursor. Restart the client if needed, open your repository, and ask it to check CodeDB’s status. If it points at the wrong folder, configure the project root explicitly. The"," ",{"tag":"a","href":"https://github.com/justrach/codedb/blob/main/docs/mcp.md","children":["MCP setup guide"]}," ","has client-specific configuration and troubleshooting."]},{"kind":"heading","text":"2. Connect CodeDB to Graff"},{"kind":"paragraph","children":["With Graff installed and your chosen provider connected, register CodeDB as a local MCP server. Replace the example path with your repository’s absolute path:"]},{"kind":"code","text":"graff mcp add codedb -- codedb mcp /path/to/your/project"},{"kind":"paragraph","children":["Then start Graff in that repository and check ",{"tag":"code","children":["/mcp"]}," before beginning a task. Graff provides the model session and execution tools; CodeDB supplies repository context through that connection. See"," ",{"tag":"a","href":"https://github.com/justrach/codegraff#install","children":["Graff’s installation guide"]}," ","if you still need the CLI. The macOS app download and the standalone CLI installer are separate installation options."]},{"kind":"paragraph","children":["If your client was detected, check that connection before adding another entry. For a manual Cursor setup, merge this entry into your project’s"," ",{"tag":"code","children":[".cursor/mcp.json"]},", keeping any other servers. Run"," ",{"tag":"code","children":["command -v codedb"]}," in your terminal to find the executable and replace both example paths:"]},{"kind":"code","text":"{\n  \"mcpServers\": {\n    \"codedb\": {\n      \"command\": \"/absolute/path/to/codedb\",\n      \"args\": [\"mcp\", \"/absolute/path/to/project\"]\n    }\n  }\n}"},{"kind":"paragraph","children":["The"," ",{"tag":"a","href":"https://github.com/justrach/codedb/blob/main/docs/mcp.md#2-client-specific-configuration","children":["client-specific setup examples"]}," ","also cover Claude Code, Codex CLI, and other MCP clients. Use the configuration format for your client."]},{"kind":"heading","text":"3. Check it with a first task"},{"kind":"paragraph","children":["Open your repository and start a fresh agent session after changing the MCP configuration. Ask it to call ",{"tag":"code","children":["codedb_status"]}," and confirm that the indexed project is yours. If the tree is empty or the root is wrong, supply the absolute repository path through the tool’s ",{"tag":"code","children":["project"]}," argument. If the process cannot launch, check the executable path."]},{"kind":"paragraph","children":["Then try this prompt, adapting the feature name to your project:"]},{"kind":"copy","text":"Use CodeDB to map how this project handles session expiry. Start with codedb_status, then codedb_context with semantic: local. Identify the validator, its callers, and the relevant tests. Cite file paths and line numbers. Explain the expected behavior before proposing a change.","label":"Copy first-task prompt"},{"kind":"paragraph","children":["You should see real repository paths and code references in the response. Check one against the source. Once the explanation is sound, ask the agent to make the smallest appropriate change and run the relevant tests."]}]},{"id":"pro","label":"When Pro helps","title":"When CodeDB Pro helps","blocks":[{"kind":"paragraph","children":["Once an agent has found the right files, the next bottleneck may be repeated tool work. A change can require several reads, searches, edits, and diff checks. CodeDB Pro keeps a local daemon available and can batch operations instead of paying for a separate process and round trip for every step."]},{"kind":"paragraph","children":["Its edit tools can address a symbol or a line range and return a compact patch result. That is useful when an agent needs to make a focused change and review what happened next. Free CodeDB remains the navigation layer; your client’s native editor can handle the changes without Pro."]},{"kind":"paragraph","children":["For the editing behavior and its checks, read"," ",{"tag":"a","href":"/blog/codedb-pro-0-2-12","children":["the CodeDB Pro edit guide"]},". You can also"," ",{"tag":"a","href":"/codedb#pricing","children":["compare the available plans"]},"."]}]},{"id":"evaluations","label":"Tasks and evaluations","title":"How this connects to the new tasks and evaluations","blocks":[{"kind":"paragraph","children":["Our previous"," ",{"tag":"a","href":"/blog/graff-frontier-harness-evals","children":["FrontierHarness evaluation post"]}," ","asks whether Graff can finish a task with a model and tools. That is the next step after finding useful context: does the resulting work actually pass?"]},{"kind":"paragraph","children":["The suite contains ",{"tag":"strong","children":["30 tasks in two groups"]},". The 21 Terminal-Bench 2.1 tasks grade the final environment with public tests. The nine DeepSWE tasks grade a submitted patch against a real project’s verifier. A working environment and a patch that applies cleanly and passes regression tests are distinct outcomes."]},{"kind":"table","caption":"Recorded Graff results from the previous post. Counts show passes / tasks.","headers":["Model with Graff","Terminal","DeepSWE","Full suite"],"rows":[["Grok 4.6","20 / 21","1 / 9","21 / 30"],["Kimi K3","17 / 21","1 / 9","18 / 30"]]},{"kind":"paragraph","children":["Both Graff runs passed the KaTeX multicolumn-array-spans patch task; the other eight DeepSWE tasks did not pass. The failures included patch-application problems and test failures. On the terminal side, the evaluation notes stress exact final-state requirements: leave only the permitted files in a directory, keep a required service running, and recover complete data rather than merely producing something parseable."]},{"kind":"paragraph","children":["Those details matter in everyday work too. For the session-expiry example, finding the validator is progress. Finishing means changing the right behavior, preserving valid sessions, and demonstrating both with tests. CodeDB can help locate that code and those tests; the harness still has to execute the work and check the final result."]},{"kind":"paragraph","children":["The recorded terminal cost per pass was $0.31 for Graff / Grok 4.6 and $0.21 for Graff / Kimi K3. Each figure divides the historical estimated cost of all 21 terminal attempts, including failures, by passes. These are recorded evaluator list-price estimates, not current pricing or a promise about your bill."]},{"kind":"paragraph","children":["Keep the conditions attached to the scores: Graff received extra"," ",{"tag":"code","children":["BENCH_APPEND"]}," evaluation instructions and ran in Docker, while the published comparison board used different instructions and a prepared VM. These are repository-reported results, not an independent rerun. They do not isolate CodeDB’s contribution or establish that adding CodeDB causes the reported pass rate. The"," ",{"tag":"a","href":"/blog/graff-frontier-harness-evals#terminal","children":["full evaluation and linked protocol"]}," ","explain the comparison."]},{"kind":"paragraph","children":["Our earlier"," ",{"tag":"a","href":"/blog/graff-rlm-token-efficiency","children":["Graff runtime article"]}," ","looks at repeated context work and token use. CodeDB’s retrieval work, Graff’s context management, and end-to-end task completion answer related but different questions. Read them together: useful evidence, efficient execution, and a verified result all matter."]}]},{"id":"measure","label":"Try your own task","title":"Give it a task you can actually grade","blocks":[{"kind":"paragraph","children":["A smaller context response is useful when it contains the evidence needed to finish the task. It is less useful if missing context sends the agent back through several searches. Compare completed work, test results, tool calls, token usage, and elapsed time together."]},{"kind":"paragraph","children":["Try CodeDB on a familiar bug or a small refactor. Ask the agent to identify the relevant code, explain the proposed change, and verify it. Keep the model and task comparable when judging the difference. Our"," ",{"tag":"a","href":"/blog/codedb-retrieval-experiments","children":["retrieval experiments"]}," ","and"," ",{"tag":"a","href":"/blog/graff-frontier-harness-evals","children":["Graff evaluation write-up"]}," ","show why the test conditions matter."]},{"kind":"list","ordered":true,"items":[[{"tag":"strong","children":["Write the task and success condition first."]}," Use a reproducible bug or a small change with a known expected result. Keep the verifier independent of the agent’s own assessment."],[{"tag":"strong","children":["Compare from the same starting point."]}," Use separate clean worktrees at the same commit, the same model and instructions, and the same test environment. Try the agent with and without CodeDB; record any extra setup or guidance."],[{"tag":"strong","children":["Record the whole attempt."]}," Capture test outcomes, input and output tokens, tool calls, elapsed time, and estimated cost, including retries and failed attempts. Note whether the index was cold or already built."],[{"tag":"strong","children":["Repeat before generalizing."]}," One successful run is a useful smoke test. Several tasks and repeated runs tell you more about whether the change helps your actual workload."]]},{"kind":"paragraph","children":["Start with the free setup above. Get the agent to show you where a behavior lives, then carry one small task through to passing checks. That gives you something concrete to judge before changing the rest of your workflow."]}]}]}