2026-09-24 · Engineering notes · Immediacy
Four engines, one backpack
The context engine that makes coding agents faster starts with an unglamorous win: getting four heavy libraries to live happily in one program.
§ Four heavy engines, one program
A graph-and-vector database, a code-parsing engine, a local decision model, and a text embedder each drag in deep, conflicting native dependencies. All of them link into a single Rust binary and run in-process on an Apple-Silicon GPU, with no Python at inference time, no separate services, and nothing on the network in the hot path. Every tool that works on your code then shares one index instead of each rebuilding its own.
§ A graph where a file is the atom
The code graph treats one file as the atomic unit. Re-indexing a file replaces its whole subtree, so the operation is idempotent and leaves no orphans. Only changed files are re-read, matched by content hash, and a file watcher re-indexes on save. A single writer owns the database, and every client, even the command line, goes through HTTP and never opens it directly.
§ The model that wouldn’t load
The first-choice code-embedding model couldn’t be loaded by the available Rust library, which didn’t support its newer architecture. Because the embedder sits behind a clean swap point, a solid general model went in instead, and a real code model can be added later without touching anything else.
§ Your git history is a signal
The engine mines your commit history for two things: which files tend to change together, and clusters of past work similar to the task at hand. Both nudge the right files up the ranking, so the brief can include a “related past work” section. It shells out to git, which the daemon already needs, instead of pulling in a whole Rust git library.
For what all this buys in practice, see the benchmark write-up. For how the local model is kept honest, see Making a small model earn its keep.