2026-09-28 · Benchmark
Cutting an agent’s orientation tax
Before an AI coding agent edits a line, it spends tokens and time just figuring out where to edit: searching, reading, following imports. I built a shared context layer to skip that step, and measured what it saved on real tasks.
§ Summary
On real feature tasks from Immediacy’s own history, giving the agent a ranked context brief from immediacy-ctx, with one-line symbol summaries and co-change history enabled, cut input tokens by 60%, tool calls by 47%, and wall-clock time by roughly the same, while raising retrieval precision by 71%.
| Condition | Input tokens | Output tokens | Tool calls |
|---|---|---|---|
| No context (baseline) | 2,852,002 | 22,238 | 36 |
| Syntactic brief only | −25% | −32% | −33% |
| + summaries & co-change | −60% | −50% | −47% |
§ Method
Tasks were harvested from real git history: each past commit that touched indexed code becomes a labelled task, where the commit subject is the “feature description” and the files it changed are the ground truth an ideal brief would surface. Each task went to a subagent as a read-only orientation job (“find the exact files and functions you’d edit; do not edit”), which isolates exactly what the context layer optimizes. Token counts came from each subagent’s own transcript; retrieval precision and recall were computed separately, with no agent runs, to avoid circularity.
§ Why it works
Summaries mainly raise precision: they drop the wrong files a purely syntactic match would surface. The sharpest gain was on a “delete” task where keyword matching chased red herrings. Tighter briefs mean fewer wasted reads, which is where the token win comes from. And because wall-clock time is dominated by sequential model round-trips, cutting tool calls buys almost the same cut in time: a matched pair showed a 43% drop in calls and a 46% drop in wall-clock.
§ The bigger idea
The engine underneath immediacy-ctx (learn context over time, store it as a graph and vector memory, retrieve the relevant slice on demand) is the same one behind Anamanti’s memory of a home. It’s one idea about memory and learning, working on two unrelated problems. Improve the core, and both a coding agent and a family assistant get better.