bonsai treats the context window as a managed budget: cache-stable prefixes,
token estimation per provider family, continuous garbage collection,
graduated compaction, and default-on task episodes that archive finished
work into compact, recallable cards. /ctx shows exactly where the window
is going and gives you manual control.
The system message is: static persona prompt + a project context block + optional operator suffix. Project context is split deliberately:
AGENTS.md → CLAUDE.md → .cursorrules, nearest-first up the
tree, 16 KiB cap each), the skills index, the subagents index, and the
repository map. All
byte-stable across turns and restarts.## Volatile state heading
(providers split their cache breakpoint on it): git branch/status (with a
run-start baseline separating pre-existing WIP from this run’s changes),
the stale-read advisory, and peer status.Connections with rolling-history caching (and Codex always) instead receive volatile state appended to history, keeping the entire prefix append-only. Memory is deliberately not in the system prompt — recalled entries arrive as per-turn background data notes.
A per-model estimator drives context-pressure decisions (billing truth stays
with provider-reported usage): local tiktoken or Qwen tokenizers, the
Anthropic count_tokens API, or a chars/4 heuristic — each tagged with a
confidence. Over-budget bails only trust high-confidence estimates; image
payloads are redacted before counting. See
Models.
Thresholds use the input window left after reserving response headroom. Outside SMOL, the reserve is 16k tokens on windows of at least 64k and scales down to 25% on smaller windows (for example, a 12k window reserves 3k rather than 8k):
| Trigger | Default | Smol |
|---|---|---|
| Continuous GC reclaim | 75% | 35% |
| Automatic compaction | 95% | 50% |
| Compaction target | 50% | 35% |
Every turn, existing stubs are re-asserted and pinned rows restored — byte-identically, so the provider prompt cache stays warm. Under pressure (the GC trigger), a reclaim pass introduces new stubs for:
Every rewrite must pay for itself: new stubs require ≥8,000 tokens or ≥8%
saved (rewriting history invalidates the provider cache, so small savings
are worse than nothing). Stubbed originals are restorable from /ctx.
Read reuse. A repeated read of unchanged content collapses to a compact
pointer ([reused previous read]) instead of resending bytes — but an
explicit re-request after that returns real bytes; reads are annotated, never
withheld. Guards watch for pathological loops: repeated identical
inspections, read storms on one target, repeated failed calls, and
implementation stalls. Read-only loop guards reject the offending call and
admit one explicit retry with a fresh budget; they do not terminate an
otherwise healthy autonomous run.
When the automatic trigger fires (or you run /compact), a graduated draft
is built:
/ctx.Always protected: the system message, the last three message groups, the two
latest user messages, and read-reuse pointer targets. Summaries carry a
trust guard marking them lossy reference data. Every compaction records a
telemetry event (tokens before/after, prefix hashes, cache impact) visible
in /ctx.
/compact [preview] [target=<tokens>] [fallback|provider|deterministic]
runs or previews one manually.
Enabled by default; BONSAI_EPISODES=0 is the emergency opt-out (zero
tracking, byte-identical disabled tool arrays, no persistence).
set_session_title
and plan_replace_draft carry an episode_action (new_topic closes the
span at the last human message; same_topic renames in place), and hard
resets (/clear, /start, /review) close unconditionally. Tiny spans
never close — a new_topic on a span without real work just renames.[Episode archived] #N "title" card (≤1,600 chars) with
recall instructions. Eviction is all-or-nothing per episode and batched
(one wire mutation per pass), and only happens when the rewrite
economics clear — ≥8,000 tokens or ≥8% saved, or the pressure batch
resolves pressure by itself. A cancelled provider call leaves context
untouched.recall tool pages an archived episode’s real messages and tool
outputs back in ({episode: N}, cursor-paged) or searches the transcript
and archives ({query: …}) — always inside an untrusted-data frame.
Repeat recalls of an unchanged page collapse to a pointer. Restoring from
/ctx re-splices originals; restored episodes are never re-evicted./episodes shows the ledger; resume repairs any inconsistent episode
state conservatively (demoted to non-evictable, archives preserved)./smol (or the Context row in /settings) switches to a compact profile
for small models: lower compaction thresholds and reserve, a reduced system
prompt (bash/terminal/todowrite/set_session_title tools), no repo map or
built-in skill/subagent/memory indices, 2 KiB steering cap, deterministic
compaction summaries only, and memory recall off.
/ctxThe context modal breaks the window down node by node with three views
(ledger / wire / turns): token counts, inclusion state, stub reasons,
compaction events, and episode reports, reconciled against actual provider
usage. Keys in the ledger view: p pin (never GC), d drop, s stub,
r restore, Enter/arrows expand. The header’s context chip (fill bar +
percent) is the always-visible summary; click it to open /ctx.
| Concern | Location |
|---|---|
| Preflight, thresholds, compaction | src/agent/compaction/ (mod, draft, gc, summary) |
| Read evidence & reuse | src/agent/read_persistence.rs, src/tool/read_evidence.rs |
| Episodes | src/episode.rs, src/agent/episodes/ |
| Project context block | src/context.rs |
| Estimators | src/provider/token.rs |
| Smol profile | src/smol.rs, src/agent/controls/smol.rs |
/ctx report & modal |
src/context_view/, src/tui/widgets/modal/context.rs |