Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions changelog.d/3247.added.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- **Injected context is now bounded and policy-driven (ADR 0108 D6, #3247)**. A new `context.budget_pct` (default 8% of the model window, never below 16k chars) caps everything injected per turn; over budget the lowest-priority parts shed first — recalled knowledge, then the prior-session digest, then skill descriptions one at a time (names never drop) — while working state and always-on memory are never shed. Always-on memory is now selected by `delivery_policy="always"` rather than the `hot` domain (every hot write already carries it), so a fact on any domain can be pinned always-on, and rejected or expired rows never enter the prompt. The prompt preview API carries the budget summary and per-section `truncated` flags; the inspector renders them in a follow-up.
13 changes: 13 additions & 0 deletions config/langgraph-config.example.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -152,6 +152,19 @@ memory:
max_sessions: 10
max_tokens: 2000

# Projected context — everything injected per turn on top of the stable prompt
# (working state, always-on memory, the skill index, the prior-session digest,
# recalled knowledge) — is bounded to this share of the model's context window
# (ADR 0108 D6). Over budget, the lowest-priority parts shed first: recalled
# knowledge, then the prior-session digest, then skill descriptions (names never
# drop); working state and always-on memory are never shed. Never below 16k chars
# (room for always-on memory + the digest) — on a 32k window or smaller that floor
# applies. 8% of a 128k window is ~10k tokens — more than a typical turn injects,
# so the default sheds nothing there; lower it to make the budget bite. 0 =
# unbounded (also unbounded when the gateway reports no window for the model).
context:
budget_pct: 8

# Skill index — SQLite FTS5 store for learned skill-v1 artifacts.
# Skills (emitted by task() or authored as SKILL.md) are indexed here and listed
# in the system prompt as an always-on <available_skills> index (name + summary);
Expand Down
14 changes: 11 additions & 3 deletions docs/adr/0108-context-architecture-v2.md
Original file line number Diff line number Diff line change
Expand Up @@ -366,9 +366,17 @@ Budget-driven delivery, relevance-gated digest, write lifecycle
(confirmation + expiration). Rollback: revert to unbounded injection and
unconditional digest.

Shipped: D7 in #3246 (creation stamps via the trust tier, `set_review_state` +
`POST /api/memory/chunks/{id}/review`, the `superseded_by:<id>` chain with
insert-then-invalidate, `expires_in_days` on `memory_ingest`).
Shipped:

- D7 in #3246: creation stamps via the trust tier, `set_review_state` +
`POST /api/memory/chunks/{id}/review`, the `superseded_by:<id>` chain with
insert-then-invalidate, `expires_in_days` on `memory_ingest`.
- D6 in #3247: `context.budget_pct` (default 8%, floored at 16k chars, `0` =
unbounded), the fixed priority/shed order in `graph/projection.py`, always-on
selected by `delivery_policy="always"`, `deliverable=True` (rejected/expired
excluded) on every store's `list_chunks`/`search`, and the budget summary on
the prompt preview API. The `delivery_order` config key named under
Compatibility was not added — the order is fixed by this decision.

### Compatibility behavior

Expand Down
74 changes: 61 additions & 13 deletions docs/explanation/memory-and-knowledge.md
Original file line number Diff line number Diff line change
Expand Up @@ -160,19 +160,24 @@ prompt layer, don't just hope the store stays clean). Three parts, in order:
one tool call away with `recall_session(session_id)`; when the id is unknown,
`session_search(query)` searches reasoning-stripped, credential-redacted transcript
content in a lazy FTS5 index and returns ids to expand.
2. **Hot memory.** `domain="hot"` chunks are always-on operator facts: the
newest 100 under a 6 000-char budget inject **every turn**, loaded fresh per
turn so a just-added fact is seen immediately. Because that makes a silent
hot write the highest-leverage poisoning move available, every hot write —
agent tool, console route, or plugin — emits a `memory.hot_written` bus
event, and an optional gate (`knowledge.hot_write_confirm`) makes the
agent's own write path refuse `domain="hot"` entirely, reserving always-on
promotion for operator surfaces. Since [ADR 0108 D4](../adr/0108-context-architecture-v2.md)
each hot chunk also carries `delivery_policy="always"` (a hot write is stamped
with it whatever the caller said), and the gate refuses that policy on *any*
domain — through `memory_ingest` and `knowledge_ingest` alike; the per-turn
reader still selects on `domain="hot"` until D6 (#3187) switches delivery to
the policy column.
2. **Always-on memory ("hot").** Chunks with `delivery_policy="always"` are
always-on operator facts: the newest 100 under a 6 000-char budget inject
**every turn**, loaded fresh per turn so a just-added fact is seen
immediately. Always-on is a *policy*, not a domain
([ADR 0108 D4 + D6](../adr/0108-context-architecture-v2.md)): every
`domain="hot"` write is stamped with it whatever the caller said, so the
legacy hot domain still works, and a row on any other domain (say a
`preferences` fact) can be pinned always-on the same way. Rows an operator
has rejected (`review_state="rejected"`) or that have passed their
`expires_at` never deliver, whatever their policy. Because always-on makes
a silent write the highest-leverage poisoning move available, every
always-on write — agent tool, console route, or plugin — emits a
`memory.hot_written` bus event, and an optional gate
(`knowledge.hot_write_confirm`) makes the agent's own write paths
(`memory_ingest`, `knowledge_ingest`) refuse `domain="hot"` and
`delivery_policy="always"` alike, reserving always-on promotion for
operator surfaces. The console's Memory → Hot memory list shows exactly the
always-on set the reader selects.
3. **RAG hits.** The store is searched with the last user message and the
top-k results (default 10) inject, each line ending with its stored date and
trust label — `(stored 2026-07-01; trust: agent)`. Two policies shape the
Expand Down Expand Up @@ -200,6 +205,48 @@ delivery ([ADR 0108](../adr/0108-context-architecture-v2.md) D2). The result is
a typed `ProjectedContext` — text, per-section labels, the injected ids, and
the sources that fed it.

### Delivery budget and priority (ADR 0108 D6)

The projection is **bounded**: it may use at most `context.budget_pct` of the
model's context window (default 8%, chars//4 — the same token heuristic the
rest of the runtime uses), and never less than **16 000 chars** — roughly the
always-on cap (6 000) plus the digest cap (~2 000 tokens), so a small-window
model keeps its standing context whole and sheds only what lies beyond it. The
ceiling is derived from the window the gateway reports for the model; no
window (logged once — the knob is inert), or `budget_pct: 0`, means unbounded.
The stable prompt is not part of it — only what is injected on top per turn.

Within the budget, the parts fill in a fixed priority (highest first):

1. **Working state** — the agent's own live commitments (trusted, operational).
2. **Always-on memory** — `delivery_policy="always"`.
3. **The skill index** — capability awareness.
4. **The prior-session digest** — cross-session continuity.
5. **RAG hits** — relevance-matched knowledge.

Over budget, the lowest-priority parts shed first, each step re-measured so
nothing is cut mid-line: RAG hits go one whole hit at a time from the
lowest-ranked end, then the digest as a unit, then the skill index gives up
descriptions one row at a time down to its identity floor — every skill's name
stays listed ([ADR 0060](../adr/0060-skill-progressive-disclosure.md), #2867).
Working state and always-on memory are **never shed**: if they alone exceed
the budget they are delivered anyway and a warning names the sizes (once per
distinct standing-context size), because a silently missing standing
instruction is worse than an oversized prompt. The order is fixed by the ADR —
there is no `delivery_order` knob to reorder it.

What was shed is visible: the prompt **preview API** (`GET /api/prompts/preview`)
carries the `budget` summary — ceiling, chars used, and an `overflow` list
(label, items and chars dropped per part) — and marks each shed section
`truncated`; the console inspector renders both in a follow-up. The injection
log records the ids that actually **entered** the turn, never those merely
retrieved. With the default 8% on a 128k-window model the ceiling is ~10k
tokens — more than a typical turn injects (6 000 chars of always-on memory, a
~2 000-token digest, ten hits, a 2%-of-window skill index) — so nothing is
shed there until you lower it; on a 32k window or smaller the 16k-char floor is
the budget, so RAG hits and skill descriptions beyond it now shed where they
used to be unbounded.

### Trust tiers

Every chunk's `source_type` ranks into three deterministic tiers
Expand Down Expand Up @@ -296,6 +343,7 @@ tuning guidance in [Tune the knowledge store](../guides/knowledge.md)):
| `hot_write_confirm` | `false` | when on, the agent's `memory_ingest` and `knowledge_ingest` refuse always-on writes (`domain="hot"` or `delivery_policy="always"`) |
| `scope` | `scoped` | tier ([ADR 0041](../adr/0041-workspaces-and-tiered-stores.md)): `scoped` (private) · `shared` (host commons) · `layered` (read commons ∪ private, write private). See [Tune the knowledge store → Sharing across a fleet](../guides/knowledge.md#sharing-knowledge-across-a-fleet-the-commons) |
| `middleware.knowledge` | `true` | turn the whole subsystem on/off |
| `context.budget_pct` | `8` | (its own `context:` block) the projected-context ceiling as a % of the model window ([D6](#delivery-budget-and-priority-adr-0108-d6)); `0` = unbounded |

Three environment knobs override paths and persistence directly:

Expand Down
37 changes: 30 additions & 7 deletions docs/guides/knowledge.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,9 +41,31 @@ knowledge:
db_path: /sandbox/knowledge/agent.db # → ~/.protoagent/knowledge/agent.db fallback
```

One knob lives outside the block — the ceiling on everything the store (and the
digest, skill index and working state) may put into a turn
([ADR 0108 D6](/adr/0108-context-architecture-v2)):

```yaml
context:
budget_pct: 8 # % of the model window the per-turn injected context may use;
# over budget, RAG hits shed first, then the prior-session
# digest, then skill descriptions (names never drop); working
# state and always-on memory are never shed. 0 = unbounded.
```

The budget never drops below 16k chars (room for always-on memory + the digest),
so on a 32k window or smaller that floor is the budget and only RAG hits / skill
descriptions beyond it shed. With the default 8% on a 128k-window model it is
~10k tokens — more than a typical turn injects — so nothing is shed there until
you lower it. The prompt preview API (`GET /api/prompts/preview`) reports the
ceiling, the chars used and what was shed; the inspector renders it in a
follow-up.

Rules of thumb:
- **Recall too thin?** raise `top_k` (more injected) and/or `vector_k` (bigger candidate
pool). **Context too noisy / off-topic?** raise `min_score` to set a relevance floor.
- **Turns too fat?** lower `context.budget_pct` — the low-priority parts (hits, digest,
skill descriptions) give way first; your standing always-on facts never do.
- `rrf_k` rebalances semantic vs keyword — lower lets semantic dominate. Tune it against
the retrieval eval harness rather than by feel.
- Chunking (`chunk_*`) and `contextual_enrichment` are ingest-time knobs — see
Expand Down Expand Up @@ -173,13 +195,14 @@ citations carry the same `trust:` label.

### Hot-memory write visibility (ADR 0069 D8)

`domain="hot"` chunks are injected in front of the model **every turn**, which
makes a silent hot write the highest-leverage poisoning move there is. Since
[ADR 0108 D4](/adr/0108-context-architecture-v2) the same promotion is spelled
out on the row as `delivery_policy="always"` (a hot write is stamped with it
automatically; rows that predate the column were classified once on the first
open after the upgrade), and both controls below cover the policy as well as
the domain. Two controls:
`delivery_policy="always"` chunks are injected in front of the model **every
turn**, which makes a silent always-on write the highest-leverage poisoning move
there is. Always-on is a policy, not a domain
([ADR 0108 D4 + D6](/adr/0108-context-architecture-v2)): a `domain="hot"` write
is stamped with it automatically (rows that predate the column were classified
once on the first open after the upgrade), the per-turn reader selects on the
policy, and a row on any other domain can be pinned always-on the same way.
Both controls below cover the policy as well as the domain. Two controls:

- **Every hot write is a visible event.** Any write that creates a hot chunk —
the agent's `memory_ingest`, the console routes, a plugin via the SDK —
Expand Down
10 changes: 9 additions & 1 deletion docs/reference/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -601,7 +601,15 @@ Only read when `middleware.knowledge` is `true`.

The bundled store is keyword-only FTS5 by default; once your gateway serves `embed_model`, opt in with `embeddings: true` for hybrid search — keyword fused with vector similarity (RRF), with an embedding circuit breaker that falls back to FTS5 on an outage. One `chunks` table; the `domain` column distinguishes operator-set notes (`memory_ingest`), always-on hot facts (`hot`), episodic summaries stored by `conversation_harvest` (`conversation`), and extracted facts (`fact`).

**Hot memory** — chunks stored under `domain='hot'` are *always-on*: `KnowledgeMiddleware` injects them into context every turn (vs. retrieved-on-relevance), re-read each turn so a freshly-added hot fact is seen immediately. Set one with `memory_ingest(content, domain="hot")` for facts the agent should never forget (operator preferences, standing constraints).
**Hot memory** — chunks with `delivery_policy='always'` are *always-on*: the projection injects them into context every turn (vs. retrieved-on-relevance), re-read each turn so a freshly-added fact is seen immediately. A `domain='hot'` write is stamped with that policy automatically ([ADR 0108 D4 + D6](../adr/0108-context-architecture-v2.md)), so `memory_ingest(content, domain="hot")` still pins a fact the agent should never forget (operator preferences, standing constraints) — and so does `delivery_policy="always"` on any domain. Rejected or expired rows never inject.

## `context`

The per-turn injected context — working state, always-on memory, the skill index, the prior-session digest, RAG hits — on top of the stable prompt ([ADR 0108 D6](../adr/0108-context-architecture-v2.md)).

| Key | Default | What |
|---|---|---|
| `budget_pct` | `8` | Ceiling for the injected context as a percentage of the model's context window (chars//4), never below 16 000 chars (room for always-on memory + the digest — on a ≤32k window the floor applies). Over budget the lowest-priority parts shed first — RAG hits, then the prior-session digest, then skill descriptions (skill names never drop); working state and always-on memory are never shed. `0` = unbounded; unbounded too when the gateway reports no window for the model (logged once). The priority order is fixed. The prompt preview API reports the budget and what was shed. |

## `skills`

Expand Down
20 changes: 8 additions & 12 deletions graph/agent.py
Original file line number Diff line number Diff line change
Expand Up @@ -272,23 +272,19 @@ def _build_middleware(
# is active, so skills work even on a KB-less agent (the store is None-tolerant).
_skills_index = skills_index if config.skills_enabled else None
if (config.knowledge_middleware and knowledge_store) or _skills_index is not None:
# ~2% of the model window as CHARS (tokens*4) for the skills index —
# Codex-style ceiling; 8KB when no window is reported (#2867).
try:
from graph.model_window import context_window_for
# ONE wiring (ADR 0108 D6/D8): every delivery knob — RAG top-k, namespace
# scope, trust floor, the skill-index caps (~2% of the model window as chars,
# 8KB when no window is reported, #2867) and the projected-context budget
# (`context.budget_pct`) — is read off the config by
# ProjectionOptions.from_config, the same reader the external runtime uses,
# so the two paths cannot drift.
from graph.projection import ProjectionOptions

_skills_window = context_window_for(config)
except Exception: # noqa: BLE001 — no profile → the 8KB fallback
_skills_window = None
middleware.append(
KnowledgeMiddleware(
knowledge_store if config.knowledge_middleware else None,
top_k=config.knowledge_top_k,
skills_index=_skills_index,
skills_top_k=config.skills_top_k,
skills_index_chars=int(_skills_window * 0.02 * 4) if _skills_window else 8192,
inject_namespaces=config.knowledge_inject_namespaces,
inject_min_trust=config.knowledge_inject_min_trust,
options=ProjectionOptions.from_config(config),
)
)

Expand Down
Loading
Loading