Skip to content

feat: stop recording capability-file history and skip approval bookkeeping in auto - #357

Merged
Max17190 merged 6 commits into
mainfrom
stop-recording-capability-history
Oct 6, 2026
Merged

Max17190 merged 6 commits into
mainfrom
stop-recording-capability-history

Conversation

@Max17190

@Max17190 Max17190 commented Oct 6, 2026 •

Copy link
Copy Markdown
Owner

Why

Every refreeze wrote change records for each tool and skill file into the hash-chained ledger and copied the bytes into its objects store. Every capture hashed and copied those files only to feed that history, and nothing enforced anything against it. In auto, which newly trusted projects now start in, each turn also re-read and hashed the whole approval log at turn start and again for each external tool after every mutating call. Unapproved read-only external tools were also kept out of concurrent batches, even though auto runs them unattended either way. Users paid disk, latency, and lost parallelism for history no one used and approval checks auto never acts on.

Summary

  • The ledger now records approvals only. A refreeze writes nothing to it, and a capture keeps paths plus an in-memory hash of each file's bytes, with no sha256 and no copies on disk.
  • Approvals no longer copy bytes into objects/. An approval is still refused when a bound file changed since its card was shown, now checked against the file's current hash. Stored objects and older change records are kept, never deleted, and the log format is unchanged, so older logs still verify and older binaries keep working.
  • The refreeze receipt lists each tool or skill file as <path> added, modified, or removed, computed by comparing the outgoing and incoming registries. A skill body edit is named as modified, a manifest added past the tool cap is named as added, and a file that is broken or past the cap is never called removed. The receipt no longer attributes an actor or includes clauses about approvals of removed files. Each entry stays on one line even when a path contains a newline.
  • The session manifest now also carries the per-file hashes its freeze read, so a resumed session's receipt names a file edited while it was closed (a SKILL.md body edit included). The field is additive: older manifests still load and compare index lines.
  • In auto, a turn never reads the ledger. Approval narration, the approval check after each mutating call, and the batch content check are all skipped, and unapproved read-only external tools batch like any other read-only call. A session started in auto that later leaves it does not announce approvals made earlier as new.
  • A batch formed under auto runs serially if the user switches to ask before it starts, so the approval card is raised instead of every call being declined silently. A call still queued behind the parallelism cap when the mode leaves auto does not start if its tool is unapproved: it is refused with a reason to request it again, and the retry takes the serial path and raises the card.
  • Approval records, chain verification, and content approval in ask are unchanged.
  • openmax --ledger is deprecated. It prints one line saying history is no longer recorded and naming the project's ledger directory, then exits 0.
  • The frozen prompt drops its --ledger pointer, and the frozen prompt budget cap ratchets from 3,500 to 3,450 bytes (payload is now 3,361, including the mcp surface name).
  • Updated docs/extending.md, docs/configuration.md, docs/stdio-protocol.md, the refrozen event doc, and the stdio --spec text to match.

Test Plan

  • an_auto_turn_batches_unapproved_tools_and_never_reads_the_ledger: a full auto turn reads the ledger zero times and batches unapproved read-only tools (fails on the old code with 14 ledger reads and serial ordering).
  • a_session_built_in_auto_starts_approval_narration_from_what_is_on_record: leaving auto does not announce existing approvals as new.
  • a_skill_body_edit_is_named_as_a_modified_file and the past-cap step in a_tool_pushed_past_the_cap_is_not_a_removed_tool: the receipt names body edits and manifests added past the cap (both fail on the old code).
  • a_batch_formed_in_auto_asks_when_the_mode_turns_to_ask_before_it_runs: switching to ask while an earlier call in the reply runs raises an approval card, and every call ends ok (fails on the old code).
  • a_queued_batch_call_does_not_start_once_the_mode_leaves_auto: with one parallel slot, switching to ask while the first call of an unapproved tool runs keeps the queued second call from starting (fails on the old code with 2 runs and the queued call ending ok).
  • a_skill_body_edited_while_the_session_was_closed_is_named_on_resume: a registry restored from its manifest names a SKILL.md body edit, and a manifest without the hashes still loads (fails on the old code with no change named).
  • a_changed_path_cannot_forge_a_line_in_the_receipt: each receipt entry is one line starting with the project-relative path.
  • an_unapproved_external_tool_batches_only_in_auto: batching depends on the mode, with zero ledger reads in auto.
  • CLI ledger_is_deprecated_and_leaves_existing_records_in_place: --ledger prints one line, exits 0, and leaves records in place (fails on the old code with 5 lines).
  • The chain, pin, and repair tests now run on change records written the way earlier builds wrote them, so older logs still verify.
  • cargo test --workspace --locked exits 0.
  • cargo +1.97.0 clippy --workspace --all-targets --locked -- -D warnings exits 0.

RetriggerConfidence Score: 5/5

The confirmed MCP shutdown issue can leave a subprocess running, but does not block merging.

What we checked:

  • Created and ran a reproducible MCP shutdown fixture to reproduce the P2 finding conditions. T-Rex
  • Captured direct-server shutdown output showing the server process terminated after each CLI exit. T-Rex
  • Captured launcher-child shutdown output documenting the launcher behavior during shutdown. T-Rex
  • Validated the lifecycle behavior across direct and launcher invocations, noting that direct runs terminate the server after CLI exit, launcher runs terminate the launcher but leave the server child alive, and all recorded runs exited with code 0. T-Rex
  • Generated a finding-comment-proof for a posted P2 finding and linked it to the corresponding review comment. T-Rex
Summary

The PR adds a one-shot MCP client alongside changes to approval handling, extension refreeze, session I/O, file tools, and release checks. When an MCP server is started through a launcher, a successful list or call can leave the server running after the CLI exits. The earlier approval-bypass and resumed-skill receipt issues are fixed. The remaining shutdown issue is non-blocking.

Reviews (3) · Last reviewed commit: "Merge remote-tracking branch 'origin/mai..."

…eping in auto

Every refreeze synced change records into the hash-chained ledger and copied
each tool and skill file into its objects store, and every capture hashed and
copied those files only to feed it, yet nothing enforced against that history.
In auto, where newly trusted projects now start, each turn also re-read and
hashed the whole approval log at turn start and once per external tool after
every mutating call, and kept unapproved read-only external tools out of
concurrent batches although auto runs them unattended either way.

The ledger now records approvals only. Captures keep paths, refreezes write
nothing, and approvals no longer copy bytes (a bound file changed since the
card is still refused, by its current hash). The refreeze receipt names each
tool or skill file added, modified, or removed from the two registries, with no
actor attribution, held-generation bookkeeping, or clauses about approvals of
removed files. Under auto a turn never reads the ledger: approval events, the
per-call approval diff, and the batch content check are skipped. Approval
records, their chain verification, and content approval in ask are unchanged.

`--ledger` stays accepted, prints one line saying history is no longer
recorded and where existing records and objects are kept, and exits 0. The log
format is unchanged, so older binaries keep working, and stored objects are
never deleted. The frozen prompt drops its `--ledger` pointer, and its budget
cap ratchets from 3,500 to 3,450 bytes.
The refreeze receipt compared skills by their index line (name and
description), but the fingerprint hashes every SKILL.md byte. A body-only
edit therefore refroze without naming any file, and the receipt fell back
to "extension files changed", so a skill rewritten by a pull arrived
unannounced. A manifest added past the tool cap was not named either.

The capture now keeps a cheap in-memory hash of each file's bytes beside
its path, and the receipt compares two captured generations file by file
from it. A registry restored from a session manifest read nothing, so that
comparison alone falls back to loaded tools and indexed skills. The
fingerprint itself is unchanged, so persisted sessions resume without a
refreeze.

Also drop the unused Actor::as_str, port the one-line-per-entry test to
the new receipt, and correct docs and comments that still described
ledger-recorded changes with actors.
…o first

A reply's tool calls are grouped into concurrent batches once, under the mode of that moment, and auto now lets unapproved read-only external tools into a batch. If the user switched from auto to ask while an earlier call in the same reply was still running, the group still ran as a batch under ask, where the content gate declined every call without raising an approval card.

Each segment now runs as a batch only when the live mode still matches the mode it was formed under; otherwise its calls take the serial path, which asks. Two test texts that described the removed receipt clauses are corrected.
Comment thread crates/core/src/agent.rs
Comment thread crates/core/src/agent.rs
…s auto

A concurrent batch admits every call up front, then starts them behind the
parallelism cap, so a call can wait in the queue long after auto admitted
it. Auto never asks whether a tool's content is approved, so if the user
switched to ask while an unapproved external tool waited, it still started,
under ask, with no approval card.

Each call now re-reads the mode as it starts. A call that auto admitted
without the content question, whose tool is unapproved, is refused once the
mode has left auto, with a reason that asks the model to request it again;
requested again, it takes the serial path, which raises the card.
A resumed session's registry is rebuilt from its manifest, which kept no
record of the bytes its freeze read. The refreeze receipt could then only
compare index lines, so a SKILL.md body edited while the session was closed
went live under a generic "extension files changed" and the file was never
named.

The manifest now carries the per-file hashes the freeze read, and a resumed
registry compares file by file like a live one. The field is additive: a
manifest written before it still loads and compares index lines.
…lity-history

The extension pointer keeps main's mcp surface and still drops the ledger history pointer; the frozen payload is 3,361 bytes under the 3,450 cap.
@greptile-apps

greptile-apps Bot commented Oct 6, 2026

Copy link
Copy Markdown

Comments Outside Diff

These findings could not be posted inline.

  • P2 MCP server survives shutdown crates/tui/src/mcp.rs:429 ▶

    If an MCP launcher spawns a server subprocess, --mcp-list and --mcp-call stop and wait for only the launcher. The CLI can exit successfully while the server keeps running and holding resources. Shut down the launched process tree, not just its direct child.

@Max17190
Max17190 merged commit 876039a into main Oct 6, 2026
16 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant