feat(bin): add Kimi crewmate adapter and agent-owned dispatch routing - #1
Open
Vhailors wants to merge 12 commits into
Open
feat(bin): add Kimi crewmate adapter and agent-owned dispatch routing#1Vhailors wants to merge 12 commits into
Vhailors wants to merge 12 commits into
Conversation
* Replace quota dispatch selector instructions * no-mistakes(review): Align bootstrap docs with agent-owned dispatch selection
* remove vestigial dispatch selector * no-mistakes(review): Synchronize isolation proof and portable shard evidence * no-mistakes(review): Correct shard history and proof archive date * no-mistakes(review): Remove reintroduced selector documentation reference * no-mistakes(document): Remove stale dispatch strategy documentation
…1039) quota-axi 0.1.13 emits schemaVersion 2 with a quotaSemantics object per provider, so the successor named in the interim rule has landed and the rule's own removal condition is satisfied. Keep the ownership clause so quota-axi remains the single owner of how model or product windows relate to bounding account windows, and drop the interim weakest-headroom instruction. The unknown-semantics case is already covered by the existing requirement to stop and report a candidate whose applicable quota data or interpretation cannot be established. Drop the matching assertion phrase from tests/fm-instruction-owners.test.sh; the retained ownership phrase still asserts.
…unchenguid#1049) * fix(tmux): scope Claude busy detection by harness * no-mistakes(review): Separate verified and fallback busy signatures * no-mistakes(test): Scope busy signatures to supplied harnesses * no-mistakes(document): Document harness-scoped busy detection
* Add verified Kimi crewmate harness adapter * no-mistakes(review): Scope Kimi moon detection to spinner lines * no-mistakes(review): Match only complete Kimi spinner rows * no-mistakes(review): Resolve Kimi binary portably before pane creation * no-mistakes(document): Align Kimi adapter documentation * no-mistakes(lint): Suppress false-positive ShellCheck warning for sourced watcher override * Fix Kimi busy spinner detection * no-mistakes(review): Recognize Kimi session-lock ancestry and holders * no-mistakes(review): Scope pending-reply Kimi busy detection by harness * no-mistakes(document): Correct Kimi spinner capture documentation * no-mistakes(document): Clarify optional Kimi spinner whitespace * no-mistakes(lint): Silence intentional pending-reply test stub warnings * test: align rebased Kimi busy fixtures * no-mistakes: apply CI fixes * Reconcile Kimi busy detection after per-harness scoping * no-mistakes(review): Clarify observed Kimi spinner whitespace contract * no-mistakes(document): Clarify Kimi harness documentation
* fix kimi pointer submission and spinner conformance * no-mistakes(review): Preserve Kimi submit target ownership guard
* Add guarded Kimi turn-end hook * no-mistakes(review): Require jq before installing Kimi turn-end hook * no-mistakes(review): Expose jq inside isolated Kimi test fixtures * no-mistakes(review): Preserve Kimi config boundaries during hook removal * no-mistakes(review): Document Kimi removal newline safeguard * no-mistakes(document): Document Kimi shared-home preservation
) * Fix structural tmux composer reading * Verify Calm compatibility with Pi 0.82 * no-mistakes(review): Harden structural composer classification boundaries * no-mistakes(review): Refresh composer and Kimi regression fixtures * no-mistakes(review): Fail closed on unbounded composer edges * no-mistakes(review): Enforce aligned composer geometry safely * no-mistakes(review): Make composer ambiguity locale-safe * no-mistakes(review): Preserve ambiguity through composer submission * no-mistakes(review): Carry composer proof through retries * no-mistakes(document): Document structural tmux composer delivery guarantees * no-mistakes: apply CI fixes
…lity in owner doc
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Implement the captain-authorized future Firstmate routing correction on the latest upstream-compatible agent-owned quota-selection design. Preserve all active homes, current task metadata, Herdr runtime and presentation-space behavior, and SceneAxi's existing override. Keep persistent secondmates pinned to Pi openai-codex/gpt-5.6-luna at high effort; set future implementation profiles to direct Grok Build grok-4.5, Claude Code opus, Codex gpt-5.6-sol, and Pi openai-codex/gpt-5.6-sol at high effort. Bug-bounty and security-engagement crewmates must hard-pin Pi zai/glm-5.2 high and report a capacity blocker rather than falling back to GLM 5.1, Opus, Sol, Fable, or random selection. Stale, unavailable, or unscorable quota must fail closed under upstream's agent-owned quota procedure. Do not resurrect or broaden the bin/fm-dispatch-select.sh script removed upstream in 3f71cdd, and preserve compatibility for existing homes and static pins. Keep the port to the smallest policy, config example, documentation, and focused test changes. Publish only through this mandatory no-mistakes pipeline after its checks are green; do not merge.
What Changed
kimias a verified crewmate harness:bin/fm-spawn.shgains Kimi binary resolution, a bare launch path with composer-based prompt delivery and readiness polling, plus the newbin/fm-kimi-turnend-hook.shguarded turn-end wake that is installed at spawn (refusing the spawn if it cannot be installed safely) and cleaned up bybin/fm-teardown.sh.bin/fm-tmux-lib.shand the composer/send path: bordered composers are now classified across all pane rows, busy detection is scoped so a crewmate's current turn is recognized instead of being read as idle, with matching updates infm-send.sh,fm-composer-lib.sh,fm-supervise-daemon.sh, andfm-watch.sh.bin/fm-dispatch-select.shand its suite, and moved routing policy intoAGENTS.md— precedence now runs captain override → bug-bounty/security hard pin (Pizai/glm-5.2high, reporting a capacity blocker instead of falling back) → configured rule → configured default → static harness, with quota evaluated per candidate and stale/unscorable quota failing closed.docs/examples/crew-dispatch.json,docs/configuration.md, and the harness-adapters skill were updated to match, andtests/fm-instruction-owners.test.shpins the new statements.Risk Assessment
✅ Low: Documentation, config-example, and test-only change that touches no shell code, satisfies every source-verifiable intent criterion (secondmate Luna pin, the four standing high-effort profiles, the GLM 5.2 hard pin with capacity-blocker semantics, fail-closed quota, no resurrected selector, preserved backward compatibility), and whose example validates cleanly against the existing bootstrap harness/effort matrix; the only residual issues are two narrow gaps in the new test's own guard coverage.
Testing
Ran the branch's mapped contract-test family plus a real operator walkthrough: copying the shipped crew-dispatch example into a firstmate home and running the actual bootstrap validator shows it is accepted and resolves the security engagement to a single pi/zai/glm-5.2/high hard pin with no fallback candidate, the implementation rule and default to the four standing high-effort profiles, and the real harness resolver reads the persistent-secondmate Pi Luna pin back correctly, with the upstream-removed dispatch selector still gone. Nine mutations against throwaway tree copies confirm each intent constraint is genuinely enforced, and I added one focused bootstrap test (mutation-checked) closing the gap where the example the docs tell captains to copy was never verified against the real validator. Two test scripts fail in this environment for reasons that predate the change and reproduce at the branch parent - the bootstrap orca case because /usr/bin/orca (the GNOME screen reader) exists on this Linux box, and a tmux-driven fm-calm-pi e2e case that names a different sub-case on each run - plus the coverage guard is locale-sensitive and passes under LC_ALL=C; none involve the routing policy.
Evidence: Operator walkthrough: shipped example accepted by the real bootstrap validator, resolved routing table, secondmate pin readback, selector still removed
### 2. Real bootstrap validation - a valid dispatch config stays silent $ bin/fm-bootstrap.sh (no output - config accepted, no CREW_DISPATCH error) ### 3. Same run with facts on - the routing table firstmate will apply $ FM_BOOTSTRAP_VERBOSE_FACTS=1 bin/fm-bootstrap.sh BOOTSTRAP_INFO: crew dispatch active config/crew-dispatch.json BOOTSTRAP_INFO: crew dispatch rule: The task is a bug-bounty or security engagement. -> pi/zai/glm-5.2/high BOOTSTRAP_INFO: crew dispatch rule: The task depends on fresh news, current events, live public facts, or recent market and product changes. -> grok/grok-4.5/high BOOTSTRAP_INFO: crew dispatch rule: The task is a genuinely trivial mechanical edit ... -> grok/grok-4.5/high BOOTSTRAP_INFO: crew dispatch rule: The task is a big or ambiguous multi-file feature ... -> quota-balanced[grok/grok-4.5/high, claude/opus/high, codex/gpt-5.6-sol/high, pi/openai-codex/gpt-5.6-sol/high] BOOTSTRAP_INFO: crew dispatch default: quota-balanced[codex/gpt-5.6-sol/high, pi/openai-codex/gpt-5.6-sol/high, claude/opus/high, grok/grok-4.5/high] ### 4. Persistent secondmate pin read back through the real resolver $ cat config/secondmate-harness pi openai-codex/gpt-5.6-luna high $ bin/fm-harness.sh secondmate pi $ bin/fm-harness.sh secondmate-model openai-codex/gpt-5.6-luna $ bin/fm-harness.sh secondmate-effort high ### 5. The selector removed upstream in 3f71cdd stays removed $ ls bin/fm-dispatch-select.sh ls: cannot access '.../bin/fm-dispatch-select.sh': No such file or directory $ grep -rl fm-dispatch-select . (excluding .git) tests/fm-instruction-owners.test.sh (the only remaining mention is the regression guard asserting AGENTS.md never resurrects it)Evidence: Mutation proof: every routing constraint in the intent is enforced, not just asserted
== baseline: unmodified tree == exit=0 ok - captain routing profiles remain agent-owned and keep persistent secondmates pinned ### security pin downgraded to GLM 5.1 in the shipped example exit=1 not ok - dispatch example lost the GLM 5.2 hard pin, a standing implementation default, or retained GLM 5.1 ### security hard pin removed from AGENTS.md exit=1 not ok - captain routing contract lost 'hard-pin Pizai/glm-5.2at high effort' ### GLM 5.1 / Opus / Sol / Fable fallback prohibition removed from AGENTS.md exit=1 not ok - captain routing contract lost 'never substitute GLM 5.1, Opus, Sol, Fable, or a random candidate' ### stale/unscorable-quota fail-closed rule removed from AGENTS.md exit=1 not ok - captain routing contract lost 'Stale, unavailable, or unscorable quota is unresolved evidence' ### standing implementation profiles removed from AGENTS.md exit=1 not ok - captain routing contract lost 'direct Grok Buildgrok-4.5, Claudeopus, Codexgpt-5.6-sol, and Piopenai-codex/gpt-5.6-sol' ### one standing implementation profile (Claude opus) dropped from the example default set exit=1 not ok - dispatch example lost the GLM 5.2 hard pin, a standing implementation default, or retained GLM 5.1 ### example security rule replaced by a random-fallback array exit=1 not ok - dispatch example lost the GLM 5.2 hard pin, a standing implementation default, or retained GLM 5.1 ### persistent secondmate Pi Luna pin removed exit=1 not ok - persistent secondmate Pi Luna pin is missing ### removed fm-dispatch-select.sh selector resurrected in AGENTS.md exit=1 not ok - agent-owned routing policy resurrected the removed selectorEvidence: New focused test (shipped example vs real validator) with its mutation check
### baseline - shipped example ok - shipped crew-dispatch example validates and renders the captain-authorized pins ### mutation: security pin downgraded to zai/glm-5.1 not ok - shipped example lost the single-profile security hard pin (missing: 'security engagement. -> pi/zai/glm-5.2/high') ### mutation: grok profile given an effort the adapter rejects (xhigh) not ok - shipped dispatch example must pass bootstrap validation, got: CREW_DISPATCH: invalid config/crew-dispatch.json - invalid effort: grok:xhighEvidence: Reproducible operator-walkthrough script (evidence generator)
Evidence: Reproducible mutation-proof script (evidence generator)
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
docs/examples/crew-dispatch.json:15- docs/examples/crew-dispatch.json:15 keeps the rule whosewhendescribes "a trivial mechanical edit such as a rote rename, formatting sweep, targeted typo fix" but now resolves it to claude/opus/high, and replaces the old rationale with "Claude Code implementation work uses Opus at high effort." The rule's selection category (cheapest fast profile for narrow, low-ambiguity work) no longer has any distinct effect - it resolves identically to the standing implementation profile - while AGENTS.md:180 and harness-adapters SKILL.md:111 still describelowas the intended level for well-understood bounded work (as a fallback that configured effort legitimately overrides, so this is not a functional conflict). Consider whether this rule should be dropped from the example or itswhen/whyreworded, rather than advertising the most expensive profile for the trivially-scoped category.AGENTS.md:166- AGENTS.md:166 still states routing precedence as "an explicit per-task captain override, then the best-fit configured rule, then the configured default, then the static crewmate harness" with no security exception, while the new AGENTS.md:174 requires security/bug-bounty crewmates to "bypass ordinary profile-array selection" and hard-pin Pi zai/glm-5.2. A local crew-dispatch.json with no security rule but a broad matching rule (e.g. the "big or ambiguous multi-file feature" array in the example) leaves an agent reading the precedence line first with a configured-rule match that appears to win - exactly the array selection the hard pin is meant to bypass. Since enforcement here is purely textual, a clause on line 166 noting that the security hard pin outranks configured rules would close the ambiguity.tests/fm-instruction-owners.test.sh:156- tests/fm-instruction-owners.test.sh:156 invokesjqunguarded and funnels any nonzero exit intofail "dispatch example lost the GLM 5.2 hard pin or retained GLM 5.1". On a host without jq (which the repo treats as an optional dependency - bin/fm-bootstrap.sh:707 reportsMISSING: jqthrough install-consent), this test reports a false policy regression instead of a missing tool. Repo convention is either a skip guard (tests/fm-crew-state.test.sh:795) or an explicit jq-specific failure (tests/fm-bootstrap.test.sh:115); adopt one of those.tests/fm-instruction-owners.test.sh:158- The new jq assertion in tests/fm-instruction-owners.test.sh:156-164 pins only the bug-bounty rule and the absence of zai/glm-5.1; the four standing implementation profiles that this change actually rewrote in the example (grok/grok-4.5, claude/opus, codex/gpt-5.6-sol, pi/openai-codex/gpt-5.6-sol, all at high) are asserted only as an AGENTS.md prose phrase. The example file could silently drift back to weaker or older profiles - the exact regression this change corrects - and the suite would still pass. Extend the same jq check to assert those profiles in.default.🔧 Fix: route trivial-edit rule to Grok, clarify security pin precedence
2 infos still open:
AGENTS.md:173- AGENTS.md:173 says the standing high-effort profiles yield to "a more specific project rule or explicit captain override" but is silent about the configureddefault, which the precedence chain on line 166 lists as its own tier. An existing home whoseconfig/crew-dispatch.jsoncarries an olderdefault(e.g.{"harness":"codex","model":"gpt-5.5","effort":"medium"}) leaves an agent with two defensible readings: follow the configured default per line 166, or substitute the standing Sol/Opus profile per line 173. The intent requires preserving compatibility for existing homes and static pins, which resolves the ambiguity in favor of the configured default, so adding "or configured default" to the exception clause on line 173 closes it. Enforcement here is purely textual, so precision is the only mechanism.docs/examples/crew-dispatch.json:16- The trivial-edit rule'swhyreads "Direct Grok Build takes narrow low-ambiguity mechanical work; Opus at high effort stays the standing profile for Claude implementation and design work, and Sol at high effort for Codex and Pi." Per docs/configuration.md:217,whyis rationale "that helps firstmate choose" this rule, and firstmate matches rules with judgment overwhenpluswhy. Two thirds of thiswhydescribes Claude/Codex/Pi implementation and design work - the category this rule is explicitly not for (itswhenrequires "no design or ambiguity to resolve") - so a design-heavy Claude implementation task can plausibly best-fit against this rule's text and get routed to Grok. The standing-profile statement already lives in its authoritative place at AGENTS.md:173; trimming thewhyto the Grok rationale alone would keep the rule's selection signal aligned with itswhen.🔧 Fix: scope trivial-edit rule rationale to its own category
2 infos still open:
tests/fm-instruction-owners.test.sh:146- The new phrase loop at tests/fm-instruction-owners.test.sh:146-153 pins five AGENTS.md statements but omits the precedence clause added at AGENTS.md:166 ("then the bug-bounty and security-engagement hard pin below"). That clause is the only text stating the hard pin outranks a matching configured rule; the hard-pin sentence itself (AGENTS.md:174) says only "bypass ordinary profile-array selection", which does not tell an agent what to do when a broad configured rule such as the example's "big or ambiguous multi-file feature" array matches a security task first. This file's own header (line 3) states it exists to guard these statements across "the AGENTS.md reduction pass", and every other statement this change added is pinned, so the precedence clause is the single unprotected one. Add 'then the bug-bounty and security-engagement hard pin below' to the same loop.tests/fm-instruction-owners.test.sh:177- tests/fm-instruction-owners.test.sh:177 asserts only that the string 'fm-dispatch-select' is absent from AGENTS.md, but its failure message is "agent-owned routing policy resurrected the removed selector" and the intent's forbidden behavior is resurrecting bin/fm-dispatch-select.sh (removed upstream in 3f71cdd). A re-added bin/fm-dispatch-select.sh that is wired through docs/configuration.md or fm-spawn.sh without any AGENTS.md mention passes this assertion unchanged, so the guard does not detect the condition it names. Add assert_absent "$ROOT/bin/fm-dispatch-select.sh" alongside the existing prose check.tests/fm-bootstrap.test.sh:351- Pre-existing environment failure, not caused by this change: tests/fm-bootstrap.test.sh::test_orca_backend_gates_orca_tool_only_when_selected assumesorcais absent from the base PATH, but /usr/bin/orca (the GNOME screen reader) ships on this Linux box, so bootstrap never emits the expected MISSING: orca line. This aborts the bootstrap suite before later cases, so I exercised the newly added dispatch-example case in a throwaway copy with that one case skipped. Reproduces identically at branch parent a5fe1bc. Fixing it would mean shadowing orca in the fake toolchain, which is outside this change's scope.tests/fm-calm-pi-extension.test.sh:1091- Pre-existing flaky test, not caused by this change: tests/fm-calm-pi-extension.test.sh fails its live tmux Pi follow-up e2e with "rendered a duplicate captain answer", but names a different sub-case on each run (absent, legacy_away, loaded_off, exact_watcher), which points at a timing race in the capture-pane polling rather than a deterministic defect. It fails the same way at branch parent a5fe1bc, and the branch touches no .pi extension or runtime code.bin/fm-test-run.sh:392- Pre-existing locale sensitivity, not caused by this change: run_coverage_guard writes its intermediate lists withLC_ALL=C sortbut invokescommunder the ambient locale, so under en_US.UTF-8 comm reports "file 2 is not in sorted order" and --check-coverage exits 1 (taking tests/fm-test-run.test.sh with it).LC_ALL=C bin/fm-test-run.sh --check-coverageprints FM_TEST_COVERAGE ok. Worth a one-line locale pin on the comm calls in a separate change.bash bin/fm-test-run.sh tests/fm-instruction-owners.test.sh- all 12 contract assertions green, including the new captain-routing casebash bin/fm-test-run.sh --changed --base a5fe1bc- the runner's own changed-file mapping selects the pure-contract-unit family (31 scripts) for these files; 29 green, 2 pre-existing failures unrelated to the changeManual operator walkthrough: copieddocs/examples/crew-dispatch.jsoninto a firstmate home config and ran the realbin/fm-bootstrap.sh(silent = accepted) andFM_BOOTSTRAP_VERBOSE_FACTS=1 bin/fm-bootstrap.sh(renders the resolved routing table)bin/fm-harness.sh secondmate/secondmate-model/secondmate-effortagainst a home pinned topi openai-codex/gpt-5.6-luna highMutation proof (throwaway tree copies, worktree untouched): GLM 5.1 downgrade in the example, security hard pin deleted from AGENTS.md, GLM 5.1/Opus/Sol/Fable prohibition deleted, stale-quota fail-closed rule deleted, standing implementation profiles deleted, Claude opus dropped from the default set, security rule replaced by a random-fallback array, Luna secondmate pin deleted, andfm-dispatch-select.shresurrected - each turnstests/fm-instruction-owners.test.shredAddedtests/fm-bootstrap.test.sh::test_shipped_dispatch_example_is_accepted_and_keeps_its_pins, then mutation-checked it (GLM 5.1 downgrade and a grokxhigheffort the adapter rejects both fail it)ls bin/fm-dispatch-select.shandgrep -rl fm-dispatch-select- selector absent, only remaining mention is the regression guardPre-existing-failure isolation: rantests/fm-calm-pi-extension.test.shandtests/fm-bootstrap.test.shagainst agit archive a5fe1bcextract of the branch parent; both fail there toodocs/configuration.md:176- docs/configuration.md:176 claims "claude, codex, opencode, pi, grok, and kimi are empirically verified for crewmate and secondmate launches", but every other evidence surface scopes Kimi to crewmates only: the upstream commit is titled "add verified Kimi crewmate adapter", .agents/skills/harness-adapters/SKILL.md states Kimi is outside the primary turn-end guard scope, docs/turnend-guard.md documents its hook as a crew wake integration, and docs/supervision-protocols/ has no Kimi primary protocol (secondmates run as primaries in their own home). I could not establish whether upstream verified Kimi secondmate launches and left the sentence unchanged rather than unilaterally narrowing a supported-scope claim.✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.