Skip to content

Swarm tranches: agentic depth, measured perf, reliability, UX/AX - #2

Merged
RasputinKaiser merged 1 commit into
mainfrom
swarm-agentic-perf-reliability
Sep 23, 2026
Merged

RasputinKaiser merged 1 commit into
mainfrom
swarm-agentic-perf-reliability

Conversation

@RasputinKaiser

Copy link
Copy Markdown
Owner

Summary

Six swarm tranches of work on Hemlock — all gates green (521/521 host · 288/291 UI · 63/63 test_server.py · build clean).

Performance — decode ~30.6 → ~75–90 tok/s (live-measured): shared per-step position array kills 24 GPU syncs/token, fused MoE Metal kernels replace the gather_qmm dispatch pair, B==1 sampler fast path. Opt-in --kv-bits lifts the practical context ceiling 24k → 61k. Persistent multi-entry prompt cache + warm prefill = cache-hot first steps.

Agent harness — batch action envelopes, plan.revise, bounded shell.exec, memory.note, experiment.suggest, world.place, agent.self/world.state, lane fallback with honest fallbackFrom receipts, input-repair retry, decide confidence+margin gates, channel rejoin. Dream trains on durable agentic action traces with a dataset preview; idle-dream proposes (never starts) cycles.

Reliability — durable_io atomic writes + quarantine recovery, wedged-TCC-dir boot freeze fixed via detached probe child, queue drain/leapfrog/restore fixes, health monitor + bounded auto-resume, pending intents persist across restarts.

UX/AX — Threads window (search/checkpoints/fork), notification center, status chips, dock badges, window cycling + focus handoff, operate-mode hardening.

Test plan

  • npm run test:agent — 521/521
  • npm run test:ui — 288/291 (3 pre-existing skips)
  • pytest tests/test_server.py — 63/63
  • npm run build — clean
  • Live-verified on maple-2bit-mlx: prompt-cache round-trip, fused-kernel A/B

Generated with Devin

Performance (all live-measured on maple-2bit-mlx):
- Decode ~30.6 -> ~75-90 tok/s: shared per-step (pos,eps) array removes
  24 GPU syncs/token, fused MoE kernels (maple_moe_upgate/down) replace
  gather_qmm dispatch pair on decode, GenerationBatch B==1 fast path.
- Opt-in --kv-bits quantizes the 6 full-attn KV caches past a start
  token; practical context ceiling 24k -> 61k (sliding layers stay exact).
- Persistent multi-entry prompt cache (manifest + per-entry safetensors,
  atomic commit) + warm prefill on server-ready/thread-switch.
- Checkpoint maple.py synced to repo impl — served code was stale.

Agent harness:
- Batch {actions:[...]} envelopes, plan.revise, bounded shell.exec,
  memory.note, experiment.suggest, world.place, agent.self/world.state
  introspection, lane fallback (receipted fallbackFrom), input-repair
  retry, decide margin+confidence gates, channel rejoin for split
  content/reasoning envelopes, failure-hint observations.
- Dream: agentic action-trace SFT rows, dream.dataset.preview, idle-dream
  proposals, dream:start budget fix.
- Threads window: search, checkpoints, fork, lifecycle — all via host
  commands.

Reliability:
- durable_io: atomic writes, quarantine + integrity.recovered on
  corruption; context_broker wedged-TCC-dir boot freeze fixed via
  detached probe child; queue drain/leapfrog/restore-payload fixes;
  health monitor + bounded auto-resume; pending intents persist.

UI: window extraction, ctxSlices/eventBuffer, notification center,
status chips, dock badges, window cycling, honest busy/disabled states,
Escape/focus hygiene, operate-mode hardening.

Eval: tool-use benchmark 11->27 tasks + calibration artifacts;
e2e_autonomy_surface + e2e_soak against real main.cjs.

Verified: 521/521 host - 288/291 UI - 63/63 test_server.py - build clean.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@RasputinKaiser
RasputinKaiser merged commit 4807ec5 into main Sep 23, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant