Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
15 commits
Select commit Hold shift + click to select a range
d47f84e
docs(add): fold v2-performance deltas -> foundation-version 2
tindangtts Jun 15, 2026
4b8a7d3
chore(add): scaffold ft-search-off-eventloop task (v2 criterion 2)
tindangtts Jun 15, 2026
6c263e5
docs(add): ft-search-off-eventloop §1 SPECIFY (cooperative-yield fram…
tindangtts Jun 15, 2026
049ffab
docs(add): ft-search-off-eventloop §2 SCENARIOS (10 Given/When/Then)
tindangtts Jun 15, 2026
6bbb0c7
docs(add): ft-search-off-eventloop §3 CONTRACT FROZEN @ v1
tindangtts Jun 15, 2026
cac32ee
test(ft-search-off-eventloop): §4 red suite — yield-seam shape + beha…
tindangtts Jun 15, 2026
785b204
chore(add): ft-search-off-eventloop advance tests->build (red confirmed)
tindangtts Jun 15, 2026
1ec2909
feat(vector): ft-search-off-eventloop BUILD batch 1 — yielding seam core
tindangtts Jun 15, 2026
6a8272d
feat(vector): ft-search-off-eventloop BUILD batch 2 — handler wiring
tindangtts Jun 15, 2026
7c4f8cd
fix(vector): make FT.SEARCH cooperative yield effective on monoio (ti…
tindangtts Jun 15, 2026
446a0ee
docs(add): close ft-search-off-eventloop §7 OBSERVE + reconcile v2 mi…
tindangtts Jun 15, 2026
0e1a91c
chore(add): close v2-performance milestone (3/3 tasks PASS)
tindangtts Jun 15, 2026
461dd43
docs(add): fold ft-search-off-eventloop deltas -> foundation-version 3
tindangtts Jun 15, 2026
2457a4b
Merge branch 'main' into feat/ft-search-off-eventloop
pilotspacex-byte Jun 15, 2026
fe74ab8
docs(changelog): add FT.SEARCH off-event-loop entry (PR #179)
tindangtts Jun 15, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions .add/CONVENTIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,3 +10,13 @@ Architecture: thread-per-core shared-nothing shards; per-shard locks only (parki
Verification: red/green TDD per task; full local CI parity via OrbStack `moon-dev` before push; benchmark numbers from Linux VM only; new parsers get fuzz targets; new atomic state machines get loom models
Red-suite shape (behavior-preserving perf work — foundation v1): split the failing-first suite into runtime-red (behavioral asserts) + compile-red (an API-shape file that must fail to compile until the new surface exists) + green pins (invariants that must stay green). Keeps a perf refactor honest without a behavior oracle. (lesson: quickwins_red.rs / quickwins_red_api.rs.)
ADD freeze unit (foundation v1): the §3 CONTRACT freeze requires the literal line `Least-sure flag surfaced at freeze: …` — the template comment alone is not a declaration; the engine refuses `advance` (unflagged_freeze) without it.
Frozen red test may be wrong (TDD, foundation v2): a frozen RED test is not sacrosanct — if it has a harness bug (e.g. double-consuming a flume `bounded(1)` single-use channel), fix it intent-preserving WITH human sign-off after confirming the implementation is actually correct; never weaken the assertion to make the build pass. (lesson: `commit_write_fail_acks_write_failed`, wal-group-commit.)
Whole-repo symbol-removal grep (TDD, foundation v2): a "symbol hard-removed repo-wide" shape test must grep the WHOLE tree (`src/` + `tests/` + `scripts/` + `benches/`), not just `src/` — a src-only scan let 15 test files + 1 script keep dangling refs that broke the full build invisibly. (lesson: `xshard_cleanup_shape`.)
VM integration binary pin (TDD, foundation v2): server-spawning integration tests must pin `MOON_BIN` (or run in-tree) on the OrbStack VM — `find_moon_binary`'s `{manifest}/target/release/moon` fallback resolves a macOS Mach-O / stale binary under an external `CARGO_TARGET_DIR`, yielding phantom "server never accepted" failures (reinforces the Mach-O binary trap).
Perf anchor must sweep the pipelined regime (TDD, foundation v2): a perf anchor measures connection-count AND pipeline depth (a synchronous spin can serialize a pipelined fan-out invisibly at c1/c100), with a flat single-shard CONTROL cell and best-of-7 to separate signal from VM drift. (lesson: xshard P16 −27.5% hidden by a c1/c100-only best-of-3 anchor.)
Confirm instrument validity before a perf Must (ADD, foundation v2): an empirical perf Must can be un-measurable on the only available instrument — the OrbStack VM's near-free virtio fsync makes fsync-bound wins (group commit) structurally invisible (writer drains before the next arrives ⇒ batch≈1; `always`≈0.9M RPS, no 11× penalty). Validate the instrument can resolve the metric BEFORE anchoring the Must; fsync-bound / real-disk numbers need a real disk or GCloud, not the VM. (lesson: wal-group-commit, §1 assumption #4 confirmed.)
Full-dual-runtime gate for deletions (ADD, foundation v2): a symbol removal's honest gate is a full `cargo test` on BOTH runtimes — scoped `cargo test --test X` runs give FALSE GREEN for cross-cutting deletions (a break hid through every scoped run, surfaced only at dual-runtime verify).
At-BUILD safety audit (ADD, foundation v2): run `scripts/audit-unsafe.sh` / `scripts/audit-unwrap.sh` during BUILD, not just verify — a new `unsafe`/`unwrap` slipping from build to verify costs a phase; an at-build audit catches it earlier.
Mechanism-proxy pass ≠ effect measured (TDD, foundation v3): a perf Must needs TWO distinct gates — a mechanism-fired proxy (a counter proving the code path ran) AND an effect-measured benchmark (proving the path achieved the intended effect). They are NOT the same test: the proxy can be green while the effect is entirely absent. (lesson: ft-search-off-eventloop — the m1 `ft_search_cooperative_yields_total` counter was green on monoio while the yield relieved zero co-located latency; only the §6 benchmark caught the no-op.)
Per-runtime EFFECTIVENESS validation, not just compile+correctness (ADD, foundation v3): when a contract guarantee rests on runtime scheduler behavior, a `#[cfg]`-split primitive that COMPILES and is CORRECT on both runtimes can still be EFFECTIVE on only one — the verify plan must MEASURE the behavior on EACH runtime, never infer parity from shared code. (lesson: ft-search-off-eventloop — identical self-wake `cooperative_yield` gave tokio p99 6ms but monoio 68ms; monoio's io_uring loop never reaps the CQ under a self-waking task. Fix: runtime-split — monoio `sleep(ZERO)` timer-park.)
Make the instrument work before deferring the measurement (ADD, foundation v3): when a GATE-DEFER would HIDE a real defect, prefer fixing the instrument (clean disk, quiesce the VM, add a deterministic proxy) over deferring — a heavier verify caught a default-runtime no-op that the easy disk-full defer would have shipped. (Complements "confirm instrument validity before a perf Must": validity-confirm picks the anchor; this says don't defer past a defect the anchor CAN resolve. lesson: ft-search-off-eventloop.)
6 changes: 5 additions & 1 deletion .add/PROJECT.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@
> UI/UX = UDD. When a loop reveals a gap here, come back and update this file —
> that is the re-entrant arrow from the engine down to the foundation.

slug: moon · stage: production · updated: 2026-06-13 · foundation-version: 1
slug: moon · stage: production · updated: 2026-06-15 · foundation-version: 3
goal: a Redis-compatible server whose thread-per-core architecture measurably out-scales Redis on multi-core hardware — without sacrificing protocol compatibility or durability semantics

---
Expand All @@ -27,6 +27,8 @@ goal: a Redis-compatible server whose thread-per-core architecture measurably ou
- Active milestone → `.add/milestones/v1-shared-nothing/MILESTONE.md` (see `add.py status`)
- Frozen contracts (living docs): RESP2/RESP3 wire compatibility with Redis (external, immutable); CI matrix (fmt, clippy ×2, tests ×2, MSRV 1.94, unsafe/unwrap audits, fuzz)
- Settled vs still open: settled — thread-per-core + SPSC mesh architecture, monoio default on Linux. Open — sub-linear multi-shard scaling (root causes mapped in 2026-06 review: leaky shared-nothing + 1ms monoio wake floor)
- [foundation v2, 2026-06-15] A contract invariant that quantifies over "all N implementations" must be verified against EACH one, not assumed uniform: group commit's `CommitOutcome.write_failed ⇒ write_error latch` held in 3 of the 4 AOF writer loops, but the tokio-TopLevel loop never carried the latch (a pre-existing gap the new contract made explicit — Finding 2, wal-group-commit). When a spec says "both writers / all loops", enumerate and check each.
- [foundation v2, 2026-06-15] Keep REJECTED-risk flags IN the frozen §3 contract, not just in discussion: a pre-named, pre-reasoned risk (xshard synchronous-spin serializing pipelined reads) was the exact failure that materialized at verify — naming it at freeze turned a surprise −27.5% P16 regression into a targeted batch-depth-gate fix instead of a redesign (xshard-read-fastpath).

## Users (UDD) — UI/UX: design before code
- No UI — surface is the **Redis wire protocol** (RESP2/RESP3) plus CLI flags (`--port --shards --appendonly --dir …`) and INFO/Prometheus metrics.
Expand All @@ -45,3 +47,5 @@ goal: a Redis-compatible server whose thread-per-core architecture measurably ou
| 2026-06-11 | FT.SEARCH off-event-loop + WAL group commit deferred to v2 | different themes (event-loop blocking; durability); keep v1 one outcome | recorded in v1 Out list |
| 2026-06-13 | CLOSE v1-shared-nothing: shared-nothing restored (locks deleted, shape-enforced), 1ms monoio wake floor gone (cross-shard p99 0.071ms), consistency 197/197 @1/4/12; s4 routed parity-or-better (+12% P16 GET) vs v0.3.0 | exit criteria met to the agreed "no-regression + honest measurement" bar | done; default-config cross-shard read regression (−85% c1 GET) RISK-ACCEPTED → follow-up: lock-free cross-shard read acceleration (waiver → next perf milestone) |
| 2026-06-13 | fold v1 deltas → foundation-version 1 | close the ADD loop so learnings outlive the milestone | DDD: lock-inventory grep → CI (PROJECT §Domain); TDD: red-suite split pattern + ADD: §3 freeze flag-line requirement (CONVENTIONS) |
| 2026-06-15 | fold v2 deltas → foundation-version 2 (9 deltas from xshard-read-fastpath + wal-group-commit) | close the loop after the first 2 v2-performance tasks; perf-measurement + cross-cutting-deletion lessons recur | SDD: "verify each impl of an all-N invariant" + "keep rejected-risk flags in the freeze" (PROJECT §Spec); TDD: whole-repo symbol-removal grep, MOON_BIN-pinned VM integration, pipelined+control+best-of-7 perf anchor, frozen-red-test-may-be-wrong (CONVENTIONS); ADD: confirm instrument validity before a perf Must, full-dual-runtime gate for deletions, at-BUILD unsafe/unwrap audit (CONVENTIONS) |
| 2026-06-15 | CLOSE v2-performance (3/3 PASS) + fold deltas → foundation-version 3 (3 deltas from ft-search-off-eventloop) | all three v1-deferred bottlenecks delivered (xshard read latency, FT.SEARCH stalls, WAL group commit); a default-runtime perf no-op slipped past green tests until the effectiveness bench | TDD: mechanism-proxy pass ≠ effect measured (CONVENTIONS); ADD: per-runtime EFFECTIVENESS validation for `#[cfg]`-split primitives + make the instrument work before deferring a defect-hiding measurement (CONVENTIONS). v2 absolute magnitudes (xshard µs, WAL throughput) GATE-DEFERRED to GCloud per the milestone's sanctioned VM bench-exception |
10 changes: 5 additions & 5 deletions .add/milestones/v2-performance/MILESTONE.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,11 +27,11 @@ Out: RCU/ArcSwap snapshot reads & client-side MOVED/CLUSTER-SLOTS routing (decli

## Tasks (breadth-first decomposition; detail lives in each TASK.md)
- [x] xshard-read-fastpath depends-on: none — re-baseline cross-shard read latency per-runtime; recover it the lock-free-safe way via adaptive spin-then-park on the SPSC reply + read coalescing for the P1 multi-client case; delete the dead `--cross-shard-fast-path` flag/metric/docstrings. Target: at least HALVE the c1 GET regression with zero memory growth. **DONE 2026-06-14 (gate PASS):** C2 idle-gated + batch-depth-gated reply spin recovers c1-GET **+22.9% same-run** (−19.4% vs 3e376a1 ≤ ~20% contract line); verify caught + fixed a −27.5% pipelined regression (batch gate, commit 7048e8a). C3 cross-connection coalescing DEFERRED → follow-up task `xshard-read-coalescing` (human-approved v2 change request).
- [ ] ft-search-off-eventloop depends-on: none — keep a pathological FT.SEARCH (large K / deep HNSW) from stalling concurrent simple commands on the same shard; cooperative-yield or snapshot-handoff execution that respects the !Send slice.
- [ ] wal-group-commit depends-on: none — batch concurrent pending writes into one fsync under appendfsync=always; close a meaningful fraction of the ~11× throughput penalty with zero data-loss regression.
- [x] ft-search-off-eventloop depends-on: none — keep a pathological FT.SEARCH (large K / deep HNSW) from stalling concurrent simple commands on the same shard; cooperative-yield or snapshot-handoff execution that respects the !Send slice. **DONE 2026-06-15 (gate PASS, commit 7c4f8cd):** owned-snapshot capture + cooperatively-yielding `search_mvcc_yielding` seam; runtime-split `cooperative_yield` (tokio `yield_now`; monoio `sleep(ZERO)` timer-park — the verify bench caught the naive self-wake as a SILENT NO-OP on monoio's io_uring loop, fixed in-task). Co-located p99 **6.6ms/27ms vs sync 48/300ms** (~7–11× relief) both runtimes; result byte-identical (append-only mutable proof); zero new lock/RSS. Cost: −22% heavy brute-force QPS at default chunk=16384 (env-tunable, transient pre-compaction window only).
- [x] wal-group-commit depends-on: none — batch concurrent pending writes into one fsync under appendfsync=always; close a meaningful fraction of the ~11× throughput penalty with zero data-loss regression. **DONE 2026-06-14 (gate PASS):** group commit wired into all 4 AOF writer loops; crash-matrix + exactly-once preserved (no data-loss regression), consistency 197/197, dual-runtime green. ⚠ ABSOLUTE RPS win GATE-DEFERRED: OrbStack virtio fsync is near-free (appendfsync=always ≈ 0.9M RPS) so a batch never forms on the VM → the group-commit throughput gain is UNMEASURABLE on the only available instrument; mechanism proven by a deterministic batching seam test, absolute win deferred to real-disk/GCloud (same instrument-validity pattern as xshard).

## Exit criteria (observable; map each to the task that delivers it)
- [x] Cross-shard c1 GET regression at least HALVED vs v0.3.0 (tag 3e376a1, pre-shared-nothing), measured per-runtime as a best-of-N RPS ratio on the SAME quiesced, core-pinned moon-dev instrument before/after (relative anchor; absolute µs untrusted). M0 anchor ESTABLISHED 2026-06-14 (monoio, clean VM): c1-GET 37580→22252 = −40.8%; target = recover ≥ half the 15.3k-RPS gap → c1-GET ≥ ~30k. `grep` confirms the `--cross-shard-fast-path` flag + `moon_cross_shard_lock_contention_total` metric are GONE; consistency 197/197 @1/4/12 unchanged; RSS not regressed; s4-c100-GET guard within noise of ~202k (← xshard-read-fastpath) **MET on the relative form 2026-06-14:** same-run +18–23% c1-GET recovery → regression −19.4% ≤ ~20%; flag+metric grep-clean; 197/197 unchanged; RSS flat; c100 +10%. ⚠ The absolute "≥30k" SUB-line was RETIRED, not met (~25k): the whole-VM baseline itself drifted 37580→31437→20243 across three clean runs, so absolute RPS is not a stable instrument here — the contract's sanctioned metric is the relative ratio (§7 OBSERVE delta records the retirement). The structural residual (the second, origin→owner cross-thread wake) needs the §1-rejected RCU path; this is the single-connection floor, documented, not a waiver. (← xshard-read-fastpath)
- [ ] During a heavy FT.SEARCH on a shard, p99 of a simple command (PING/GET) on that same shard stays under a recorded bound; FT.SEARCH recall/correctness unchanged vs current (← ft-search-off-eventloop)
- [ ] appendfsync=always write throughput improves measurably at pipeline depth >1 with N concurrent writers (recorded before/after); crash-matrix green + exactly-once preserved (no data-loss regression) (← wal-group-commit)
- [ ] Cross-cutting per task: dual-runtime green, `clippy -D warnings` ×2 featuresets + `fmt` clean, zero new `unsafe`, zero new cross-thread lock (← all three)
- [x] During a heavy FT.SEARCH on a shard, p99 of a simple command (PING/GET) on that same shard stays under a recorded bound; FT.SEARCH recall/correctness unchanged vs current (← ft-search-off-eventloop) **MET 2026-06-15:** co-located PING p99 6.6ms (1-thread) / 27ms (3-thread) vs sync 48 / ~300ms — bounded, ~7–11× below the M0 stall, BOTH runtimes (tokio 6ms); FT.SEARCH result byte-identical to sync (G-IDENTITY, m2/m3 + MVCC suites green). Anchored on the relative before/after ratio + the deterministic `ft_search_cooperative_yields_total` proxy (absolute VM p99 jittery per the §1 instrument flag, as predicted).
- [x] appendfsync=always write throughput improves measurably at pipeline depth >1 with N concurrent writers (recorded before/after); crash-matrix green + exactly-once preserved (no data-loss regression) (← wal-group-commit) **MET on durability + mechanism 2026-06-14; throughput-magnitude GATE-DEFERRED:** crash-matrix green + exactly-once preserved (zero data-loss regression) + the batching seam deterministically forms one fsync per concurrent-writer batch. The measurable-throughput-gain SUB-clause is un-instrumentable on the OrbStack VM (virtio fsync near-free → no batch pressure → group-commit ≈ no-op there); deferred to real-disk/GCloud, same sanctioned pattern as the xshard absolute-µs retirement above.
- [x] Cross-cutting per task: dual-runtime green, `clippy -D warnings` ×2 featuresets + `fmt` clean, zero new `unsafe`, zero new cross-thread lock (← all three) **MET:** all three tasks closed dual-runtime green with clippy ×2 + fmt clean, audit-unsafe 100% (zero new unsafe), zero new cross-thread lock, RSS not grown — the milestone's hard constraints held end-to-end.
Loading
Loading