Skip to content

Phase 4a: execution context — documentation and phase record - #59

Merged
Wahbeh-Mohammad merged 13 commits into
mainfrom
14-phase-4a-execution-context-docs
Sep 16, 2026
Merged

Wahbeh-Mohammad merged 13 commits into
mainfrom
14-phase-4a-execution-context-docs

Conversation

@Wahbeh-Mohammad

Copy link
Copy Markdown
Contributor

Closes #14. Third of three phase 4a PRs — the documentation and the phase record for #57 and #58. Targets #58's branch. The same rebase-and-reprove note as #57 applies: CLAUDE.md, the READMEs, architecture.md and the roadmap's status notes here were written against main before 4b landed and are re-derived on the rebased tip.

What lands

15 files, +934 / −46, all Markdown.

  • The checklist, docs/work/mvp/phase4/phase4a/2026-09-08-phase4a-execution-context-checklist.md, written from the build: 20 rows, CTX-1–CTX-20, every one ✅ with its task and the test that proves it (the design's disposition table defers nothing). Plus what was built, the guards run red (eighteen mutations with their messages, and the rows the two fix rounds added), the audit groups run, seventeen deviations from the plan's text (none lowers a gate), the findings routed, and the postponed work re-read (the store cap's configuration source → phase 5a Task 13; the no-op span and tracer protocols → phase 5c Tasks 3–5; nothing new postponed).
  • The phase 4a design's ledger gains an "As built" addendum: P4-3 (the two private constants' sig/ mirrors), P4-8 (#tracer's five-parameter arity re-read from the gem), P4-10 (the eighth cop, not the seventh), P4-11 (the private #validate_context!), and one new row, P4-60 — _ContextHost, the self-type interface strict Steep needs for Context#close. It was numbered P4-40 while 4a and 4b ran as parallel lanes; 4b merged first with P4-40–P4-49, P4-50–P4-59 are 4c's, and the addendum, the checklist and the roadmap note say so, so no two rows share a number on main. docs/sdk-design-ruby/ §5.4, §8.1 and §10 are frozen; their addenda and the consolidation are a human's, stated in the roadmap note as 3a, 3b and 4b did.
  • docs/sdk-documentation/execution-context.md — new as-built page for the correlation model (65 annotations, every fence run verbatim on 4.0.6 and 3.2.11 by the reviewer); architecture.md, quality-gates.md (the eighth cop), the core README, README.md and docs/README.md updated to point at it.
  • The roadmap's execution step 5 and the phase 3, 4, 5, 6 and 8 segmentation designs — the "a phase returns to main as one phase-level pull request" sentence corrected in place, dated 2026-09-16, old wording kept visible, the way the mvp → main correction of 2026-09-14 was made: each sub-phase returns as its own code → tests → docs stack, merged bottom-up. That is what phases 1, 2, 3a, 3b and 4b actually did and what 4a, 4b and 4c are doing in parallel. Phase 8's design leans on "the phase-level PR" as a mechanism in three further sentences; those are recorded in the checklist for phase 8's execution to restate against its three stacks and are not edited here.
  • CLAUDE.md — the built-phases sentence, the lib-file count (63 → 75 under lib/dexpace/), the sig/ and test/ mirror clauses (hooks.rb, bounded_map.rb and context/call_key.rb the files with no test/ mirror), the cop count where the seventh was named, and the "Constraints that will bite" list. The roadmap gains the phase 4a status note and a thirty-eighth phase-10 inbound bullet (phase 0's plan amendment still says the keyword-splat cop carries no ordinal and 4a's is the seventh; the documentation half only).

Verification

  • bundle exec rake on Ruby 4.0.6 at this tip: exit 0, all seventeen gates green (1,272 / 6,889, 99.96%, YARD 0 undocumented).
  • ruby .claude/skills/housekeeping/probe.rb: no drift, all eight checks — including after the renumber commit; verify_knowledge_structure.rb OK; the housekeeping and knowledge suites green.
  • Empty diff under docs/product-spec/, docs/sdk-design-ruby/, docs/knowledge/ (harvested and notes — the design's two notes stand; notes/observability.md's 2026-09-13 entry quotes P4-8's two-positional recommendation where the build ships the gem's five, and is left for phase 10's plan Task 16, which already owns its next rewording), docs/deviations.md (phase 10 flips it), docs/first-release.md (the IO-38 row's CTX-7/CTX-8 sentence, added at planning, reads true of the build and is unchanged) and every earlier phase's documents beyond the dated step-5 sentences.
  • The final reviewer re-derived every count CLAUDE.md and the status note state against the tree, read 22 checklist rows against their tests (all proven), re-ran the CTX-13 and CTX-19 experiments and the two thread tests, and ran every execution-context.md fence by hand on both interpreters.

Known follow-up from the review (not blocking)

  • R2-1 — "Guards run red" rows 2, 3 and 8 carry the first build's failure counts (4, 6 and 3), one short of what the same mutation yields since round 1 added the real-promotion CTX-13 case (5, 7, and 3 failures plus 7 errors on 4.0.6). Refresh the three counts, or qualify them "on the first build" the way rows 19–24 date themselves.
  • When this stack is rebased onto main for the merge, the two lanes' shared prose — the built-phases paragraph, the file counts, the constraints list, the status notes — is reconciled with 4b's, and the surface manifest is regenerated on the rebased tip rather than hand-merged.

Phase 4a (issue #14). One module, Dexpace::Context, that the three
flat Data flavours include: CTX-1's one-way chain is DispatchContext
-> RequestContext -> ExchangeContext, with #promote_to_request and
exchange stage terminal by the absence of a third (P4-1). Each
promotion builds the successor through its .build, carries the bundle,
the call key and the store forward by identity, adds exactly one
artefact, and registers the successor with store.set -- the first
store entry a chain ever has, because construction registers nothing
(CTX-2, CTX-3, CTX-17). #close is store.release(self): a frozen Data
cannot carry a latch, and CTX-9's identity-conditional eviction makes
it idempotent through the store instead (P4-4). CTX-16's operation
name is nil or a non-empty frozen String, carried forward unchanged,
advisory only.

Dexpace::CallKey (private_constant) mints CTX-4's key, a frozen
"traceId:spanId:n" with one process-wide counter under one
::Thread::Mutex, so two default-constructed contexts are never ==
and a pinned call_key: is the one way to make them so (CTX-5, CTX-6).
Dexpace::ContextStore HAS the private_constant Dexpace::BoundedMap --
the one bounded-keyed-map implementation XCUT-14 and AUTH-19 will
share, a strong Hash behind one mutex with XCUT-14's drain loop inside
the same synchronize as the insert -- and exposes CTX-8's #set and
after the mutex is released), CTX-18's #[] and the identity-checked
assigned .default. The cap must be a positive Integer, since a
negative one would spin the drain forever.

Dexpace::Instrumentation is roadmap obligation 1 discharged: Bundle,
a frozen Data of eight members with #valid? derived (P4-6) and NONE
carrying OBS-26's sentinels; TraceIdFlavour, a closed set of NONE,
W3C and DATADOG with the trace-id sentinel on the flavour and the
span-id sentinel on Bundle (P4-7); and NO_SPAN, NO_TRACER and
NO_TRACER_FACTORY, three frozen singletons whose classes are private
so phase 5c can widen them without an NFR-4 diff, the factory's
(P4-8). Every sig/ mirror declares every method for the strict core
Steep target, the two private constants included as hooks.rbs is;
_ContextHost is the self-type Context needs (P4-40). The entry file
requires the twelve files in dependency order.
Dexpace/NoWeakReferences is CTX-19 as a lint rule (P4-10): a constant
reference to ObjectSpace::WeakMap or ::WeakKeyMap -- qualified,
cbase-qualified, or bare inside `module ObjectSpace` -- to WeakRef, or
a `require "weakref"` in any form tools/require_scan.rb reads, is an
offense naming CTX-19 and the cap CTX-11 makes the leak backstop. A
source-text cop, because RuboCop parses and never evaluates, so the
WeakKeyMap rows parse at TargetRubyVersion 3.2 where the constant does
not exist. It is the eighth custom cop, not the plan's seventh:
CLAUDE.md, quality-gates.md, phase 2's checklist and cops_test.rb all
count Dexpace/NoKeywordSplat, which phase 0's plan says is uncounted.
Its cases live in a nested class of cops_test.rb under the 100-line
cap, fourteen rejected and fifteen accepted on RuboCop 1.91.0; the
.rubocop.yml entry scopes it to every gem's lib/, since the require
allowlist covers core alone.

The runtime surface manifest of dexpace-core grows from 515 to 580
lines through `rake surface:regenerate`, once, and every one of the
65 added rows was read against the design's object model: the seven
flat constants and the instrumentation subsystem, their .build and
.of and .default singletons, the Data readers of the five value types
(which the shipped walker holds, closing the plan's finding against
phase 0), Bundle's #valid? and #remote?, TraceIdFlavour's two
predicates and four constants, ContextStore's five methods and its
cap, and the three no-op singletons typed by their private classes.
Nothing removed; no private_constant present. Context's construction
validation is a private instance method rather than the plan's public
Context.validate!, so the module's public surface is the one method
the design gives it. The core smoke suite's constant list gains the
seven public constants and asserts BoundedMap and CallKey are as
unreachable as Hooks. Both are files a gate reads at run time, so
they travel with the code.
Review round 0, R0-4 and R0-5. dispatch_context.rbs still named the
plan's public Context.validate!, which the build made the private
instance method #validate_context! (checklist deviation 5); the comment
now names what lib defines. The cop's head comment claimed to match
`require "weakref"` "in the forms tools/require_scan.rb reads", but the
scanner's loader list includes autoload and the cop's matcher does not:
`autoload :WeakRef, "weakref"` produces no offense on its own line. The
comment now states the three require spellings the matcher reads and
why the autoload form needs no row -- it names the constant as a Symbol,
and the reference that later triggers it is the one on_const flags,
verified against a scratch adapter file on 4.0.6.
Context#validate_context! accepted any non-nil object as a pinned
call_key and the two flavours that carry an operation_name checked only
for "": a Symbol passed silently, keying a slot no String lookup could
find, and an Integer escaped as a NoMethodError from #empty? where every
core model raises a field-named InvalidArgumentError. The design says a
given key "must be a non-empty String" and that operation_name is nil or
a non-empty frozen String, and sig/ already types both that way.

Add the "must be a String" guard to #validate_context!, between the
required! check and the empty check, and hoist CTX-16's two-state rule
into one private Context#validate_operation_name! the RequestContext and
ExchangeContext initializers share, so the message form lives in one
place the way Model.required! keeps SEAM-29's. The .build paths pass a
frozen Symbol or Integer through Model.frozen_string untouched, so the
guard in initialize covers .build, #with and both promotions alike.
No public signature changes; context.rbs declares the private helper.
The three phase-4 lanes each added their layer's resolution case to
dexpace_test.rb; alone each kept the class under Metrics/ClassLength's
hundred lines, together they reach 116. The six "requiring dexpace
alone makes the whole <layer> resolve" cases move into a nested
DexpaceTest::Layers, the split every larger suite in this gem already
uses, so the honest RuboCop run is clean again. No case changes.
Ten suites mirroring the ten public lib/ files one for one, each
opening with the IDs it exercises, 112 cases in all, identical counts
on 3.2.11, 3.3.12, 3.4.10 and 4.0.6, and one fake, FakeContext, a
real in-memory Dexpace::Context of two members required explicitly by
the two suites that drive the store through a keyed occupant.

The tests a reader would otherwise write wrong, written the way the
design says: CTX-9's trap over two DispatchContexts with one pinned
key -- == and not equal?, the only way the pair exists -- so a
release by value equality fails where two minted contexts would pass
it; CTX-2's carried-forward members asserted with assert_same, never
assert_equal, so a promotion that rebuilds the bundle fails; CTX-15's
two keys minted from the SAME Bundle::NONE object; CTX-16 asserted as
a negative over a named and an unnamed promotion sharing one key;
CTX-8's race with 32 threads released from one ::Thread::Queue onto
one key, every loser's error asserted on; CTX-12's discriminating
single-threaded form, exactly one eviction per insert from cap and
the size never above it, with the aggregate count it does not use
explained in the comment; CTX-19's 1000 contexts surviving three
GC.starts with every local dropped, plus a real Request and Response
kept reachable through the store; and the Fiber[] boundary with its
setup guard. CTX-6's case asserts the counter suffix increases
strictly across flavours in build order rather than that three keys
are distinct, because a counter per flavour passed the weaker form in
the guard battery. Every thread a case starts is joined, which
DexpaceTestCase's teardown asserts; every suite is one class per
behaviour group under the 100-line cap.

test/gates/rubocop_config_test.rb pins Dexpace/NoWeakReferences's
require: line, its enablement and its every-gem scope, so a narrowing
of Include: to core alone is the one mutation of .rubocop.yml the
gate suite catches.
…uard

Review round 0, R0-1, R0-2 and R0-3, all in context_store_test.rb.

R0-1: the CTX-13 no-refresh claim was proven only through ContextStore#set
on a FakeContext, so a #promote_to_exchange that released its source
before setting the successor -- refreshing the chain's eviction position
and leaving the slot transiently empty between two mutex acquisitions --
survived every suite. A BoundTest case now drives the policy through the
real path: cap 3, three chains at the request stage, a's promoted to
exchange, a fourth chain promoted, a still the victim and both of its
links closing false. The survivor assertion is a key list so the red
never dumps a whole context. Under the mutation: Expected ["b","c","d"]
Actual ["a","c","d"] on 4.0.6 and 3.2.11.

R0-2: a memoised .default (`@default ||= new`, load-time line deleted)
survived, because the identity case runs after some .build has already
called .default in-process. A ReachabilityTest case asks a fresh process
-- RbConfig.ruby -w -W:deprecated, `require "dexpace"` alone -- whether
the ivar is set before any call. Under the mutation: Expected "true"
Actual "false" on both interpreters.

R0-3: the Fiber[] boundary test's setup guard covered the main fiber and
a child Fiber; the design names three carriers. The guard now writes a
fiber-storage slot and a fiber-local slot on the main fiber and asserts
[:fiber_storage, nil] inside a child Fiber, a new ::Thread and an
Enumerator's internal fiber -- observability/016d9154's fact, re-run on
4.0.6 and 3.2.11 -- and resets both slots in an ensure.

Store suite 18 -> 20 runs; ten seeded runs on 4.0.6 and one on 3.2.11
green.
Four cases for review round 1's R1-1, each asserting the field-named
InvalidArgumentError message on :sym and 5 (or :GetUser and 7): the
probe includer in context_test.rb drives #validate_context! directly and
the surface case now names #validate_operation_name! as private too;
dispatch_context_test.rb covers DispatchContext.build(call_key:) and a
#with derivation; request_context_test.rb covers RequestContext.build
and the dispatch->request promotion, and asserts the failed promotion
registered nothing; exchange_context_test.rb covers the terminal
flavour's own build path. Four mutations run red on 4.0.6 and 3.2.11:
the call_key guard deleted (1 failure in each of the first two suites),
the operation_name guard deleted (1 in each of the last two), and the
helper's call removed from either flavour's initializer (2 failures in
that flavour's suite alone, the other suite staying green).
The checklist, written from the build: twenty rows, CTX-1 to CTX-20,
every one implemented and tested, each naming the task and the suite
that proves it; what was built; the guard battery -- eighteen
single-edit mutations, seventeen caught on the first run and the
per-flavour counter caught once CTX-6's case asserted the counter's
order across flavours, with the CTX-9 trap and the WeakMap swap run
red on 3.2.11 as well; the audit groups; sixteen departures from the
plan's text, four of them phases 0-3b as built overriding the plan
(the eighth cop, the CountKeywordArgs: false already in the baseline,
the sig/ mirrors hooks.rbs sets the precedent for, the Data readers
the shipped walker already holds); the findings routed; and the two
postponed items keeping their owners.

The design's ledger gains an "As built" addendum: P4-40 for the
_ContextHost self-type interface, numbered from the tree because 4b's
design filed P4-12-P4-25 and 4c's P4-26-P4-39, and P4-3, P4-8, P4-10
and P4-11 as built. docs/sdk-documentation/execution-context.md is
the as-built page, every fence run verbatim on 4.0.6 and 3.2.11 with
identical output; architecture.md, quality-gates.md (eight cops, 129
cases), the core README, README.md and docs/README.md point at it.

CLAUDE.md's built-phases paragraph, its lib-file count (seventy-five
beside version.rb, three private constants without a test/ mirror),
its phase-directory sentence (six checklists) and its constraints list
(the no-latch close, CTX-9's equal?, the eighth cop) are rewritten
from the tree. The roadmap gains the phase 4a status note and a
thirty-eighth inbound bullet for phase 10 (phase 0's plan calls the
keyword-splat cop uncounted; every as-built document counts it), and
its execution step 5 is corrected in place, with the phase 3, 4, 5, 6
and 8 segmentation designs' restatements of it: each sub-phase returns
to main as its own code -> tests -> docs stack, which is what every
phase since 1 did and what "one phase-level pull request" said none
would. The frozen chapters' addenda and the consolidation into design
§10 are a human's, as they were for 3a and 3b.
…unts

Review round 0's three should_fix findings landed as test cases; this is
their record. The checklist's CTX-13 row now says the no-refresh claim is
driven through a real promotion as well as the fake, its guards table
gains rows 19 (release-then-set promotion) and 20 (memoised .default),
both red on 4.0.6 and 3.2.11, and its prose says which round found each
survivor; the Fiber scheduler audit row and deviation 13 describe the
boundary test's three-carrier guard and the new run counts (store suite
20, execution-context suites 114). The design's As-built addendum gains
one bullet for the three testing-strategy cases as built in round 1, and
the roadmap's 4a status note reads 1,268 runs and twenty mutations.
Review round 1's R1-1: a pinned call_key or an operation_name was never
checked to be a String. The checklist's CTX-4 and CTX-16 rows now say
the type is enforced and where, the "What was built" paragraph names the
second private helper, guards 21-24 record the four mutations run red on
4.0.6 and 3.2.11, deviation 17 states the departure from the plan's
emptiness-only fences with its reason, and deviation 13 carries the new
per-suite and total run counts (118 execution-context cases). The
design's As-built addendum gains the round-2 bullet and the P4-11 bullet
names the helper; the roadmap's status note reads 1,272 runs,
twenty-four mutations and seventeen departures; the as-built page's
operation-name and pinned-key passages state the refusal and its
message.
Phase 4a and 4b ran as parallel stacks off main, and both numbered
their first as-built ledger row P4-40 -- 4a's one row (_ContextHost)
and 4b's ten (P4-40-P4-49). 4b merged first (#53-#55), so its numbers
stand on main; 4a's row moves to P4-60, with P4-50-P4-59 reserved for
phase 4c, which builds on 4b's tip. The design addendum, the checklist
and the roadmap's status note say what the row was called while the
lanes overlapped and why it moved, so the history stays readable.
Phase 4a's stack was cut from main at 419aace and built in parallel
with 4b; 4b and 4c merged first, so the three branches were rebased
onto their main before merging. The rebase kept both sides of every
shared file -- the entry file's three phase-4 blocks in sub-phase
order, the smoke suite's three layer pins, the merged built-phases
sentences and counts in CLAUDE.md, the READMEs and architecture.md --
and regenerated the surface manifest on the merged tree (665 + 65 =
730 rows, none of 4b's or 4c's changed). The status note keeps the
build's own numbers and says what they became.
@Wahbeh-Mohammad
Wahbeh-Mohammad force-pushed the 14-phase-4a-execution-context-docs branch from dfde33f to 733de8a Compare September 16, 2026 20:04
@Wahbeh-Mohammad
Wahbeh-Mohammad force-pushed the 14-phase-4a-execution-context-tests branch from 409d2c7 to 5752d86 Compare September 16, 2026 20:04
@Wahbeh-Mohammad
Wahbeh-Mohammad changed the base branch from 14-phase-4a-execution-context-tests to main September 16, 2026 20:10
@Wahbeh-Mohammad
Wahbeh-Mohammad merged commit 993c439 into main Sep 16, 2026
10 checks passed
@Wahbeh-Mohammad
Wahbeh-Mohammad deleted the 14-phase-4a-execution-context-docs branch September 16, 2026 20:11
@Wahbeh-Mohammad

Copy link
Copy Markdown
Contributor Author

Review record for the phase 4a stack (#57 → #58 → #59)

Three independent reviews, each by a fresh agent with no memory of the previous one, each re-running the gates itself on every tip and both interpreters; two fix rounds between them. Every finding got a stable id, a severity, a file:line and the command output that was its evidence. approve required zero blocking and zero should-fix; the run was capped at four reviews and stopped at the third because it approved. The stack was then rebased onto main (4b and 4c had merged first) and re-proven by the manager — that pass, with its one extra fix commit, is recorded on #57.

Round Verdict Blocking Should-fix Nits Disposition
0 changes required 0 3 2 all 5 fixed
1 changes required 0 1 0 fixed
2 (final) approve 0 0 1 carried into the PR bodies

No finding was ever blocking; nothing was skipped or disputed. Three of the four should-fixes were tests that did not reach far enough — a claim resting on the fake rather than a real chain, a load-time property asserted only by identity, a guard covering one carrier of three — and the reviewers found each by a mutation that survived or by reading the test against the design's own sentence. The fourth was a real argument-boundary gap in shipped code.

Round 0 → fixed in round 1

  • R0-1 should-fix — CTX-13's "a promotion re-setting a key does not refresh its eviction position" was proven only through ContextStore#set on the fake; a #promote_to_exchange that called store.release(self) before store.set(ctx) survived all four context suites. Fixed on Phase 4a: execution context — tests and the fake #58: a real-chain case (cap 3, chains a/b/c promoted to request, a promoted to exchange before d arrives, survivors asserted as a key list) that goes red under the release-then-set mutation on 4.0.6 and 3.2.11 — Expected ["b", "c", "d"] Actual ["a", "c", "d"].
  • R0-2 should-fix — a memoised ContextStore.default (@default ||= new, the load-time assignment deleted) survived every suite, although the design says eager assignment is what keeps CTX-17's "construction registers nothing" inert and the source comment names XCUT-11 as the reason. Fixed on Phase 4a: execution context — tests and the fake #58: a fresh-process case — Open3.capture3(RbConfig.ruby, …, 'require "dexpace"; print Dexpace::ContextStore.instance_variable_defined?(:@default)') — red under the memoised form on both interpreters.
  • R0-3 should-fix — the Fiber[] boundary test's setup guard asserted the Fiber[:probe] / Thread.current[:probe] pair on the main fiber and in a child Fiber only, where the design's testing strategy names three carriers. Fixed on Phase 4a: execution context — tests and the fake #58: the guard writes both slots on the main fiber and asserts [:fiber_storage, nil] inside a child Fiber, a new ::Thread and an Enumerator's internal fiber; 8 assertions alone, no warning under -w -W:deprecated.
  • R0-4 nit — dispatch_context.rbs's comment still named the plan's public Context.validate!, renamed in the build to the private #validate_context!. Reworded on Phase 4a: execution context — the context layer #57. R0-5 nit — the cop's head comment claimed the require forms tools/require_scan.rb reads, but its matcher does not see autoload :WeakRef, "weakref" (verified over a scratch adapter-gem file: no offense). Closed by narrowing the comment to the three require spellings it matches, per the design's R1 list; any later use must name WeakRef, which on_const flags.

Round 1 → fixed in round 2

  • R1-1 should-fix — a pinned call_key or an operation_name was never checked to be a String: by hand against the docs tip, DispatchContext.build(bundle: NONE, call_key: :sym, store: st) succeeded with a Symbol key (st[:sym] resolving, st["sym"] nil), promote_to_request(operation_name: :GetUser) succeeded, and .build(call_key: 5) / operation_name: 7 leaked NoMethodError (Integer#empty?) rather than InvalidArgumentError — Model.frozen_string passes a frozen Symbol or Integer through untouched. Fixed on Phase 4a: execution context — the context layer #57: Context#validate_context! raises the field-named call_key must be a String between the required and empty checks, and CTX-16's two-state rule is hoisted from the two flavours' initializers into one private Context#validate_operation_name! (nil / non-String → operation_name must be a String; "" → must not be empty), declared in context.rbs; because the guard lives in initialize, it covers .build, #with and both promotions. Four cases on Phase 4a: execution context — tests and the fake #58 across context_test.rb, dispatch_context_test.rb, request_context_test.rb and exchange_context_test.rb, each asserting the exact message on :sym/5 and :GetUser/7; four mutations (each guard deleted; the helper's call removed from either flavour alone) red on 4.0.6 and 3.2.11 on exactly those cases.

Round 2 (final) — open, carried into the PR bodies

  • R2-1 nit — "Guards run red" rows 2, 3 and 8 in the checklist carry the first build's failure counts (4, 6 and 3), one short of what the same mutation yields since round 1 added the real-promotion CTX-13 case (5, 7, and 3 failures plus 7 errors on 4.0.6). Listed on Phase 4a: execution context — documentation and phase record #59.

What the final reviewer verified, in its own runs

  1. Gates — Phase 4a: execution context — the context layer #57 tip d8cf132: all seventeen individually on 4.0.6, all green including the SimpleCov floor (1,154 runs / 6,094 assertions, 97.61%; the tolerated red the layering rule permits was not needed), the honest RuboCop run 240 files clean, cops:test 129 cases, steep clean over the strict core target, YARD 402 methods / 0 undocumented; the matrix set and cops:test on 3.2.11 (97.58%). Phase 4a: execution context — tests and the fake #58 tip 409d2c7: full bundle exec rake green on 4.0.6 in 245 s (1,272 runs / 6,889 assertions, 3,231 / 3,232 lines, 99.96%; test:gates 130 / 538), RuboCop 251 files clean, the matrix set on 3.2.11, 3.3.12 and 3.4.10, the core suite on seeds 1, 99991, 424242 on 4.0.6 and 8675309 on 3.2.11 at 1,255 runs every time, each of the ten suites alone under -w -W:deprecated on both interpreters (4/9/20/14/3/5/20/21/14/8 = 118 runs, 0 warnings). Phase 4a: execution context — documentation and phase record #59 tip 90199be: full rake green, probe exit 0 on all eight checks and on --only registers,citations, knowledge verifier OK.
  2. Every prior finding individually — R1-1 by hand on both interpreters across .build, #with and both promotions (every path raising the field-named message; a failed promotion registering nothing) and by re-applying the four round-2 mutations; round 1 had done the same for all five round-0 findings (the two new store cases red under the re-applied mutations, the three-carrier guard read and run alone, the sig comment by grep, the cop comment by scratch files).
  3. The mutation battery — 24 single-edit mutants applied by the reviewer itself, 23 caught on 4.0.6 and 3.2.11; the one survivor is the plan's documented while→if drain equivalent, and it survived exactly as the plan's "open questions, resolved" item 6 says. Across the run: 18 by the implementer, then 18, 20 and 24 by the three reviewers. The ones that matter: #release comparing == (the CTX-9 trap); the drain evicting two per insert, and deleted; #put as an overwrite; a promotion rebuilding the bundle (only assert_same catches it); the key derived from the bundle alone; one counter per flavour (caught by the strictly-increasing case the implementer strengthened after the weaker form survived its own battery); the store's Hash swapped for ObjectSpace::WeakMap (3 failures, 7 errors, the floor-specific NoMethodError on 3.2.11); Bundle#valid? ignoring the sentinels; an uppercase span id; a W3C pattern without timeout:; #release raising KeyError on an absent slot; #tracer allocating per call; the release-then-set promotion and the memoised .default (round 0's two survivors); the operation name folded into the store key; .build registering what it builds; the cop's Include narrowed to core (caught by test/gates/rubocop_config_test.rb, since CopCase structurally cannot see .rubocop.yml).
  4. Concurrency — the 32-thread CTX-8 race and the 16×1,000 CTX-7 burst run individually ten times on 4.0.6 and five on 3.2.11 under timeout 120: 1 winner / 31 ContextConflictErrors each naming the key, and 16,000 entries, every run; no flake, no hang, no warning. Every synchronize block read: BoundedMap's insert-and-drain, CallKey's increment, the store's lookup and release — each a hash write or a counter, never across a callback.
  5. The layer by experiment, identical on 4.0.6 and 3.2.11: CTX-19 — 1,000 contexts registered inside a method, locals dropped, three GC.starts, size 1,000 and a sampled key resolving; the same script with bounded_map.rb's Hash swapped for ObjectSpace::WeakMap — size 0, key nil, and the shipped test red on the same edit; ObjectSpace::WeakKeyMap undefined on 3.2.11 while cops:test still parses and rejects the spelling there, and the cop over a scratch adapter-gem file flags WeakKeyMap.new, WeakRef.new and ::ObjectSpace::WeakMap.new; #with re-validating on all five Data types (<name> is required on both interpreters — Data#with skips initialize on 3.2 and Model#with is what re-validates); CTX-5/6/15/16/17 by hand (two default contexts not ==; the pinned pair ==, eql?, hash-equal; three flavours from one counter with consecutive suffixes; two keys from the same equal? Bundle::NONE differing; store.size 0 after every .build, the dispatch→request promotion the registering one); CTX-13 with cap 3 through real promotions; Bundle::NONE's pair exactly OBS-26's 32 hex zeros and 16 zeros, DATADOG.invalid_trace_id "0" and its rendering 2**64 − 1, the renders?/valid_trace_id? property over 256 seeded draws, Bundle.build(**bundle.to_h) == bundle, "B" * 16 refused with span_id must be 16 lowercase hex chars (OBS-26), all five patterns carrying timeout: 1.0 with Regexp.timeout nil; NO_TRACER_FACTORY#tracer called five ways returning the same NO_TRACER with #parameters matching opentelemetry-api 1.11.0's five; the private constants absent from constants and refused qualified (and const_get reaching them, as Ruby does for phase 2's Hooks — not a finding).
  6. Layering — main ⊂ d8cf132 ⊂ 409d2c7 ⊂ 90199be; ten commits, the right prefixes, no attribution lines; every file in each diff classified (30 code, 12 tests, 15 docs); sig/ mirrors lib/ 77/77 with the two private constants carrying hooks.rbs's comment; the manifest's +65 rows exactly the reviewer's own derivation from the object model (2+2+8+6+8+1+14+3+13+8), 0 removed; no 4b file touched (error.rb, hooks.rb, closeable.rb unchanged from main); empty diffs under the frozen trees, docs/deviations.md, docs/first-release.md and tools/; the only earlier-phase edits the dated step-5 sentences in the roadmap and the five segmentation designs, plus the roadmap's own status note and 38th inbound bullet.
  7. Global constraints by script over the 24 new .rb — both header lines; only require_relative under core lib/; no weakref/WeakMap/WeakKeyMap/WeakRef or ObjectSpace in any gem lib/ outside comments (the seventh and eighth cops each 0 over all 87 gem lib files); every Regexp.new with timeout: and no Regexp.timeout=; no Enumerator.new; no Fiber[]/Thread.current[] write in lib/; no Closeable on a context; Dexpace::ArgumentError never defined.
  8. Plan and spec coverage — the 9 tasks' files; each Scope bullet of Phase 4a: Execution Context #14 mapped; all 20 rows ✅ with the cited test read for 17 of them (all proven; CTX-13 had been partially proven in round 0 and is proven since round 1); the corpus consulted for the sample (--req CTX-9,CTX-13,CTX-16,CTX-19: 20 of 20 substantive, 0 roll-ups, 0 gaps); the ledger addendum's numbering re-derived from the tree; the findings routed as the design asks — the Data-readers finding recorded as closed by phase 0 as built and not routed, the ParameterLists count already applied by phase 0, the drain measurement already on the IO-38 row.
  9. Phase record — execution-context.md's fences executed statement by statement, 65 annotations, 0 mismatches on both interpreters; CLAUDE.md counts re-derived from git ls-tree (75 lib/ files beside version.rb, 77 sig mirrors, test mirrors for all but hooks.rb, bounded_map.rb and context/call_key.rb, 8 cops, 6 checklists, 11 phase directories, 40 harvested topics, the 580-row manifest); the status note's numbers against the runs; the step-5 correction present and dated in all six documents with the old wording kept visible, saying nothing false about 4b.
  10. Report discrepancies — one, explained: the fix report and round 1 claimed the tests tip's 3.2.11 matrix run at 3,189 / 3,189 (100.00%); the final reviewer's run of the same command at 409d2c7 gave 3,188 / 3,189, and 3.3.12 the same. The uncovered line is phase 2's registry.rb:253, a thread-scheduling-dependent branch main's own suite covers or not run to run — not 4a code and not a 4a test. Every other number in the fix report reproduced exactly.

Run: workflow wf_a6129983-b32, 6 agents (1 implementer, 3 reviewers, 2 fixers), 2.3M tokens, 3.8 h.

@Wahbeh-Mohammad Wahbeh-Mohammad mentioned this pull request Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Phase 4a: Execution Context

1 participant