Repository navigation
feat: multi-network support with runtime network registration - #19
Merged
Merged
Conversation
Brings the role-centric `roles` config in as the base for network registration: registered networks describe their DNAs with the same per-role shape the static config uses.
Networks are keyed by `happ_id` -- the same identity space as a conductor installed-app id -- so a registered network needs no separate slug. Each record carries the per-role DNA config in the same shape as the static `roles` config, optional hApp metadata for the info endpoint, and an optional `allowed_agents` list. The three backends mirror the session and allowed-agent stores: memory for tests and ephemeral deployments, sqlite for single-node, KV for Workers.
#13) A joining service could serve exactly one network -- the one in its static config. Networks registered at runtime through POST /v1/admin/networks are now first-class: join, provision, and the info endpoint all take a `network` naming a registered happ_id, and resolve roles, membrane proofs, and allowed_agents from that network's record instead of the static config. The static network keeps working untouched. Naming it explicitly by its own happ_id collapses to the same scope as omitting `network` (normalizeNetwork), so one agent is never split across two spellings of one network. Registering a network under the service's own happ id is rejected, and startup warns if one is already stored that way. Role `dna_hash` is required only when membrane proofs are enabled, matching the rule the static roles config already follows.
Network identity (happ_id) stands in for DNA hash but nothing enforced it: registering a network with a dna_hash already used elsewhere let one agent join both and receive membrane proofs for the same cell twice, a chain-fork risk. POST /v1/admin/networks now rejects a candidate dna_hash that collides with any other registered network or the service's static roles with 409 duplicate_dna_hash, naming the hash and the owning network. Duplicates within one registration's own roles remain allowed (same-DNA multi-role apps exist), and re-registering a happ_id against its own prior hashes is unaffected.
JoiningClient.join() and the provision CLI can now name a network. The client defaults `network` to the happ_id it discovers from GET /v1/info, so a client pointed at a service backed by a registered network joins that network without the caller spelling it out; passing an explicit value overrides the discovered id, and an explicit null omits the key entirely.
…ng (#13) Documents the admin network API, the `network` parameter on join and provision, per-network `GET /v1/info/:happ_id`, and dynamic-only mode (network_registration with no static roles). A network's identity is its happ_id -- the same identity space as a conductor installed-app id -- so `.well-known`'s existing happ_id is what routes a client to its network. Covers the conditional `dna_hash` rule (required only under membrane proofs) and the cross-network uniqueness rule with its 409 duplicate_dna_hash error, including why the check is best-effort on KV. CLI.md gets the same treatment rather than just a `--network` flag row: its prose still said role names come from the service config and that the service "must be configured with roles", both of which stop being true once a session names a network -- provision reads roles, modifiers, and happ_bundle_url from that network's registration record instead, and a dynamic-only service has no static roles to read.
…ry (#13) Session uniqueness was keyed on agent key alone, so an agent that joined one network got 409 agent_already_joined when joining another -- a single-network assumption left over in the session layer. Lookups are now scoped to (agent_key, network): findByAgentKey matches the exact network scope (undefined network is its own no-network scope, byte-identical to prior behavior), and a new findAnyByAgentKey covers call sites like reconnect where linker and gateway URLs are service-wide rather than per-network. Reconnect builds on that scoping to close a recovery gap: an agent that completed join and crashed before calling provision had no way back, since a fresh join only returns 409. POST /v1/reconnect now also returns the session token for the requested network's ready session, and the client gains reconnectAndProvision() for the one-call version of that path. The URL refresh and the token lookup are gated separately, so an agent whose only session names a non-static network still gets its URLs. The sqlite `sessions` table gains a `network` column and a composite (agent_key, network) index. No migration ships for it: there is no deployed data at this stage, and sessions are ephemeral anyway, so an old sessions.db should be deleted rather than upgraded.
Session uniqueness is per (agent, network): a second join on a different network succeeds instead of 409ing, and agent_already_joined now means re-joining a network the agent already has a live session on. Documents reconnect's `network` parameter, the `session` token it returns, and the crash-recovery path that token exists for. The URL refresh and the session-token lookup are gated separately, so the three outcomes are spelled out rather than left to be inferred from the error table.
|
Yes, this is too big to review in a sensible amount of time. Accepting it as it is and I'll test it when it's made available downstream. |
# Conflicts: # JOINING_SERVICE_API.md # src/cli/provision.ts
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Context
I admit that this is a too big AI generated PR. But I've been around the block with it a number of times and looked at it pretty carefully.
Note that it brings in one part of what the unyt fork was about, improving handling how to reconnect during a partial join. The second of that fork is addressed in the followup branch
feat/worker-kv-parity.For review I think the most important part is making sure that the API matches what's needed by the automation pipeline.
Description
Second half of #13 (with the allowed-agent registration PR): one joining service serves many networks, each identified by its happ_id. A pipeline registers a network — per-role DNA config, optional hApp metadata, and its progenitor — in one call:
.well-known/holo-joiningalready carrieshapp_id;JoiningClient.discover()now passes it through tojoin(), so an app on a registered network just hosts the standard document.GET /v1/info/<happ_id>serves per-network info; bare/v1/info(and/v1/info/<static happ id>) means the statically configured network.config.rolesneeded — register networks at runtime only.networkbody field (a happ_id; the static id is equivalent to omitting it). Sessions are scoped per(agent, network): one key can join any number of networks; same-network re-joins 409. Non-emptyallowed_agentsgates a network (403join_rejectedfor unlisted agents) — and gated networks do NOT expose their role modifiers via info; joiners get them at provision.happ_bundle_url.dna_hashper role is required only when membrane proofs are enabled.POST /v1/reconnectaccepts an optionalnetworkand returns the agent'ssessiontoken for that scope, so an agent that crashes between joining and installing can recover instead of being permanently stuck behind409 agent_already_joined. Recovery is keyed on the one thing the client durably holds first — its agent private key — via the existing signed-timestamp check;/v1/joinis deliberately NOT made idempotent (that would turn any public agent key into a provisioning credential).JoiningClientgainsreconnectAndProvision().NetworkStore(memory / sqlitenetworks.db/ Cloudflare KV) andnetwork_registration.admin_secretconfig; admin middleware is path-scoped (three-way coexistence with linker and agent admin surfaces regression-tested).dna_hashuniqueness is enforced at registration (409duplicate_dna_hash, checked against other registered networks and the static roles) — duplicate hashes would let one agent be provisioned twice for the same cell, forking the chain. Best-effort on KV (eventual consistency), noted in docs.Breaking / operational notes:
discover(),join()defaultsnetworkto the document'shapp_id. Verify an app domain's.well-knownhapp_idmatches the service's static happ id (or a registered network) before deploying; passnetwork: nullto suppress the default.hc_auth.requiredwant the hc-auth client misreads /request-auth status codes #15 status-mapping fix first.network_config, the bundle URL, and a mintable membrane proof, where previously it yielded only public infra URLs. Deliberately unchanged here (changing the signed payload breaks existing signers); worth a follow-up to coveragent_keyand a nonce.network_config(bootstrap/relay/auth-server URLs) remains service-wide.sessionstable gains anetworkcolumn andnetworks.dbis new, but nothing upgrades an existing file — there is no deployed data to preserve at this stage. Anyone who ran an earlier build locally should deletesessions.dbandnetworks.dbbefore running this one. Sessions are ephemeral (pending ones expire on TTL), so there is nothing worth carrying across.Stacks on the role-centric config PR (#18) and the allowed-agent registration PR (#17) — merge those first (this branch is based on both).
Followed by
feat/worker-kv-parity, which brings the Cloudflare Worker entry point up to parity with Node. That work was split out of this branch: it touches onlydeploy/and the auth-plugin construction shared by both entry points, and reviews independently.Commits (7)
feat: NetworkStore with memory, sqlite, and KV backendsfeat: multi-network registration with network-aware join and provisiondna_hashfeat: enforce dna_hash uniqueness across registered networksduplicate_dna_hashfeat: pass network through JoiningClient and provision CLIdocs: network registration, happ_id identity, and network-aware joiningfeat: scope join sessions by agent and network, with reconnect recoverydocs: per-network session semantics and reconnect recoveryDiff vs base: 37 files, +3479/-103.