Skip to content

feat: multi-network support with runtime network registration - #19

Merged
zippy merged 9 commits into
mainfrom
feat/multi-network
Aug 21, 2026
Merged

zippy merged 9 commits into
mainfrom
feat/multi-network

Conversation

@zippy

@zippy zippy commented Aug 17, 2026

Copy link
Copy Markdown
Member

Context

I admit that this is a too big AI generated PR. But I've been around the block with it a number of times and looked at it pretty carefully.

Note that it brings in one part of what the unyt fork was about, improving handling how to reconnect during a partial join. The second of that fork is addressed in the followup branch feat/worker-kv-parity.

For review I think the most important part is making sure that the API matches what's needed by the automation pipeline.

Description


Second half of #13 (with the allowed-agent registration PR): one joining service serves many networks, each identified by its happ_id. A pipeline registers a network — per-role DNA config, optional hApp metadata, and its progenitor — in one call:

curl -X POST https://join.example.com/v1/admin/networks \
  -H "Authorization: Bearer $JOINING_ADMIN_SECRET" \
  -d '{
    "happ_id": "acme-net",
    "happ": { "name": "Acme", "happ_bundle_url": "https://..." },
    "roles": { "main": { "dna_hash": "uhC0k...", "modifiers": { "network_seed": "..." } } },
    "allowed_agents": ["uhCAk...progenitor..."]
  }'
  • Discovery routes by happ_id with no format change: .well-known/holo-joining already carries happ_id; JoiningClient.discover() now passes it through to join(), so an app on a registered network just hosts the standard document. GET /v1/info/<happ_id> serves per-network info; bare /v1/info (and /v1/info/<static happ id>) means the statically configured network.
  • Purely-dynamic services are first-class: no config.roles needed — register networks at runtime only.
  • Joins name a network via the network body field (a happ_id; the static id is equivalent to omitting it). Sessions are scoped per (agent, network): one key can join any number of networks; same-network re-joins 409. Non-empty allowed_agents gates a network (403 join_rejected for unlisted agents) — and gated networks do NOT expose their role modifiers via info; joiners get them at provision.
  • Provision serves the session's network's roles, membrane proofs, and its own happ_bundle_url. dna_hash per role is required only when membrane proofs are enabled.
  • Session recovery: POST /v1/reconnect accepts an optional network and returns the agent's session token for that scope, so an agent that crashes between joining and installing can recover instead of being permanently stuck behind 409 agent_already_joined. Recovery is keyed on the one thing the client durably holds first — its agent private key — via the existing signed-timestamp check; /v1/join is deliberately NOT made idempotent (that would turn any public agent key into a provisioning credential). JoiningClient gains reconnectAndProvision().
  • New NetworkStore (memory / sqlite networks.db / Cloudflare KV) and network_registration.admin_secret config; admin middleware is path-scoped (three-way coexistence with linker and agent admin surfaces regression-tested).
  • Cross-network dna_hash uniqueness is enforced at registration (409 duplicate_dna_hash, checked against other registered networks and the static roles) — duplicate hashes would let one agent be provisioned twice for the same cell, forking the chain. Best-effort on KV (eventual consistency), noted in docs.

Breaking / operational notes:

  • Client behavior change: after discover(), join() defaults network to the document's happ_id. Verify an app domain's .well-known happ_id matches the service's static happ id (or a registered network) before deploying; pass network: null to suppress the default.
  • Each network join re-recheks the key with hc-auth — deployments with hc_auth.required want the hc-auth client misreads /request-auth status codes #15 status-mapping fix first.
  • Reconnect's replay window now has a larger blast radius: the signed payload is the timestamp alone (300s tolerance) and ready sessions never expire, so an observed reconnect body replayed within the window now yields a session token — i.e. network_config, the bundle URL, and a mintable membrane proof, where previously it yielded only public infra URLs. Deliberately unchanged here (changing the signed payload breaks existing signers); worth a follow-up to cover agent_key and a nonce.
  • network_config (bootstrap/relay/auth-server URLs) remains service-wide.
  • No sqlite migration ships, by design. The sessions table gains a network column and networks.db is new, but nothing upgrades an existing file — there is no deployed data to preserve at this stage. Anyone who ran an earlier build locally should delete sessions.db and networks.db before running this one. Sessions are ephemeral (pending ones expire on TTL), so there is nothing worth carrying across.

Stacks on the role-centric config PR (#18) and the allowed-agent registration PR (#17) — merge those first (this branch is based on both).

Followed by feat/worker-kv-parity, which brings the Cloudflare Worker entry point up to parity with Node. That work was split out of this branch: it touches only deploy/ and the auth-plugin construction shared by both entry points, and reviews independently.


Commits (7)

Commit Concern
feat: NetworkStore with memory, sqlite, and KV backends The store, happ_id-keyed, three backends
feat: multi-network registration with network-aware join and provision Admin routes, app wiring, per-network info, conditional dna_hash
feat: enforce dna_hash uniqueness across registered networks 409 duplicate_dna_hash
feat: pass network through JoiningClient and provision CLI Client + CLI
docs: network registration, happ_id identity, and network-aware joining API + CLI docs; corrects CLI.md's single-network assumptions about where roles come from
feat: scope join sessions by agent and network, with reconnect recovery Session scoping + reconnect token
docs: per-network session semantics and reconnect recovery

Diff vs base: 37 files, +3479/-103.

zippy added 8 commits August 17, 2026 14:19
Brings the role-centric `roles` config in as the base for network
registration: registered networks describe their DNAs with the same
per-role shape the static config uses.
Networks are keyed by `happ_id` -- the same identity space as a conductor
installed-app id -- so a registered network needs no separate slug. Each
record carries the per-role DNA config in the same shape as the static
`roles` config, optional hApp metadata for the info endpoint, and an
optional `allowed_agents` list.

The three backends mirror the session and allowed-agent stores: memory for
tests and ephemeral deployments, sqlite for single-node, KV for Workers.
#13)

A joining service could serve exactly one network -- the one in its static
config. Networks registered at runtime through POST /v1/admin/networks are
now first-class: join, provision, and the info endpoint all take a
`network` naming a registered happ_id, and resolve roles, membrane proofs,
and allowed_agents from that network's record instead of the static config.

The static network keeps working untouched. Naming it explicitly by its own
happ_id collapses to the same scope as omitting `network` (normalizeNetwork),
so one agent is never split across two spellings of one network. Registering
a network under the service's own happ id is rejected, and startup warns if
one is already stored that way.

Role `dna_hash` is required only when membrane proofs are enabled, matching
the rule the static roles config already follows.
Network identity (happ_id) stands in for DNA hash but nothing enforced it:
registering a network with a dna_hash already used elsewhere let one agent
join both and receive membrane proofs for the same cell twice, a chain-fork
risk. POST /v1/admin/networks now rejects a candidate dna_hash that collides
with any other registered network or the service's static roles with 409
duplicate_dna_hash, naming the hash and the owning network.

Duplicates within one registration's own roles remain allowed (same-DNA
multi-role apps exist), and re-registering a happ_id against its own prior
hashes is unaffected.
JoiningClient.join() and the provision CLI can now name a network. The
client defaults `network` to the happ_id it discovers from GET /v1/info, so
a client pointed at a service backed by a registered network joins that
network without the caller spelling it out; passing an explicit value
overrides the discovered id, and an explicit null omits the key entirely.
…ng (#13)

Documents the admin network API, the `network` parameter on join and
provision, per-network `GET /v1/info/:happ_id`, and dynamic-only mode
(network_registration with no static roles). A network's identity is its
happ_id -- the same identity space as a conductor installed-app id -- so
`.well-known`'s existing happ_id is what routes a client to its network.

Covers the conditional `dna_hash` rule (required only under membrane
proofs) and the cross-network uniqueness rule with its 409
duplicate_dna_hash error, including why the check is best-effort on KV.

CLI.md gets the same treatment rather than just a `--network` flag row:
its prose still said role names come from the service config and that the
service "must be configured with roles", both of which stop being true
once a session names a network -- provision reads roles, modifiers, and
happ_bundle_url from that network's registration record instead, and a
dynamic-only service has no static roles to read.
…ry (#13)

Session uniqueness was keyed on agent key alone, so an agent that joined one
network got 409 agent_already_joined when joining another -- a single-network
assumption left over in the session layer. Lookups are now scoped to
(agent_key, network): findByAgentKey matches the exact network scope
(undefined network is its own no-network scope, byte-identical to prior
behavior), and a new findAnyByAgentKey covers call sites like reconnect where
linker and gateway URLs are service-wide rather than per-network.

Reconnect builds on that scoping to close a recovery gap: an agent that
completed join and crashed before calling provision had no way back, since a
fresh join only returns 409. POST /v1/reconnect now also returns the session
token for the requested network's ready session, and the client gains
reconnectAndProvision() for the one-call version of that path. The URL refresh
and the token lookup are gated separately, so an agent whose only session names
a non-static network still gets its URLs.

The sqlite `sessions` table gains a `network` column and a composite
(agent_key, network) index. No migration ships for it: there is no deployed
data at this stage, and sessions are ephemeral anyway, so an old sessions.db
should be deleted rather than upgraded.
Session uniqueness is per (agent, network): a second join on a different
network succeeds instead of 409ing, and agent_already_joined now means
re-joining a network the agent already has a live session on.

Documents reconnect's `network` parameter, the `session` token it returns,
and the crash-recovery path that token exists for. The URL refresh and the
session-token lookup are gated separately, so the three outcomes are spelled
out rather than left to be inferred from the error table.
@ThetaSinner

Copy link
Copy Markdown

Yes, this is too big to review in a sensible amount of time. Accepting it as it is and I'll test it when it's made available downstream.

Base automatically changed from feat/dynamic-progenitor-registration to main August 21, 2026 12:49
# Conflicts:
#	JOINING_SERVICE_API.md
#	src/cli/provision.ts
@zippy
zippy merged commit aedf1ad into main Aug 21, 2026
1 check passed
@zippy
zippy deleted the feat/multi-network branch August 21, 2026 15:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants