Skip to content

ci(apk): cache SDK dl/staging/build trees per target + job timeouts (#370 follow-up) - #385

Merged
c03rad0r merged 3 commits into
OpenTollGate:mainfrom
felixfelix-bot:ci/sdk-cache-on-main
Sep 13, 2026
Merged

c03rad0r merged 3 commits into
OpenTollGate:mainfrom
felixfelix-bot:ci/sdk-cache-on-main

Conversation

@felixfelix-bot

@felixfelix-bot felixfelix-bot commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

Lands #370's content on main.

#370 is marked MERGED on GitHub (2026-08-30T22:43:30Z, merge commit
3f53827a) but none of it is on main:

$ git merge-base --is-ancestor 3f53827 origin/main        # false, exit 1
$ git diff 3f53827^1 3f53827 -- .github/workflows/build-package.yml
# 16 insertions(+), 0 deletions(-)

Why the merge did not carry. #370's base was the stacked
ci/trigger-hygiene branch (#369). #369 landed on main as the squash
db8af35 at 2026-08-30T22:42:36Z — 54 seconds before #370 merged into
that same, already-merged branch, so #370's commits ended up on a dead
branch head and main never saw them. main today has no actions/cache
step in this workflow and a timeout-minutes on publish-metadata only,
so every package-apk leg runs against the 360-minute default.

This is the gap flagged on #339 (the v0.6.0 release request): a tag cut from
main builds without the SDK build-tree cache and without the apk timeout.

What this PR changes

#370's 16 lines, minus one line that is stale on main:

job change
compile-binaries timeout-minutes: 30
build-portal timeout-minutes: 15
package-ipk timeout-minutes: 30
package-apk timeout-minutes: 90
package-apk SHA-pinned actions/cache@v5 step restoring /builder/dl, /builder/staging_dir and /builder/build_dir, keyed apk-sdk-<matrix.sdk>-25.12.0-<hashFiles('packaging/**')> with restore-keys fallback

The one omitted line, and why. git cherry-pick 3f53827 does apply
cleanly onto main and produces 16 insertions — but one of them re-declares
a job-level timeout-minutes: 15 on publish-metadata, which main
already carries for that job (present since 218cd2d, #97; #370's base
branched before it). Taken verbatim that is a duplicated mapping key:

$ actionlint -shellcheck= .github/workflows/build-package.yml
key "timeout-minutes" is duplicated in "publish-metadata" job.
   previously defined at line:682,col:5 [syntax-check]

So this branch applies the four new job timeouts plus the cache step:
15 insertions, 0 deletions, actionlint clean. If you would rather have
the literal 16 lines, cherry-pick 3f53827 directly and close this PR — a
duplicated YAML key is not something I wanted to land on main on the
strength of "the parser probably tolerates it".

Verification

$ git diff --numstat origin/main HEAD -- .github/workflows/build-package.yml
15	0	.github/workflows/build-package.yml
$ python3 -c "import yaml;yaml.safe_load(open('.github/workflows/build-package.yml'))"
# OK - 8 jobs
$ actionlint -shellcheck= .github/workflows/build-package.yml   # exit 0, no messages

timeout-minutes lands on compile-binaries (30), build-portal (15),
package-ipk (30), package-apk (90) and the pre-existing
publish-metadata (15) — nowhere else, and there are no duplicated mapping
keys anywhere in the file.

Every added line is one of #370's own added lines: a multiset diff of this
branch's added lines against git diff 3f53827^1 3f53827 shows 15 of the 16
present and zero lines that are not in #370's delta — the one difference
is timeout-minutes: 15, exactly the publish-metadata duplicate
described above.

No GitHub Actions run: this repository's last workflow run of any kind is
from 2026-08-27 and nothing is queued for this PR, so the checks above are
the available evidence (local actionlint 1.7.x + PyYAML, both clean).

Second commit (docs-only, pushed)

CHANGELOG.md and RELEASE-NOTES.md both described #370's change as
already shipped — the changelog bullet came from #373 (as an [Unreleased]
entry, moved into [v0.6.0-alpha2] by #380) and the release-notes bullet
"the SDK build tree is cached and every heavy job has a timeout" was
rewritten by #380. Neither is true of the tree those documents describe, so
both entries now name this PR as the change that carries #370's content onto
main, and the changelog no longer lists a publish-metadata timeout among
#370's additions. Kept as a separate commit so the CI change can be
cherry-picked on its own, and easy to split into its own PR if you prefer.

Whole-branch diffstat:

 .github/workflows/build-package.yml | 15 +++++++++++++++
 CHANGELOG.md                        | 22 ++++++++++++--------
 RELEASE-NOTES.md                    |  7 +++++--
 3 files changed, 34 insertions(+), 10 deletions(-)

Overlap note

Open PR #383 also edits .github/workflows/build-package.yml (and
CHANGELOG.md). This branch touches only the four job headers, the one
cache step and two documentation paragraphs — flagging the overlap so the
two can merge in whichever order is convenient.

…e#370 follow-up)

PR OpenTollGate#370 is marked MERGED on GitHub but none of its content is on `main`:

  git merge-base --is-ancestor 3f53827 origin/main   # false
  git diff 3f53827^1 3f53827 -- .github/workflows/build-package.yml
  # 16 insertions(+), 0 deletions(-)

Cause: OpenTollGate#370's base was the stacked `ci/trigger-hygiene` branch (OpenTollGate#369). OpenTollGate#369
landed on `main` as the squash db8af35 54 seconds before OpenTollGate#370 merged into
that already-merged branch, so OpenTollGate#370's commits were orphaned and `main` never
saw them. `main` still has no SDK cache and no timeout on any heavy job
except `publish-metadata` (pre-existing since 218cd2d), which means
`package-apk` falls back to the 360-minute default.

This applies OpenTollGate#370's own delta. One of its 16 lines is NOT applied: OpenTollGate#370 also
re-declares a job-level `timeout-minutes: 15` on `publish-metadata`, which
`main` already carries at the same job (introduced by 218cd2d, OpenTollGate#97). Taking
it verbatim duplicates the mapping key - actionlint:

  key "timeout-minutes" is duplicated in "publish-metadata" job.

Net change: 15 insertions, 0 deletions.

  - compile-binaries: timeout-minutes: 30
  - build-portal:     timeout-minutes: 15
  - package-ipk:      timeout-minutes: 30
  - package-apk:      timeout-minutes: 90
  - package-apk: SHA-pinned actions/cache@v5 step restoring /builder/dl,
    /builder/staging_dir and /builder/build_dir, keyed
    apk-sdk-<matrix.sdk>-25.12.0-<hashFiles('packaging/**')> with
    restore-keys fallback.

No other changes.
…n via OpenTollGate#385

CHANGELOG.md's `[v0.6.0-alpha2]` entry (moved out of `[Unreleased]` by
OpenTollGate#373) and RELEASE-NOTES.md's "the SDK build tree is cached and every heavy
job has a timeout" bullet (rewritten by OpenTollGate#380) both describe OpenTollGate#370 as
shipped. It is not on `main`: OpenTollGate#370 merged into the stacked
`ci/trigger-hygiene` branch *after* OpenTollGate#369 had already been squash-merged as
`db8af35`, so none of OpenTollGate#370's 16 lines were in the tree either document
describes.

Both entries now name OpenTollGate#385, the PR that carries OpenTollGate#370's content onto
`main`, and the changelog entry no longer lists a `publish-metadata`
timeout among OpenTollGate#370's additions (that job has carried
`timeout-minutes: 15` since `218cd2d`). The release headings, section
order and every other line are untouched.

Refs OpenTollGate#370, OpenTollGate#339.
@c03rad0r

Copy link
Copy Markdown
Contributor

@felixfelix-bot the actions runners rely on github secrets. How is that handled in the ngit version of the actions runner?

@felixfelix-bot

Copy link
Copy Markdown
Contributor Author

Short answer: ngit-ci does support ${{ secrets.NAME }} — it just sources them from the coordinator host instead of from a GitHub secret store, and only a confirmed maintainer of the repo can set them.

The mechanism (per the ngit-ci repo). Per-repo secrets are exposed to jobs of a repo whose #ALIAS appears in NGIT_CI_REPOS. Values are supplied out of band, keyed by that alias:

  • Docker/compose deploy: NGIT_CI_SECRET_<ALIAS>__<NAME>=<value> lines in a git-ignored ngit-ci-secrets.env, loaded as the coordinator's env_file. The compose file carries no secret names itself (docker-compose.yml, "Per-repo secrets" comment).
  • NixOS deploy: one runtime secret file per value, e.g. repoSecrets.NGIT.API_TOKEN = "/run/agenix/...", or a secretsEnvironmentFile of NGIT_CI_SECRET_<ALIAS>__<NAME>=*** lines (agenix/sops-nix).
  • Never on argv: secret values are deliberately rejected as command-line flags because /proc/<pid>/cmdline is world-readable — env, .env, or systemd credentials only (AGENTS.md, secret-handling note).

Jobs then reference them exactly as on GitHub — env: NSEC_HEX: ${{ secrets.NSEC_HEX }} — so a ported workflow does not change its syntax.

Authorization is the difference from GitHub. docs/secret-authorization.md defines a secret-handling boundary built on a confirmed-maintainership rule with an authorization scenario table and a regression index; docs/configuration.md documents the per-repo secrets section. Maintainers can also push encrypted Repository Secret Updates over Nostr inbox relays (disabled when the relay list is empty), optionally with bunker-sealed values (an established nbunksec, NIP-46), so rotation does not require shell access to the runner host.

What that means for this repo. The lane already ported to .ngit/act/workflows/test.yml needs no secrets (no container:, no GITHUB_TOKEN), which is exactly why it was the clean port. build-package.yml is a different story and is deliberately not ported: it uses secrets.NSEC_HEX in five places, a REPO_ACCESS_TOKEN for repository-dispatch, and container: jobs — the last of which the ngit-ci default refuses (its job containerisation is owned by act, and nested-docker is a separate opt-in backend). So the package/publish lane would need, in order: container support on the runner, a confirmed maintainer to provision NGIT_CI_SECRET_<ALIAS>__NSEC_HEX (and the dispatch token) on the coordinator, and then the workflow port.

One caveat specific to us: the nsec in this checkout is not a maintainer of the ngit mirror of this repo, so I can't provision those secrets myself — it needs a maintainer identity (or the maintainer-role mapping established first). Worth knowing before anyone promises the apk/Blossom lane on ngit CI.

None of this affects this PR: the cache and timeout changes are GitHub-Actions-side (actions/cache, timeout-minutes), and they land the content #370 lost when its stacked base was squash-merged 54 seconds earlier — that diagnosis checks out.

@c03rad0r

Copy link
Copy Markdown
Contributor

Short answer: ngit-ci does support ${{ secrets.NAME }} — it just sources them from the coordinator host instead of from a GitHub secret store, and only a confirmed maintainer of the repo can set them.

The mechanism (per the ngit-ci repo). Per-repo secrets are exposed to jobs of a repo whose #ALIAS appears in NGIT_CI_REPOS. Values are supplied out of band, keyed by that alias:

* Docker/compose deploy: `NGIT_CI_SECRET_<ALIAS>__<NAME>=<value>` lines in a git-ignored `ngit-ci-secrets.env`, loaded as the coordinator's `env_file`. The compose file carries no secret names itself (`docker-compose.yml`, "Per-repo secrets" comment).

* NixOS deploy: one runtime secret file per value, e.g. `repoSecrets.NGIT.API_TOKEN = "/run/agenix/..."`, or a `secretsEnvironmentFile` of `NGIT_CI_SECRET_<ALIAS>__<NAME>=***` lines (agenix/sops-nix).

* Never on argv: secret _values_ are deliberately rejected as command-line flags because `/proc/<pid>/cmdline` is world-readable — env, `.env`, or systemd credentials only (`AGENTS.md`, secret-handling note).

Jobs then reference them exactly as on GitHub — env: NSEC_HEX: ${{ secrets.NSEC_HEX }} — so a ported workflow does not change its syntax.

Authorization is the difference from GitHub. docs/secret-authorization.md defines a secret-handling boundary built on a confirmed-maintainership rule with an authorization scenario table and a regression index; docs/configuration.md documents the per-repo secrets section. Maintainers can also push encrypted Repository Secret Updates over Nostr inbox relays (disabled when the relay list is empty), optionally with bunker-sealed values (an established nbunksec, NIP-46), so rotation does not require shell access to the runner host.

What that means for this repo. The lane already ported to .ngit/act/workflows/test.yml needs no secrets (no container:, no GITHUB_TOKEN), which is exactly why it was the clean port. build-package.yml is a different story and is deliberately not ported: it uses secrets.NSEC_HEX in five places, a REPO_ACCESS_TOKEN for repository-dispatch, and container: jobs — the last of which the ngit-ci default refuses (its job containerisation is owned by act, and nested-docker is a separate opt-in backend). So the package/publish lane would need, in order: container support on the runner, a confirmed maintainer to provision NGIT_CI_SECRET_<ALIAS>__NSEC_HEX (and the dispatch token) on the coordinator, and then the workflow port.

One caveat specific to us: the nsec in this checkout is not a maintainer of the ngit mirror of this repo, so I can't provision those secrets myself — it needs a maintainer identity (or the maintainer-role mapping established first). Worth knowing before anyone promises the apk/Blossom lane on ngit CI.

None of this affects this PR: the cache and timeout changes are GitHub-Actions-side (actions/cache, timeout-minutes), and they land the content #370 lost when its stacked base was squash-merged 54 seconds earlier — that diagnosis checks out.

Does this mean that the builds can only be published by a maintainers npub if the CI runs on the maintainers infrastructure? I'm fine with that constraint, but I want to understand what the constraints are and why they are that way...

@felixfelix-bot

Copy link
Copy Markdown
Contributor Author

Following up on the secrets question with the exact rules, since they're stricter than "a maintainer sets it".

Who may cause a secret to be injected (docs/configuration.md#maintainer-gating, docs/secret-authorization.md#authorization-rule): secrets are injected only on runs whose trigger event was authored by a confirmed maintainer of the repo, and recursive maintainership including invited maintainers is insufficient — the trigger author needs a current kind-30617 announcement completing a reciprocal path into the selected view. The rule is three-party: the trigger author, the secret owner, and the acting maintainer must all be in one recursive maintainership excluding invited maintainers. Under request-required the run is additionally scoped by an accepted Service Request; mixed-identifier triggers receive no secrets.

Consequences worth stating plainly:

  • A contributor's PR can never receive secrets, wherever the runner runs. Not a gap — that's the point: untrusted code can't exfiltrate a publishing key by opening a PR.
  • The publishing key's npub (the "secret owner") must itself be inside that maintainership. It need not be the same person as the trigger author, but all three have to share one confirmed view. So in practice: yes, publishing identity must be a maintainer-controlled npub.
  • A job that fails these checks still runs — the secret scope is withheld, not the workflow.
  • Relay confidence is part of it: at least one relay announced by that maintainership must be connected with completed control-history EOSE, or the secret is withheld.

Where the plaintext lives (#provisioning-secret-values, #receiving-secrets-over-nostr, #bunker-sealed-secret-values): on the coordinator host — NGIT_CI_SECRET_<ALIAS>__<NAME> in env/.env (chmod 600, double underscore is the delimiter) or systemd LoadCredential=<ALIAS>__<NAME>:/path files (agenix/sops-nix; credentials win over env) — or pushed at runtime as encrypted Repository Secret Updates over configured inbox relays, accepted only while the signer is in the current maintainership excluding invited. There is deliberately no CLI flag for secret values because argv is world-readable via /proc/<pid>/cmdline. Bunker sealing goes further: values are stored sealed to a per-scope keypair whose private half is wrapped to a NIP-46 bunker's user pubkey, so the coordinator cannot decrypt on its own — each run costs exactly one nip44_decrypt immediately before runner handoff, with plaintext only in RAM.

So what are the constraints, and why? Two separate trust boundaries:

  1. Untrusted code with a key — anyone can open a PR. Answered by maintainer-gated triggers, per-repo alias scoping (secrets never cross repos), and no default GITHUB_TOKEN.
  2. Untrusted host with a key — whoever operates the CI host can read a job's environment, so hosting elsewhere hands that operator the ability to publish as that npub. Answered by keeping values out of argv, 600/credential-file provisioning, Nostr-encrypted delivery, and bunker sealing to shrink exposure to one decrypt per run.

Neither constraint requires the CI to run on the maintainer's own machine. It requires that the publishing identity is a maintainer-controlled key and that the plaintext only exists on infrastructure the maintainers trust, for the shortest possible time. Self-hosting satisfies that trivially; using someone else's runner means trusting that operator with the key unless values are bunker-sealed.

For this repo specifically: the package/publish lane would still need container: job support (containerisation is owned by act; nested docker is a separate opt-in backend) plus a maintainer provisioning NGIT_CI_SECRET_<ALIAS>__NSEC_HEX and the dispatch token — and note the nsec in the working checkout here is not a maintainer of the ngit mirror, so that provisioning needs a maintainer identity (or the maintainership extended). The ported test lane needs none of this, which is why it was the clean port.

@c03rad0r

Copy link
Copy Markdown
Contributor

2.

Neither constraint requires the CI to run on the maintainer's own machine. It requires that the publishing identity is a maintainer-controlled key and that the plaintext only exists on infrastructure the maintainers trust, for the shortest possible time. Self-hosting satisfies that trivially; using someone else's runner means trusting that operator with the key unless values are bunker-sealed.

For this repo specifically: the package/publish lane would still need container: job support (containerisation is owned by act; nested docker is a separate opt-in backend) plus a maintainer provisioning NGIT_CI_SECRET_<ALIAS>__NSEC_HEX and the dispatch token — and note the nsec in the working checkout here is not a maintainer of the ngit mirror, so that provisioning needs a maintainer identity (or the maintainership extended). The ported test lane needs none of this, which is why it was the clean port.

Are any secrets being leaked to the actions runners that we use on ngit right now in this repo? How is the publishing step of the action handled?

@felixfelix-bot

Copy link
Copy Markdown
Contributor Author

Following this up with what's actually live, since the ngit publish lane is about to be relied on. Audited the coordinator and the workflow trees rather than reasoning from the docs.

Are secrets currently exposed to the ngit runners? No — the surface is one secret and it is gated

  • Exactly one secret exists for this repo: NGIT_CI_SECRET_TMBG__NSEC_HEX (mode-600 file, one line). No PAT, no dispatch token, nothing else. The #TMBG alias is bound to this repo's coordinate in NGIT_CI_REPOS.
  • Nostr secret intake is disabled (NGIT_CI_NOSTR_SECRET_RELAYS empty) → no recipient is advertised and the secrets-key capability is omitted, so no new secret can be pushed over Nostr.
  • Bunker sealing is not enabled for this alias (no WORKFLOW_SECRETS_DECRYPTION_BUNKER), so the coordinator holds decryptable values.
  • No key material lives in the workflow files. ${{ secrets.NSEC_HEX }} appears at 13 sites across .ngit/act/workflows/* (the binaries stage, both halves of build-package.yml, the three apk files and the four upx files), all on publish/upload steps. A scan for literal nsec1…/64-hex values finds nothing.
  • Injection is gated, so "leaked to the runners" isn't the right frame: the trigger must be maintainer-authored, the three-party maintainership rule must hold, and a failed gate withholds the secret while still running the job. Per-repo alias scoping means the value cannot cross into another repo's jobs.
  • Consequence worth knowing: a contributor branch never receives it. So the apk lanes failing on ci/ngit-split are failing without any secret — their failures are code/env problems, not credential problems.

Residual risks, stated plainly:

  1. The runner host can read the plaintext — the coordinator decrypts and hands it into the job's env, so anything with root/docker on the coordinator host can read it. Bunker sealing exists to remove exactly that capability and is currently off.
  2. The gate keys on the trigger author, not the workflow content, so a maintainer-authored push can add a .ngit/act/workflows/*.yml that exfiltrates the key. Nothing reviews .ngit/** today — worth a required review/contract check on .ngit/** and .github/**.
  3. Worth confirming the publish steps fail loudly rather than silently skipping when the secret is withheld.
  4. tests/cloud-lab/configs/{reseller,upstream}-identities.json contain "privatekey": "nsec1fgaau8m9xd7…q3qj6q3qj6…". The repeating pattern says it's a stub, but a stub shaped like a real nsec trips secret scanners and masks real findings — an obviously-invalid form would be better.

How the publishing step is handled on the ngit lane

It's already ported, and the handover between stages is Nostr records instead of CI artifacts — that's what makes the same publish path work on both substrates:

  • The release pipeline is split into two workflow files because ngit-ci has no cross-file needs and actions/upload-artifact fails there (Unable to get the ACTIONS_RUNTIME_TOKEN env variable). Stage 1 build-package-binaries.yml: cross-compile the five GOARCH/GOARM/GOMIPS targets, build the portal assets, mirror both to Blossom, then publish build-id records. Stage 2 build-package.yml: the ipk legs, publish-metadata, os-handoff.
  • Stage 2 discovers stage 1's output by polling nak req -k 30078 --tag d=tollgate-build/<BUILD_ID>/… rather than downloading artifacts — so the two files meet on a Nostr record keyed by the build id.
  • The publish call itself is nak blossom upload --server <mirror> --sec "$NSEC_HEX" <file>, fanned across four mirrors (blossom.primal.net, blossom.psbt.me, blossom2.orangesync.tech, drive.cashu.email; blossom1 was dropped as unreliable) with a quorum required, followed by the build-id records and then the apk/ipk announcements.
  • No GITHUB_TOKEN and no repository-dispatch on the ngit side. The REPO_ACCESS_TOKEN the GitHub twin needs for dispatch has no counterpart here — the record handover replaces it.

One operational gotcha for anyone iterating on these lanes

The .ngit/act/workflows/*.yml files declare on: push: branches: [main] (plus a bare pull_request:). A push of a PR branch therefore triggers nothing — which is why a GitHub PR for this repo has no ngit CI results at all. The two ways to get a run: a pull_request event (no branch filter), or a manual replay, which ignores on: filters entirely — the latter needs a recent ngit CLI (v3+; v2.6.1 has no ci subcommand). Historic runs here are almost all trigger=manual for that reason.

@c03rad0r

Copy link
Copy Markdown
Contributor

The publish call itself is nak blossom upload --server <mirror> --sec "$NSEC_HEX" <file>, fanned across four mirrors (blossom.primal.net, blossom.psbt.me, blossom2.orangesync.tech, drive.cashu.email; blossom1 was dropped as unreliable) with a quorum required, followed by the build-id records and then the apk/ipk announcements.

Is it using a random nsec_hex in this step or how does it do the publishing step?

@felixfelix-bot

Copy link
Copy Markdown
Contributor Author

One more clarification, since the publish identity matters for provenance: $NSEC_HEX is not a random key. It is a fixed, provisioned identity — the same key on every run.

What it is. It's injected from the per-repo secret NGIT_CI_SECRET_TMBG__NSEC_HEX (a hex private key held mode-600 on the coordinator, bound to this repo by the #TMBG alias in NGIT_CI_REPOS). Its public identity is npub1dn798sztmf743h2d6pr36ehk5n4860sj8euqqmswpsdtcysg4sxs6eyvuj. Every upload and every announcement the lane makes is therefore attributable to that one persistent npub — which is the point: consumers verify provenance by filtering on author.

What that single identity does (six steps):

  • Upload binaries to Blossom (per-arch, mirrored) → nak blossom upload --server <mirror> --sec "$NSEC_HEX" <file>. The key signs the Blossom auth event, so the stored blob is attributed to that pubkey; fanned across four mirrors with a quorum required.
  • Publish portal assets to Blossom and record the hash — same pattern for the captive-portal assets.
  • Publish the stage-2 rendezvous record → a kind-30078 addressable record keyed by build id, which is precisely what stage 2 polls for (nak req -k 30078 --tag d=tollgate-build/<BUILD_ID>/…). The inter-stage handover is a signed Nostr record, not a CI artifact.
  • Upload to Blossom and announce → kind-1063 NIP-94 file-metadata events tagged n, x/ox, A, v, c, compression, m, filename.
  • Verify the announcements against the build records → nak req -k 1063 --tag n=$PACKAGE_NAME -a "$(nak key public "$NSEC_HEX")", i.e. it verifies the announcement set authored by that same key.
  • Publish the tollgate-os handoff record → kind-30078.

Four consequences worth designing around:

  1. The publisher npub is not the mirror owner. The mirror's maintainer is npub1x677kgul…; the publisher is npub1dn798…. Under the authorization rule the secret owner must sit in the same confirmed maintainership as the trigger author and the acting maintainer, so that publisher identity has to be a confirmed co-maintainer of the repo. If it is not, the secret is withheld, the job still runs, and the publish step fails without credentials — a plausible contributor to the currently-red apk lanes.
  2. The GitHub twin's NSEC_HEX cannot be verified from outside (Actions secrets are write-only). If the two substrates hold different keys, artifacts published from GitHub and from ngit CI carry different provenance. That should be aligned deliberately rather than assumed.
  3. Rotation is a breaking change for consumers. Because everything is attributed to one npub, replacing the key leaves previously published artifacts attributed to the old identity and downstream filters must be updated — and anyone holding the key can publish as it. That is the practical argument for bunker-sealing the value, which is currently not enabled for this alias.
  4. As a sanity check, a -k 1063 query by that publisher returned nothing on relay.ngit.dev, gitnostr.com and relay.primal.net — consistent with the package/publish lanes never having completed, though it may also mean the announcements go to other relays.

@c03rad0r

Copy link
Copy Markdown
Contributor

One more clarification, since the publish identity matters for provenance: $NSEC_HEX is not a random key. It is a fixed, provisioned identity — the same key on every run.

What it is. It's injected from the per-repo secret NGIT_CI_SECRET_TMBG__NSEC_HEX (a hex private key held mode-600 on the coordinator, bound to this repo by the #TMBG alias in NGIT_CI_REPOS). Its public identity is npub1dn798sztmf743h2d6pr36ehk5n4860sj8euqqmswpsdtcysg4sxs6eyvuj. Every upload and every announcement the lane makes is therefore attributable to that one persistent npub — which is the point: consumers verify provenance by filtering on author.

What that single identity does (six steps):

* `Upload binaries to Blossom (per-arch, mirrored)` → `nak blossom upload --server <mirror> --sec "$NSEC_HEX" <file>`. The key signs the Blossom auth event, so the stored blob is attributed to that pubkey; fanned across four mirrors with a quorum required.

* `Publish portal assets to Blossom and record the hash` — same pattern for the captive-portal assets.

* `Publish the stage-2 rendezvous record` → a **kind-30078** addressable record keyed by build id, which is precisely what stage 2 polls for (`nak req -k 30078 --tag d=tollgate-build/<BUILD_ID>/…`). The inter-stage handover is a signed Nostr record, not a CI artifact.

* `Upload to Blossom and announce` → **kind-1063** NIP-94 file-metadata events tagged `n`, `x`/`ox`, `A`, `v`, `c`, `compression`, `m`, `filename`.

* `Verify the announcements against the build records` → `nak req -k 1063 --tag n=$PACKAGE_NAME -a "$(nak key public "$NSEC_HEX")"`, i.e. it verifies the announcement set authored by that same key.

* `Publish the tollgate-os handoff record` → kind-30078.

Four consequences worth designing around:

1. **The publisher npub is not the mirror owner.** The mirror's maintainer is `npub1x677kgul…`; the publisher is `npub1dn798…`. Under the authorization rule the secret owner must sit in the same confirmed maintainership as the trigger author and the acting maintainer, so that publisher identity has to be a confirmed co-maintainer of the repo. If it is not, the secret is withheld, the job still runs, and the publish step fails **without credentials** — a plausible contributor to the currently-red apk lanes.

2. **The GitHub twin's `NSEC_HEX` cannot be verified from outside** (Actions secrets are write-only). If the two substrates hold different keys, artifacts published from GitHub and from ngit CI carry different provenance. That should be aligned deliberately rather than assumed.

3. **Rotation is a breaking change for consumers.** Because everything is attributed to one npub, replacing the key leaves previously published artifacts attributed to the old identity and downstream filters must be updated — and anyone holding the key can publish as it. That is the practical argument for bunker-sealing the value, which is currently not enabled for this alias.

4. As a sanity check, a `-k 1063` query by that publisher returned nothing on `relay.ngit.dev`, `gitnostr.com` and `relay.primal.net` — consistent with the package/publish lanes never having completed, though it may also mean the announcements go to other relays.

How about we just use a different npub for the time being to unblock the CI and we figure out how to safely use the maintainer's npub at a later point in time?

@Amperstrand

Copy link
Copy Markdown
Collaborator

Review pass against PR-REVIEW.md:

The forensic story checks out — I re-ran the checks myself: git merge-base --is-ancestor 3f53827a origin/main exits 1, and main's workflow carries zero actions/cache steps, so #370's content genuinely never landed and every package-apk leg has been running against the 360-minute default. Recovering it as this PR does is the right fix, and the CHANGELOG/RELEASE-NOTES corrections tell the story so the next person doesn't re-derive it.

The mechanics are sound: the cache action is SHA-pinned (immutable ref), and the cache key includes hashFiles('packaging/**') — which matters more than it looks here, because packaging/build-inputs.json lives under that glob, so an SDK digest bump automatically invalidates warm trees. Timeouts (30/15/30/90) match reality: our measured cold SDK apk builds run 15–25 min on big iron, so 90 min leaves headroom for a slow runner. Caching /builder/dl, staging_dir, and build_dir should eliminate the toolchain-package (libc/libgcc) rebuild cost we measured as the long pole.

One cross-PR verification ask (not blocking, but it should be demonstrated once Actions is back, or in an ngit run): show that a warm-cache apk build is byte-identical to a cold-cache one. #383 just established byte-level reproducibility for the apk lane with cold, fresh containers; this PR reintroduces carried state into /builder. #383's normalization covers the staged package tree and propagates SOURCE_DATE_EPOCH, so I expect warm trees to be inert for the tollgate-wrt artifact — but "expected" is the word #383 exists to delete. A one-off repro-test.sh apk run with a pre-warmed builder directory (or an explicit warm/cold A-B) would close it.

Disposition: land, with the warm/cold byte-equality check as a follow-up.

@c03rad0r
c03rad0r self-requested a review September 13, 2026 19:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants