fix(cnpg): detect a fully-down cluster instead of showing it as starting - #1614
Conversation
CNPG omits status.readyInstances when it is 0 (omitempty), so a cluster with zero ready instances serializes identically to one whose status was never written. The badge and the issue detector both gated the all-down path on the field being present, so a fully-down cluster was unreachable there: it fell through to the transient-phase branch, showed amber Starting Instances forever, and raised no issue. Distinguish the cases by whether the operator has reported and whether a primary was ever elected (status.currentPrimary, which CNPG sets on first election and never clears): - no status yet -> Unknown (unchanged; absence is still not zero) - reported, zero ready, no primary yet -> first bootstrap, amber, no alarm - reported, zero ready, primary present -> was up and now down, red badge and a Critical issue Hibernated and fully-fenced clusters are intentionally at zero ready, so they read neutral. A short grace on the Ready=False transition absorbs a routine single-instance restart before alarming. One shared availability read drives the badge, the drawer banner, the row count, and the cell colour, and the Go detector mirrors it, so the four surfaces agree. The golden fixtures now use the real omitted-ready wire shape rather than a fabricated zero. Ticket: RAD-403. Claude-Session: https://claude.ai/code/session_01KxZ3xt2G4KpexrKSoXQ91S
PR Summary by QodoFix detection of fully-down CNPG clusters with omitted ready counts
AI Description
Diagram
High-Level Assessment
Files changed (10)
|
Code Review by Qodo
1.
|
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit fcdcdbb. Configure here.
| ) { | ||
| return null | ||
| } | ||
| return 'down' |
There was a problem hiding this comment.
Wake-up treated as cluster outage
Medium Severity
Waking a hibernated cluster or lifting a full fence still matches the was-up outage path. currentPrimary is never cleared and Ready=False keeps its old lastTransitionTime, so the 5-minute grace is already elapsed. Both the badge and the detector then raise a Critical CNPGClusterDegraded for an operator-intended restart.
Additional Locations (1)
Reviewed by Cursor Bugbot for commit fcdcdbb. Configure here.
There was a problem hiding this comment.
Confirmed reachable, and verified against CNPG source: hibernation keeps currentPrimary and leaves Ready=False with a stale lastTransitionTime, and the hibernation marker (annotation and condition) is removed on wake — so a wake-up is indistinguishable from a genuine outage on the Cluster status alone. No CR-only signal separates them without either missing real node-loss outages or adding detector state the stateless badge cannot share. Decision: accept the transient window (it self-clears once the first instance is Ready) to keep every real outage covered; documented at the verdict in ef7b2ea. A Go-only observation-time debounce is the follow-up if on-call noise proves unacceptable.
… change history
The comments explaining CNPG's readyInstances-at-zero omission referenced the
prior implementation ("old okR gate", "old presence gate", "made unreachable").
Restate them as the current WHY: CNPG omits readyInstances (omitempty), so
absence on a cluster that reported anything else is a real 0, and currentPrimary
distinguishes a was-up regression from a first bootstrap.
Claude-Session: https://claude.ai/code/session_01KxZ3xt2G4KpexrKSoXQ91S
A cluster woken from hibernation or lifted from a full fence briefly presents the same shape as an outage; CNPG has dropped the hibernation marker by then, so it cannot be told apart from the Cluster status alone. Document the limitation where the verdict is decided, on both the badge and the detector. Claude-Session: https://claude.ai/code/session_01KxZ3xt2G4KpexrKSoXQ91S
…ing (skyhook-io#1614) ## Problem CNPG omits `status.readyInstances` when it is 0 (`omitempty` int), so a cluster with **zero ready instances serializes identically to one whose status was never written** — the field is simply absent in both. Radar's badge (`getCNPGClusterStatus`) and the Go issue detector (`detectCNPGClusterIssues`) both gated the all-down path on the field being *present*, so a fully-down cluster never reached it: it fell through to the transient-phase branch, showed amber "Starting Instances" forever, and raised no issue. A dead database, silent. (The existing "all-down" tests hid this by fabricating `readyInstances: 0`, a shape CNPG never emits.) ## Fix Distinguish the cases the absent field collapses together, using `status.currentPrimary` (CNPG sets it on first primary election and never clears it — the version-robust "was up" signal, present on 1.27/1.28): - **No status yet** → Unknown (unchanged — absence is still not zero) - **Reported, 0 ready, no primary elected** → first bootstrap → amber, no alarm - **Reported, 0 ready, primary present** → was up and now down → red badge + **Critical** issue - **Hibernated / fully-fenced** (`cnpg.io/hibernation`, `cnpg.io/fencedInstances: ["*"]`) → intentionally 0 ready → neutral, no alarm A **5-minute grace** on the `Ready=False` transition absorbs a routine single-instance restart before alarming — the same on-read, stateless grace mechanism CNPG's existing timers already use (`conditions.FindFalseConditionWithTime`). It escalates immediately if there is no `Ready` condition to time from. One shared availability read drives the **badge, drawer banner, row count, and cell colour**, and the **Go detector mirrors it field-for-field**, so all surfaces agree. Golden fixtures now use the real *omitted*-ready wire shape rather than a fabricated zero. ## Tests Rewrote the shared `testdata/cnpg/badge-issue-matrix.json` to realistic omitted-ready shapes and added: was-up fully down (→ red + Critical), first bootstrap (→ amber, no issue), slow restore past grace (→ amber), hibernated, all-fenced, plus the unchanged healthy / partial / status-never-written cases. The was-up-down / hibernated / fenced cases fail against the pre-fix code. Full k8s-ui suite and `go test ./internal/issues/...` pass. ## Two known, non-blocking edges 1. A 1-instance, was-up cluster whose only pod is restarting *within* the grace while its phase still lags at "healthy" can briefly show `0/1` next to a green "Healthy" badge. Narrow, self-heals in ≤5 min, not a false page or a silent outage. 2. The cell colour branch has no dedicated unit test (brittle class-name assertion avoided); the underlying `getCNPGClusterAvailability` is thoroughly covered. Ticket: RAD-403 (Velero/CNPG/Kyverno review). https://claude.ai/code/session_01KxZ3xt2G4KpexrKSoXQ91S <!-- CURSOR_SUMMARY --> --- > [!NOTE] > **Medium Risk** > Changes core CNPG health and alerting semantics across UI and the issue detector; incorrect parity could cause missed outages or false criticals, though behavior is heavily pinned by the shared golden matrix. > > **Overview** > Fixes **false “starting” and silent outages** when CloudNativePG omits `status.readyInstances` at zero — the same wire shape as “status not written yet.” > > **Shared availability logic** (`getCNPGClusterAvailability` in the UI, mirrored in `detectCNPGClusterIssues`) now treats omitted ready as **0 only after the operator has reported** (phase, `currentPrimary`, or conditions), uses **`currentPrimary` as “was up”** to separate first bootstrap from regression, and applies carve-outs for **hibernation** and **all-instance fencing** plus a **5-minute `Ready=False` grace** (kept in sync with `cnpgDownGrace` / `CNPG_DOWN_GRACE_MS`). > > **Surfaces aligned:** list badge, instances cell (including red styling), cluster drawer “down” banner, and Go **`CNPGClusterDegraded` Critical** for total outage; partial shortfalls still need an explicit ready count above zero. > > **Tests:** shared `badge-issue-matrix.json` uses realistic omitted-ready fixtures and optional `metadata.annotations`; Go golden tests and TS golden/unit tests extended for bootstrap, hibernated, fenced, grace, and was-up-down cases. > > <sup>Reviewed by [Cursor Bugbot](https://cursor.com/bugbot) for commit ef7b2ea. Bugbot is set up for automated code reviews on this repo. Configure [here](https://www.cursor.com/dashboard/bugbot).</sup> <!-- /CURSOR_SUMMARY -->


Problem
CNPG omits
status.readyInstanceswhen it is 0 (omitemptyint), so a cluster with zero ready instances serializes identically to one whose status was never written — the field is simply absent in both. Radar's badge (getCNPGClusterStatus) and the Go issue detector (detectCNPGClusterIssues) both gated the all-down path on the field being present, so a fully-down cluster never reached it: it fell through to the transient-phase branch, showed amber "Starting Instances" forever, and raised no issue. A dead database, silent.(The existing "all-down" tests hid this by fabricating
readyInstances: 0, a shape CNPG never emits.)Fix
Distinguish the cases the absent field collapses together, using
status.currentPrimary(CNPG sets it on first primary election and never clears it — the version-robust "was up" signal, present on 1.27/1.28):cnpg.io/hibernation,cnpg.io/fencedInstances: ["*"]) → intentionally 0 ready → neutral, no alarmA 5-minute grace on the
Ready=Falsetransition absorbs a routine single-instance restart before alarming — the same on-read, stateless grace mechanism CNPG's existing timers already use (conditions.FindFalseConditionWithTime). It escalates immediately if there is noReadycondition to time from.One shared availability read drives the badge, drawer banner, row count, and cell colour, and the Go detector mirrors it field-for-field, so all surfaces agree. Golden fixtures now use the real omitted-ready wire shape rather than a fabricated zero.
Tests
Rewrote the shared
testdata/cnpg/badge-issue-matrix.jsonto realistic omitted-ready shapes and added: was-up fully down (→ red + Critical), first bootstrap (→ amber, no issue), slow restore past grace (→ amber), hibernated, all-fenced, plus the unchanged healthy / partial / status-never-written cases. The was-up-down / hibernated / fenced cases fail against the pre-fix code. Full k8s-ui suite andgo test ./internal/issues/...pass.Two known, non-blocking edges
0/1next to a green "Healthy" badge. Narrow, self-heals in ≤5 min, not a false page or a silent outage.getCNPGClusterAvailabilityis thoroughly covered.Ticket: RAD-403 (Velero/CNPG/Kyverno review).
https://claude.ai/code/session_01KxZ3xt2G4KpexrKSoXQ91S
Note
Medium Risk
Changes core CNPG health and alerting semantics across UI and the issue detector; incorrect parity could cause missed outages or false criticals, though behavior is heavily pinned by the shared golden matrix.
Overview
Fixes false “starting” and silent outages when CloudNativePG omits
status.readyInstancesat zero — the same wire shape as “status not written yet.”Shared availability logic (
getCNPGClusterAvailabilityin the UI, mirrored indetectCNPGClusterIssues) now treats omitted ready as 0 only after the operator has reported (phase,currentPrimary, or conditions), usescurrentPrimaryas “was up” to separate first bootstrap from regression, and applies carve-outs for hibernation and all-instance fencing plus a 5-minuteReady=Falsegrace (kept in sync withcnpgDownGrace/CNPG_DOWN_GRACE_MS).Surfaces aligned: list badge, instances cell (including red styling), cluster drawer “down” banner, and Go
CNPGClusterDegradedCritical for total outage; partial shortfalls still need an explicit ready count above zero.Tests: shared
badge-issue-matrix.jsonuses realistic omitted-ready fixtures and optionalmetadata.annotations; Go golden tests and TS golden/unit tests extended for bootstrap, hibernated, fenced, grace, and was-up-down cases.Reviewed by Cursor Bugbot for commit ef7b2ea. Bugbot is set up for automated code reviews on this repo. Configure here.