Skip to content

Compare federation UTXOs against guardian wallet claims (backend only) - #131

Open
Juwon-Ogunseye wants to merge 2 commits into
fedimint:masterfrom
Juwon-Ogunseye:guardian-utxo-reconciliation
Open

Juwon-Ogunseye wants to merge 2 commits into
fedimint:masterfrom
Juwon-Ogunseye:guardian-utxo-reconciliation

Conversation

@Juwon-Ogunseye

Copy link
Copy Markdown

Backend implementation for reconciling the observer's own on-chain UTXO view against what each guardian's wallet module actually reports, surfacing genuine discrepancies.

What this does
Fetches each guardian's live wallet_summary (using the existing per-guardian request pattern from guardians.rs)
Compares the observer's own reconstructed UTXO set (/federations/:id/utxos) against guardian claims from the spendable and unconfirmed_change buckets
Deliberately excludes unsigned_peg_out/unsigned_change (in-flight withdrawals) from the comparison, since mismatches there are expected and not a sign of a real problem
Surfaces four categories of discrepancy: unclaimed by any guardian, claimed by only one guardian, amount mismatches, and genuine guardian conflicts (guardians reporting different amounts for the same UTXO)
New endpoint: GET /federations/:federation_id/utxo_report
Tested against real, live federations

Ran this against several federations already tracked by a local dev instance. One early version of the comparison logic mislabeled "multiple guardians agreeing on something the observer's own index hadn't caught up on yet" as a guardian disagreement — fixed by splitting that into its own honest category (missed_by_observer), separate from genuine guardian_conflict (which only fires on real amount mismatches between guardians). After the fix, tested against a real federation mid-sync and got a clean, correctly-categorized result: 96 missed_by_observer (explained entirely by observer sync lag) and 7 unclaimed_by_any_guardian, zero false guardian_conflict results.

What's intentionally not included yet

No frontend. This PR is backend-only — the new endpoint is only reachable directly (e.g., via curl), nothing is surfaced in the UI yet. Flagging this explicitly rather than leaving it implicit, since the related work in #116 does include frontend changes. Wanted to get the backend logic and its output reviewed first before building UI on top of it, happy to follow up with a frontend PR once the approach here looks right.

No automated tests yet — everything's been verified manually against live federations so far.

The 7 unclaimed_by_any_guardian cases haven't been individually root-caused yet — plausible benign explanations (recent deposits, in-flight spends) but worth digging into a couple concretely before considering this fully validated.

@bansalayush247

Copy link
Copy Markdown
Member

Thanks for working on this. I think the separation between missed_by_observer and a genuine guardian_conflict is especially useful.

I’ve been working on the same guardian-vs-observer UTXO reconciliation problem in #116, but my approach also addresses the correctness of the observer-side UTXO inventory itself.

While investigating this, I found a historical transaction ID byte-order issue in the existing database. On an existing observer database, some UTXOs currently appear unspent in the observer view but match recorded withdrawal inputs once the stored txid bytes are normalized. Because of that, I’m concerned that some missed_by_observer or unclaimed_by_any_guardian results could be caused by the observer’s reconstruction rather than an actual guardian discrepancy.

That’s why #116 includes the reconstruction changes and the v11 migration. The goal isn’t to make reconciliation more complicated for its own sake, but to make sure the observer inventory being reconciled is itself correct for existing historical data.

I think your reconciliation logic is useful and could fit well with the reconstruction work in #116. In particular, I like the explicit distinction between observer lag and genuine guardian conflicts.

It would also be useful to add automated tests around these cases, especially observer lag, guardian amount conflicts, and the historical txid byte-order case.

One additional suggestion: instead of having specific claimed_by_guardian_only / unclaimed_by_any_guardian classifications, could we represent how many guardians claim and do not claim each UTXO?

For example, with 5 guardians:

claimed_by: 3
unclaimed_by: 2

This seems more general than treating “claimed by only one guardian” as a special case, and it scales naturally across different federation sizes.

I’d also suggest calculating the denominator from successfully responding guardians. An unreachable guardian should be represented separately rather than treated as “not claiming” the UTXO.

That would give us a more useful reconciliation signal, such as 3/4 responding guardians claim this UTXO, while still keeping the distinction between missing guardian responses and actual inventory differences.

@Juwon-Ogunseye

Juwon-Ogunseye commented Sep 28, 2026 •

Copy link
Copy Markdown
Author

@bansalayush247 thanks for this genuinely appreciate the depth here, and the txid byte-order find is important. You're right that some of my missed_by_observer/unclaimed_by_any_guardian results could be artifacts of that rather than real guardian discrepancies, so I don't think this PR's data can be considered fully validated until that's accounted for.

Agree on all three suggestions:

Switching to a claimed_by/unclaimed_by count instead of the binary categories you're right that scales better and is the more honest representation
Separating unreachable guardians from "didn't claim it" good catch, conflating those was a real gap in my logic
Adding automated tests for observer lag, real conflicts, and the byte-order case

On sequencing since #116 includes the actual reconstruction fix and v11 migration, does it make sense for me to rebase this on top of #116 once it's further along, rather than duplicating that fix here? Happy to adjust the categorization and add tests in the meantime regardless, but want to avoid us solving the byte-order issue twice in parallel PRs.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants