Skip to content

fix(treatment): province beds, occupancy and admissions on one capped basis - #969

Open
seabbs-bot wants to merge 22 commits into
mainfrom
feat/province-admissions-forecast
Open

seabbs-bot wants to merge 22 commits into
mainfrom
feat/province-admissions-forecast

Conversation

@seabbs-bot

@seabbs-bot seabbs-bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Province beds, occupancy and admissions now use one capped basis. Each province quantity sums to the national one. This PR carries the changes from #942.

What

Fit

  • Admissions are scored as a plain negative binomial, and admission_headroom is removed. The old free-bed censor never bound on a fitted day.
  • Each day's 24h province admissions are fitted as a split of the national admissions. The split is on each patch's modelled admissions, A_bvd_p + w_p·A_bg after the admission delay, with its own ρ (province_admissions).

Cut-off beds and occupancy

  • Each province's beds are its modelled capacity floored at its last recorded effective beds.
  • Each province holds its demand share of the reported-scale occupancy (demand plus the reclassification offset, floored at zero), capped at its beds. The rest is its shortfall (cutoff_occupancy).
  • The national beds, occupancy and shortfall are the province sums, so utilisation never exceeds 100%. With one patch, the beds are the modelled capacity floored at the last recorded capacity.

Forecast

  • Each province is a stock capped at its beds (capped_stock_forecast), starting from its cut-off occupancy.
  • Its beds are its modelled capacity floored at its cut-off beds, taken as a running maximum so they never fall.
  • The in-care deaths, recoveries, rule-outs and absconds are scaled by the occupied beds over the uncapped occupancy mean the day before, and shared across provinces by occupancy.
    • Below the beds the scale is 1, so the stock follows the fitted occupancy mean and the flows are the fitted ones.
  • A province admits its modelled admissions up to its free beds: its beds, less its previous occupancy, plus its exits.
  • Occupancy is drawn by province censored at its beds, and admissions censored at the free beds of the mean path, not each draw's own free beds. The national counts are the sums (forecast_isolation.province.obs, forecast_admissions.province.obs, forecast_province_beds).

Data

  • Effective province beds are the printed beds. Where the patients held exceed them, they are raised to the larger of the patients and the rate-implied beds. Only Nord-Kivu changes.
  • Occupancy rates checked against the PDFs: SitRep 092 Nord-Kivu 80.1, 093 Tshopo 66.7 and 095 Nord-Kivu 87.9. The scan had taken a ward's rate.
  • SitReps 081 to 083 Nord-Kivu admissions checked against the PDFs: 22, 54 and 56. The scan had taken the confirmed admissions only.
  • New province_admissions_history block and loader, regenerated with scripts/province_care_manifest.jl from main's CSVs to SitRep 135.

Why

Checks

  • Scoped TestItemRunner runs on 4eb46ae7 (2773 pass, 0 fail): all of test_forecast_provinces.jl and test_province_care.jl, plus the touched items in test_forecast_horizon.jl, test_isolation.jl and test_mooncake_rules.jl. No fits were run locally.
  • A new test checks the forecast stock below its beds: exit scale 1, and the stock tracks the uncapped occupancy.
  • Tests now cover the forecast shortfall, both over the beds and under them.
  • All seven review-bot findings on 126d07fb are taken in 301fc5fe.
  • The earlier red checks, the test_zone_plots.jl legend test and the province page's province_capacity_share_sd KeyError, came from the old base and are fixed on main. The Windows job lost its runner.

Review findings

First review (5345882617, on 126d07fb): all seven findings fixed in 301fc5fe.

Second review (5361409245, on 7d8225b1), all in 4eb46ae7 unless stated:

  • Exit scale above 1, or silently 1 on a non-positive uncapped occupancy: fixed. The scale is capped at 1, documented, and tested in both cases.
  • forecast_bed_shortfall duplicating cutoff_occupancy: fixed. It now calls cutoff_occupancy, and a test covers the bed_shortfall column of forecast_reported. The shortfall stays on the uncapped occupancy by design, since it measures unmet demand, not the stock.
  • have_cap == false untested and underdocumented: fixed. The docstring is cut and a test covers that path.
  • The silent nothing in forecast_provinces: fixed. It throws on any length other than one national row or one row per patch. The existing test already checks that the province isolation and admissions sum to the national forecast.
  • The inline one-patch bed floor: fixed. It is now _national_bed_floor, with a zero-derivative rule and tests.
  • Narration comments and long docstring lines, including the missing comma in plot_province_split_ppc: fixed.
  • The comment on censored draws pointing at callers: fixed.
  • The test including scripts/province_effective_beds.jl: rejected. Ten other test files include scripts the same way (for example test_scoring.jl and test_province_lab_parser.jl). The rule prepares data and is not model code, so it stays in scripts/.

Fit comparison with main

The current CI fit on 4eb46ae7 (run 36674507618) passes the convergence gate with warnings.
This head also includes #1013, which changes the onset reporting walk.

fit max R-hat min bulk / tail ESS divergences
joint 1.06 49 / 32 2
zone 1.03 146 / 109 0

For context on the slow capacity-step mixing (#1016), the earlier CI fit on 7d8225b1 (run 36615565487) is set against main's fit on 73b1b398 (run 36603384866), the commit merged in there.

#969 7d8225b1 main 73b1b398
Max R-hat 1.144 (cap_state.steps[3]) 1.087 (cap_state.steps[3])
Min bulk / tail ESS 16 / 22 23 / 20
Divergences 0 0
Chain step sizes 0.0006, 0.0013 0.0004, 0.0015
cap_share_state τ_cap and z_cap bulk ESS 41–44 70–96
C_T mean by chain 18.8k, 18.8k 17.9k, 18.4k
  • Both fits fail the gate on the same parameter, an early-June knot of the national capacity walk.
    The PR does not change the national bed data or the walk.
  • In both fits steps[3] is split between chains by about 0.3 sd, and the chains agree on the log density.
    Main's fit on 442b0592 mixed the same knot at ESS 168, so its mixing depends on the draw.
  • The capacity-share group mixes worse here but stays above the failure threshold.
    The effective Nord-Kivu beds move the shares (τ_cap 0.75 against 0.98), and one chain visits the neck of the non-centred τ·z funnel.
  • The slow mixing of the capacity steps and shares on main is Capacity walk steps and capacity shares mix slowly in the joint fit #1016.

Closes #640
Closes #918
Closes #920
Closes #958

This was opened by a bot. Please ping @seabbs for any questions.

seabbs-bot and others added 9 commits September 26, 2026 23:17
…scale

Admissions are scored as a plain negative binomial in the fit, its
predictive checks and the forecast. The free-bed headroom censor never
bound a fitted day but capped the predictive draws (#918).

The cut-off occupancy is the reported-scale mean, the demand plus the
reclassification offset, and the bed shortfall is the demand above the
modelled capacity (#640).

Province occupied beds are the national occupancy split on the demand
shares, as the split likelihood and the forecast use, rather than the
demand capped at the printed beds (#920).
The fit scores admissions uncensored. The forecast censors them at the
modelled capacity less the previous day's occupancy, so forecast
admissions cannot exceed the beds available. Guard the province
occupancy split against a zero demand total and test that province
occupied beds sum to the national occupancy.

Co-authored-by: Sam Abbott <contact@samabbott.co.uk>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Forecast admissions are capped at the capacity less the previous day's
occupancy plus the day's modelled deaths, recoveries, rule-outs and
absconds, since a bed freed during the day can be refilled.

Co-authored-by: Sam Abbott <contact@samabbott.co.uk>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e 100%

Move the forecast admissions ceiling into `_admission_ceilings` and test
that it follows the previous day's occupancy, adds the day's exits and
floors at half a patient. State that the national utilisation can exceed
100% now the cut-off occupancy is on the reported scale.

Co-authored-by: Sam Abbott <contact@samabbott.co.uk>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each province bed entry is the largest of the printed beds, the beds
implied by the printed occupancy rate and the patients held (#958). The
CSVs keep the printed counts.

Co-authored-by: Sam Abbott <contact@samabbott.co.uk>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The cut-off occupancy is the reported-scale mean, demand plus the
reclassification offset, capped at the modelled capacity. Patients above
the beds are unmet demand, carried by the bed shortfall, so utilisation
stays at or below 100% as the national data do.

Co-authored-by: Sam Abbott <contact@samabbott.co.uk>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nd utilisation

The cut-off beds are the modelled capacity floored at the recorded cap on
the last occupancy day, the bound the forecast starts from. The cut-off
occupancy is capped at them, and the shortfall and utilisation are taken
against them, so the occupancy never sits below an observed count.

Co-authored-by: Sam Abbott <contact@samabbott.co.uk>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ince's free beds

With more than one patch, forecast admissions are drawn per province at
its modelled admissions, censored at its beds less its previous day's
occupancy plus its exits that day, and the national forecast is their sum
(#958). Province beds are the modelled capacity floored at the last
recorded effective beds. The national cap adds the day's exits.

Co-authored-by: Sam Abbott <contact@samabbott.co.uk>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seabbs-bot
seabbs-bot requested a review from seabbs as a code owner September 28, 2026 18:08
@seabbs-bot
seabbs-bot marked this pull request as draft September 28, 2026 18:09
seabbs-bot and others added 6 commits September 28, 2026 19:25
The forecast reports the floored beds its caps use, the shortfall is on
the reported scale the occupancy is, and the last recorded cap is
computed once in the treatment state.

Co-authored-by: Sam Abbott <contact@samabbott.co.uk>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ons-forecast

# Conflicts:
#	docs/pages/methods.jl
#	src/models/joint.jl
#	src/models/observations.jl
#	test/test_mooncake_rules.jl
…ove the printed beds

The implied beds count only when the patients held exceed the printed
beds, so a rounded rate no longer inflates them (#958). The three
occupancy rates the scan took from a ward (SitReps 092, 093, 095) are
settled against the PDFs and recorded in the scanned CSV, and the
manifest reconciles the rates like the other cells.

Co-authored-by: Sam Abbott <contact@samabbott.co.uk>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`province_admissions_history` holds each province's 24h admissions,
reconciled from the two reads by the manifest builder and read by
`load_observations` (#958). The scan's Nord-Kivu counts for SitReps
081-083 took the confirmed admissions alone; the totals are settled
against the PDFs and recorded in the scanned CSV.

Co-authored-by: Sam Abbott <contact@samabbott.co.uk>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…f and in the forecast

Each patch's beds are its modelled capacity floored at its recorded
beds. At the cut-off each patch holds its demand share of the reported
occupancy up to its beds, the rest its shortfall, and the national
figures are the sums. The forecast runs each patch as a stock capped at
its beds from that cut-off occupancy: it admits up to its free beds and
loses its share of the in-care exits, each scaled by the occupied beds
over the demand. Occupancy and admissions are drawn by patch and summed
to national, and the in-care deaths and rule-outs are the scaled flows
(#958).

Co-authored-by: Sam Abbott <contact@samabbott.co.uk>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Benchmark comparison vs main

Minimum time per call, per model component. Both revisions were benchmarked by AirspeedVelocity in one job on one machine, so a ratio here compares two revisions and not two runners.

Buckets are PR time as a % of main, so lower is faster (🟢 faster, ⚪ within 20%, 🔴 slower).

The memory column is bucketed to 0.5% rather than that band. Allocations are deterministic, so a change there is real rather than sampling noise.

Warning

The band is capped. 90% of benchmarks varied by less than 20.3% across their own samples, which is wider than this comment will call neutral. Treat everything below as unresolved.

Note

That band is a lower bound on what this harness can resolve. AirspeedVelocity times each revision once, so the spread is dispersion within a revision's own samples, not drift between the two revisions, which are separated by the time it takes to compile the second one's gradients.

Counts of benchmarks per bucket:

Group 🟢 <50% 🟢 50–75% 🟢 75–80% ⚪ 80–120% 🔴 120–125% 🔴 125–150% 🔴 >150%
Log density · · · 17 · · ·
Mooncake · · · 17 · · ·

Across all 34 benchmarks the ratio runs 0.99× to 1.09×, median 1.0×.

Log density — 17 benchmarks (by time change)
Benchmark main PR time spread memory
Log density / Latent / patch_infection_model (uncoupled) 29.46 μs 30.77 μs ⚪ 1.04× 8.3% ⚪ 1.0×
Log density / Latent / latent (infection + onset staging) 15.26 μs 15.06 μs ⚪ 0.99× 10.0% ⚪ 1.0×
Log density / Composer / confirmed_only_model 29.52 μs 29.19 μs ⚪ 0.99× 6.3% ⚪ 1.0×
Log density / Joint / bvd_joint 80.94 μs 81.73 μs ⚪ 1.01× 6.0% ⚪ 1.0×
Log density / Composer / treatment_only_model 42.56 μs 42.88 μs ⚪ 1.01× 10.1% ⚪ 1.0×
Log density / Latent / patch_infection_model (coupled) 24.81 μs 24.93 μs ⚪ 1.0× 28.3% ⚪ 1.0×
Log density / Composer / onsets_only_model 23.3 μs 23.18 μs ⚪ 1.0× 7.8% ⚪ 1.0×
Log density / Submodel / treatment_flow_model 17.53 μs 17.59 μs ⚪ 1.0× 9.2% ⚪ 1.0×
Log density / Submodel / onset_reporting_model 6.78 μs 6.8 μs ⚪ 1.0× 8.7% ⚪ 1.0×
Log density / Composer / cases_only_model 22.62 μs 22.68 μs ⚪ 1.0× 15.0% ⚪ 1.0×
Log density / Composer / deaths_only_model 37.23 μs 37.32 μs ⚪ 1.0× 6.0% ⚪ 1.0×
Log density / Submodel / exports_model 4.79 μs 4.78 μs ⚪ 1.0× 35.5% ⚪ 1.0×
Log density / Submodel / deaths_model 15.17 μs 15.14 μs ⚪ 1.0× 10.9% ⚪ 1.0×
Log density / Composer / exports_only_model 23.5 μs 23.47 μs ⚪ 1.0× 6.4% ⚪ 1.0×
Log density / Submodel / province_composition_model 1.26 μs 1.26 μs ⚪ 1.0× 11.4% ⚪ 1.0×
Log density / Submodel / confirmed_cases_model 4.02 μs 4.02 μs ⚪ 1.0× 13.3% ⚪ 1.0×
Log density / Submodel / reported_cases_model 4.68 μs 4.68 μs ⚪ 1.0× 5.0% ⚪ 1.0×
AD gradients — 17 benchmarks (by time change)
Benchmark main PR time spread memory
AD gradients / Composer / treatment_only_model / Mooncake 317.01 μs 345.23 μs ⚪ 1.09× 20.3% 🔴 1.37×
AD gradients / Joint / bvd_joint / Mooncake 656.6 μs 697.36 μs ⚪ 1.06× 10.8% 🟢 0.88×
AD gradients / Submodel / onset_reporting_model / Mooncake 69.92 μs 72.71 μs ⚪ 1.04× 21.5% ⚪ 1.0×
AD gradients / Submodel / province_composition_model / Mooncake 5.02 μs 5.16 μs ⚪ 1.03× 9.4% ⚪ 1.0×
AD gradients / Submodel / treatment_flow_model / Mooncake 119.9 μs 123.07 μs ⚪ 1.03× 9.3% 🔴 1.01×
AD gradients / Composer / confirmed_only_model / Mooncake 137.28 μs 140.27 μs ⚪ 1.02× 9.9% ⚪ 1.0×
AD gradients / Latent / patch_infection_model (coupled) / Mooncake 115.28 μs 117.29 μs ⚪ 1.02× 5.9% ⚪ 1.0×
AD gradients / Latent / latent (infection + onset staging) / Mooncake 50.98 μs 51.85 μs ⚪ 1.02× 12.2% ⚪ 1.0×
AD gradients / Submodel / confirmed_cases_model / Mooncake 39.81 μs 40.27 μs ⚪ 1.01× 8.0% ⚪ 1.0×
AD gradients / Composer / cases_only_model / Mooncake 153.45 μs 151.85 μs ⚪ 0.99× 11.8% ⚪ 1.0×
AD gradients / Submodel / reported_cases_model / Mooncake 15.86 μs 15.71 μs ⚪ 0.99× 4.5% ⚪ 1.0×
AD gradients / Submodel / exports_model / Mooncake 17.92 μs 17.77 μs ⚪ 0.99× 5.0% ⚪ 1.0×
AD gradients / Latent / patch_infection_model (uncoupled) / Mooncake 107.65 μs 108.54 μs ⚪ 1.01× 5.5% ⚪ 1.0×
AD gradients / Composer / exports_only_model / Mooncake 151.47 μs 150.29 μs ⚪ 0.99× 12.3% ⚪ 1.0×
AD gradients / Composer / deaths_only_model / Mooncake 179.78 μs 180.9 μs ⚪ 1.01× 11.4% ⚪ 1.0×
AD gradients / Composer / onsets_only_model / Mooncake 138.84 μs 139.32 μs ⚪ 1.0× 8.1% ⚪ 1.0×
AD gradients / Submodel / deaths_model / Mooncake 45.53 μs 45.57 μs ⚪ 1.0× 4.0% ⚪ 1.0×

spread is the interquartile range of that benchmark's own samples, as a fraction of its median.

…ional admissions

Each day's printed province admissions are scored as a split of their sum
on each patch's modelled admissions, with their own overdispersion, and
the recovery generator simulates them (#958). The methods, news and the
in-sample province page cover the split and the capped cut-off and
forecast.

Co-authored-by: Sam Abbott <contact@samabbott.co.uk>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Parameter recovery: unconverged (3 seeds)

1 of 3 fits converged (R-hat at most 1.1, bulk ESS at least 30).
An unconverged seed's recovery is shown but not judged.
The truth is inside the 90% interval in 88 of 96 quantity-seed pairs (92%) and outside the 99% interval in 1.
Forecasts: 12 scored, median CRPS relative to persistence 0.12 (below one beats it).

National
Quantity Truth Relative error z Truth in 90% Largest z
CFR 0.11 to 0.45 -0.06 (-0.10 to 0.05) -0.41 (-0.68 to 0.29) 3/3 seed 3 (-0.68)
C_T 2,796 to 9,256 -0.06 (-0.17 to 0.23) -0.04 (-0.80 to 1.18) 3/3 seed 3 (1.18)
R_T 0.77 to 1.06 0.10 (0.01 to 0.30) 0.67 (0.01 to 1.28) 3/3 seed 1 (1.28)
T 197.2 to 231.5 0.00 (-0.01 to 0.02) 0.17 (-0.04 to 0.46) 3/3 seed 2 (0.46)
growth_state.G 14.1 to 14.2 0.02 (0.01 to 0.04) 0.24 (0.13 to 0.61) 3/3 seed 2 (0.61)
lambda_bg 0.34 to 11.6 0.20 (-0.06 to 0.64) 0.58 (0.16 to 1.26) 3/3 seed 2 (1.26)
onset_report_state.τ 0.68 to 1.18 0.02 (-0.02 to 0.02) 0.43 (-0.65 to 0.79) 3/3 seed 1 (0.79)
p_drc 0.60 to 0.91 0.15 (-0.12 to 0.21) 0.60 (-1.07 to 1.04) 3/3 seed 3 (-1.07)
r -0.0179 to 0.00445 0.98 (0.27 to 1.53) 0.66 (-0.03 to 1.40) 3/3 seed 1 (1.40)
rt_state.intervention_effect -0.69 to -0.46 0.05 (-0.08 to 0.55) 0.12 (-0.11 to 1.68) 3/3 seed 1 (1.68)
rt_state.sigma_rw 0.0333 to 0.26 -0.16 (-0.62 to 0.01) 0.01 (-4.39 to 0.14) 2/3 outside 99% in 1 seed 3 (-4.39)
tau_test 0.54 to 0.96 -0.05 (-0.15 to 0.04) -0.64 (-1.72 to 0.26) 2/3 seed 3 (-1.72)
Province
Quantity Truth Relative error z Truth in 90% Largest z
CFR_patch[Haut-Uele] 0.11 to 0.45 -0.06 (-0.14 to 0.13) -0.37 (-0.72 to 0.56) 3/3 seed 3 (-0.72)
CFR_patch[Ituri] 0.12 to 0.45 -0.08 (-0.08 to 0.02) -0.52 (-0.54 to 0.15) 3/3 seed 3 (-0.54)
CFR_patch[Nord-Kivu] 0.12 to 0.46 -0.07 (-0.07 to 0.01) -0.44 (-0.45 to 0.11) 3/3 seed 2 (-0.45)
CFR_patch[Other provinces] 0.11 to 0.45 -0.03 (-0.12 to 0.05) -0.14 (-0.77 to 0.28) 3/3 seed 3 (-0.77)
C_T_patch[Haut-Uele] 0.46 to 52.4 -0.16 (-0.62 to 0.35) -0.36 (-0.43 to 1.10) 3/3 seed 3 (1.10)
C_T_patch[Ituri] 2,491 to 8,953 -0.05 (-0.17 to 0.23) -0.00 (-0.80 to 1.18) 3/3 seed 3 (1.18)
C_T_patch[Nord-Kivu] 0.90 to 128.0 -0.13 (-0.57 to 0.32) -0.34 (-0.46 to 1.05) 3/3 seed 3 (1.05)
C_T_patch[Other provinces] 0.84 to 124.9 -0.10 (-0.58 to 0.27) -0.23 (-0.45 to 0.89) 3/3 seed 3 (0.89)
R_T_patch[Haut-Uele] 0.70 to 1.03 0.29 (0.02 to 0.43) 1.29 (0.16 to 1.45) 1/3 seed 3 (1.45)
R_T_patch[Ituri] 0.77 to 1.07 0.09 (0.01 to 0.30) 0.64 (-0.00 to 1.28) 3/3 seed 1 (1.28)
R_T_patch[Nord-Kivu] 0.77 to 1.03 0.29 (0.01 to 0.46) 0.99 (0.08 to 1.98) 2/3 seed 3 (1.98)
R_T_patch[Other provinces] 0.81 to 1.04 0.13 (0.01 to 0.21) 0.78 (0.02 to 0.82) 3/3 seed 1 (0.82)
province_ascertainment[Haut-Uele] 0.66 to 1.07 -0.02 (-0.08 to 0.53) -0.41 (-1.07 to 1.26) 2/3 seed 1 (1.26)
province_ascertainment[Ituri] 0.98 to 1.31 -0.09 (-0.23 to 0.02) -0.70 (-0.78 to 0.17) 3/3 seed 1 (-0.78)
province_ascertainment[Nord-Kivu] 0.94 to 1.08 0.03 (-0.07 to 0.08) 0.48 (-0.30 to 0.95) 3/3 seed 3 (0.95)
province_ascertainment[Other provinces] 0.84 to 1.08 -0.07 (-0.08 to 0.16) -0.35 (-0.72 to 1.53) 3/3 seed 2 (1.53)
region_drift_sd[Haut-Uele] 0.00124 to 0.0967 0.08 (-0.64 to 8.26) 0.47 (-1.63 to 0.92) 3/3 seed 1 (-1.63)
region_drift_sd[Ituri] 0.000798 to 0.0619 -0.32 (-0.56 to 12.93) -0.34 (-0.37 to 0.98) 2/3 seed 2 (0.98)
region_drift_sd[Nord-Kivu] 0.000575 to 0.0429 -0.22 (-0.62 to 17.26) -0.08 (-0.64 to 0.96) 2/3 seed 2 (0.96)
region_drift_sd[Other provinces] 0.00114 to 0.0199 0.71 (-0.44 to 9.53) 0.66 (-0.13 to 0.94) 3/3 seed 2 (0.94)
Fit per seed
Seed Verdict 90% coverage Outside 99% Fit (min) Max R-hat Min bulk ESS Divergences
1 unconverged 0.94 none 185.4 1.123 14 28
2 pass 0.94 none 185.9 1.013 378 48
3 unconverged 0.88 rt_state.sigma_rw 197.5 1.801 6 87
How the verdict is reached

A seed fails when more quantities lie outside their 99% interval than the 99th percentile of a Binomial with a 1% chance each (2 of 32), or fewer than 60% lie inside their 90% interval.
Relative error is (posterior median − truth) / |truth| and z is (posterior mean − truth) / posterior SD, each the median across seeds with the range in brackets.

This was opened by a bot. Please ping @seabbs for any questions.

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Fit convergence: ⚠️ passed with warnings

Every gated fit converged. The diagnostics below are past the warning thresholds, so the posterior is thinner than it should be, but the build is not failed on it.

Commit 1271355 — build log.

fit max R-hat min ESS bulk min ESS tail divergences draws
joint 1.06 49 32 2 2000
local 1.03 146 109 0 1600

joint: ⚠️ passed with warnings

  • warn — max R-hat is 1.06, past the warning threshold 1.05
  • warn — min bulk ESS is 49, past the warning threshold 100
  • warn — min tail ESS is 32, past the warning threshold 100
Worst-mixing parameters of `joint`
parameter rhat ess_bulk ess_tail
cases_state.positivity 1.062 49 149
confirmed_state.s_test 1.056 63 171
deaths_state.cfr_state.CFR 1.021 74 125
treatment_state.sev_state.δ_iso 1.047 76 109
recovered_state.rec_state.recovery_offset 1.021 79 131
background_total 1.04 79 181
treatment_state.isolation_bvd_admission 1.045 84 141
treatment_state.recovery_los_state.delay_sd 1.028 86 32
onset_report_state.σ_h0 1.032 89 132
treatment_state.β_iso 1.015 89 125

The same parameters grouped, one row per parameter:

parameter elements max_rhat min_ess_bulk above_1.05
cases_state.positivity 1 1.062 49 1
confirmed_state.s_test 1 1.056 63 1
deaths_state.cfr_state.CFR 1 1.021 74 0
treatment_state.sev_state.δ_iso 1 1.047 76 0
recovered_state.rec_state.recovery_offset 1 1.021 79 0
background_total 1 1.04 79 0
treatment_state.isolation_bvd_admission 1 1.045 84 0
treatment_state.recovery_los_state.delay_sd 1 1.028 86 0
onset_report_state.σ_h0 1 1.032 89 0
treatment_state.β_iso 1 1.015 89 0

Where the divergent transitions sit, in standard deviations of the full posterior:

parameter all_draws divergent_draws separation
growth_state.m 0.099–2.21 2.19–3.79 3.14
growth_state.T 1.32–29.1 27.6–47.1 2.94
T 182.0–210.0 209.0–228.0 2.94
growth_state.τ 8.09–20.1 17.5–29.6 2.92
rt_state.sigma_rw 0.0681–0.237 0.211–0.304 2.2
expected_onset_reported_T 5740.0–6060.0 6030.0–6130.0 1.92
deaths_state.asc_state.p_death 0.798–0.953 0.729–0.891 -1.72
confirmed_state.ρ 0.0368–0.0594 0.0542–0.0605 1.54
rt_state.log_R0 0.407–0.944 0.276–0.454 -1.53
growth_state.r 0.0345–0.0857 0.0238–0.0403 -1.43

Thresholds — fail: R-hat above 1.1, bulk ESS below 25, tail ESS below 25, divergences above 0.05 of draws; warn: 1.05, 100, 100, 0.01.

The full per-parameter breakdown is on the sensitivity page, under Fit diagnostics by parameter.

…and the national sums

Co-authored-by: Sam Abbott <contact@samabbott.co.uk>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@seabbs-bot seabbs-bot changed the title model(forecast): per-province forecast admissions capped at each province's free beds fix(treatment): province beds, occupancy and admissions on one capped basis Sep 28, 2026
@seabbs-bot
seabbs-bot marked this pull request as ready for review September 28, 2026 23:26

@seabbs-review-bot seabbs-review-bot Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The PR caps each province's occupancy at its beds at the cut-off and in the forecast. It drops the admissions censoring from the fit and adds a province admissions split. The logic reads as sound and the AD-relevant code uses min/max only. The gaps are one untested forecast quantity, a weak assertion, an untaped data helper and over-long docstrings.

Automated first pass by seabbs-review-bot (Claude sonnet), triggered by: re-review requested by seabbs-bot. Not a human review. Comment @seabbs-review-bot to ask for another pass: @seabbs any time, the author's agent once it has pushed changes. Add the no-review label to opt this PR out. Ping @seabbs with any questions.

Comment thread src/models/joint.jl
Comment thread test/test_forecast_horizon.jl Outdated
Comment thread src/models/observations.jl
Comment thread src/models/joint.jl Outdated
Comment thread src/models/joint.jl
Comment thread test/test_isolation.jl
Comment thread test/test_forecast_provinces.jl
Comment thread src/models/joint.jl
seabbs-bot and others added 3 commits September 29, 2026 18:36
…dmissions-forecast

The province blocks are regenerated with scripts/province_care_manifest.jl
from the merged CSVs, which adds SitReps 131-135 to the admissions block and
the effective Nord-Kivu beds for 26 September.

Co-authored-by: Sam Abbott <contact@samabbott.co.uk>
…ds fall

The forecast scaled the in-care exits by the occupied beds over the latent
demand, which carries no reclassification offset, so the scale was below one
even with no province at its beds. The in-care deaths and rule-outs were
scaled down and the stock drifted off the fitted occupancy mean. The scale is
now over the uncapped occupancy mean, so it is one below the beds. Forecast
beds are a running maximum, so the stock stays within them.

Also from review: a zero-derivative rule for `_province_bed_floors`, the
forecast shortfall computed directly, property tests for the cut-off
occupancy and shortfall, tests of the forecast shortfall and the forecast
beds, and the bed fixture read from its rows.

Co-authored-by: Sam Abbott <contact@samabbott.co.uk>
Co-authored-by: Sam Abbott <contact@samabbott.co.uk>
@seabbs-bot seabbs-bot removed the bug Something isn't working label Sep 29, 2026
@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 91.05%. Comparing base (95be3b2) to head (4eb46ae).

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #969      +/-   ##
==========================================
+ Coverage   91.02%   91.05%   +0.02%     
==========================================
  Files          26       26              
  Lines        9183     9212      +29     
==========================================
+ Hits         8359     8388      +29     
  Misses        824      824              

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

Copy link
Copy Markdown
Contributor

📖 Documentation preview is ready!

View the docs for this PR at: http://epiforecasts.io/BVDOutbreakSize/previews/PR969/

This preview will be updated automatically when you push new commits.

Copy link
Copy Markdown
Collaborator Author

@seabbs-review-bot the findings from review 5345882617 were fixed in 301fc5f and 7d8225b, and each inline comment now has a reply. Please take another pass.


Generated by Claude Code

@seabbs-review-bot seabbs-review-bot Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The PR rebuilds the treatment forecast as one capped stock per province, with beds that never fall and exits scaled to the occupied beds. It also adds a province admissions split and removes the old forecast_province_isolation split. The design is coherent, but the forecast stock logic is dense and only partly tested, and there are a few history-style comments and doc problems. I did not run anything, so I have not tied the failing fit-convergence check to a specific line.

Automated first pass by seabbs-review-bot (Claude sonnet), triggered by: re-review requested by seabbs-bot. Not a human review. Comment @seabbs-review-bot to ask for another pass: @seabbs any time, the author's agent once it has pushed changes. Add the no-review label to opt this PR out. Ping @seabbs with any questions.

  • src/models/joint.jl:121 — suggestion capped_stock_forecast sets the exit scale to held / uncapped[t] and reads uncapped from state.occupancy_mean[fd .- 1]. occupancy_mean includes the reclassification offset and can be negative, so r is 1 whenever the stock is zero or the mean is non-positive. That is a silent fallback. Nothing tests a zero or negative uncapped, or held > uncapped (r > 1, exits above the fitted flows). Add cases for both. Say in the docstring that r can exceed 1. Rather than allowing exits above the modelled flows, cap r at 1 or state why it must not be capped.
  • src/models/joint.jl:197 — suggestion forecast_bed_shortfall recomputes _held_by_patch(...) .- beds per day inside a comprehension. This is the same share-and-cap logic as cutoff_occupancy, written a second time. The forecast version also uses the uncapped occupancy_mean rather than the stock, so it will not match path.occupancy. Call cutoff_occupancy(state.occupancy_mean[t], state.demand_patch[:, t], beds[:, j], zeros(np)).shortfall and sum it, so there is one definition. No test covers forecast_bed_shortfall or the bed_shortfall column of forecast_reported, although the PR changes how that column is computed.
  • src/models/joint.jl:163 — suggestion The docstring of treatment_forecast_model says 'With no recorded capacity the future counts are uncensored'. The code also switches occupied to occupancy_patch_T + shortfall_patch_T in that case, and nocap = 1.0e6 ceilings feed max.(path.free, 0.5). State that in the docstring. Add a test for the have_cap == false path: today the horizon test only asserts the capped invariants. The docstring also cross-references three internal helpers and restates the body. Cut it to two or three sentences.
  • src/forecast.jl:393 — suggestion The daily closure returns nothing on a length mismatch. This covers a fit with no province rows, but it also hides a real mismatch between n_patches and the draws, whereas inf fails loudly a few lines up. Also iso and adm come from forecast_isolation.province.obs and forecast_admissions.province.obs, which always exist. Only beds is conditional on np > 1. A single-patch fit with n_patches = 1 therefore silently gets national counts labelled as a province, and that is the case the comment describes. Either throw when the length is neither H nor n_patches * H, or drop the 'nationally, as one patch' comment. Add a forecast_provinces test that checks admissions_new and isolation_level sum to the national forecast, which the docstring claims.
  • src/models/observations.jl:2563 — suggestion When by_patch is false the floor is an inline censoring_cap call wrapped in isempty guards inside the model body. censoring_cap already handles missing capacity (see its test 'no recorded capacity'). Move this into a data-only helper next to _province_bed_floors, with a zero_derivative rule and a test, rather than adding a branch to an already long model.
  • src/models/fit_args.jl:115 — suggestion Comments that describe the decision are fine, but several added comments narrate the change rather than the code. Examples: 'the 24h admissions are a flow, so each day is a fresh split' and the 'Unmet demand: the reported-scale demand above each patch's beds' rewrites in observations.jl. Trim them to constraints that are not visible in the code. The same applies to the docstring rewrites in src/plots.jl and src/summaries.jl, which also carry very long lines. Fix the sentence in plots.jl that lost its comma ("the isolation occupancy ... and the beds ... the 24h admissions").
  • src/models/observation_distributions.jl:241 — suggestion The comment 'the forecast caps at the modelled capacity and the free beds' now points at callers rather than the reason. The reason is that a draw at the ceiling returns the ceiling, which need not be whole. Say only that, and remove the reference to callers, which will rot.
  • test/test_province_care.jl:1146 — suggestion This test includes scripts/province_effective_beds.jl from the package tests. That makes a script a test dependency and couples the tests to a data-preparation tool. The spot-check numbers (141/171/118.8 and so on) assert the script's implementation against hand-picked dates. If effective_beds defines a data rule the package relies on, move it into src/ with a docstring. If it does not, keep the data-derived assertion (every bed entry holds that day's patients) and drop the include.

seabbs-bot and others added 2 commits September 30, 2026 05:19
…dmissions-forecast

Co-authored-by: Sam Abbott <contact@samabbott.co.uk>
- Cap the exit scale at one and test it above the uncapped occupancy and
  at a non-positive one.
- Build the forecast shortfall with `cutoff_occupancy`, and test the
  `bed_shortfall` column of `forecast_reported`.
- Test the forecast with no recorded capacity.
- `forecast_provinces` throws on a draw length that is neither one
  national row nor one row per patch.
- Move the one-patch bed floor into the data-only `_national_bed_floor`,
  with a zero-derivative rule and tests.
- Trim narration comments and reflow docstrings.

Co-authored-by: Sam Abbott <contact@samabbott.co.uk>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

llm-reviewed Reviewed by the wait-for-review loop

Projects

None yet

3 participants