Hi, and thanks for maintaining HRLDAS.
While setting up a multi-stage spin-up I ran into a case where a run finishes normally but no restart file is written for the final state, because the run ended between two scheduled restart times. It took a while to track down, so I wanted to write it up in case it is worth changing.
What happens
Restart writing is keyed on the timestep counter, and there is no unconditional write after the loop finishes:
hrldas/IO_code/module_NoahMP_hrldas_driver.F#L1393-L1397
if (NoahmpIO%restart_frequency_hours > 0) then
if (mod(ITIME, int(NoahmpIO%restart_frequency_hours*3600./nint(NoahmpIO%dtbl))) == 0) then
call lsm_restart()
endif
endif
The main loop is do ITIME = 0, NTIME (hrldas/IO_code/main_hrldas_driver.F#L17-L19). So when NTIME is not an exact multiple of the restart interval expressed in timesteps, the last scheduled restart is the final one written, and the model state at the end of the run is not saved anywhere.
Reproducer
With KHOUR = 61368, NOAH_TIMESTEP = 3600, RESTART_FREQUENCY_HOURS = 4320, SPINUP_LOOPS = 0 (SPINUP_LOOPS matters here because NTIME is multiplied by spinup_loops + 1 at module_NoahMP_hrldas_driver.F#L758):
NTIME = 61368 timesteps
restart interval = 4320 timesteps (4320 h at dtbl = 3600 s)
61368 = 4320*14 + 888 -> remainder 888, so no restart at ITIME = NTIME
last restart written at ITIME 60480
final 888 h (37 days) of the run is not captured in any restart file
KHOUR = 60480 or 64800 both write a final restart, so whether the end state is saved depends on divisibility.
Why it is easy to miss
The run itself completes normally, and the log does show Write restart at ... for each restart that is written and Found restart file: ... on the next startup. What is missing is any specific indication that the final checkpoint was skipped. For a staged spin-up where the next stage is meant to continue from the end state, the second stage instead picks up from the last scheduled restart, which is up to one restart interval earlier than intended, and the run proceeds without complaint.
To be fair to the current design, <LATEST> does not fall back silently in the general case: find_restart_file searches backward hourly from olddate + KHOUR + 24 h and stops with an error if it reaches startdate without finding anything. The quiet case is specifically the one where an earlier scheduled restart does exist and is used in place of the intended final one.
Possible improvements
Either of these would have saved the debugging:
- Write a restart unconditionally on the final timestep (
ITIME == NTIME), so the end of a run is always recoverable.
- Or, if periodic-only checkpointing is intended, warn at init when
mod(NTIME, restart_step) /= 0, naming the last timestep that will actually be written.
I appreciate that option 2 may be the better fit if RESTART_FREQUENCY_HOURS is meant strictly as a periodic checkpoint setting with no final-state guarantee. Happy to open a PR for whichever you prefer.
Environment: offline Noah-MP single-point run, HRLDAS at commit 79f783aa2fe58bccd1c727105cecfaeaa6dff412, submodule noahmp at b0e8fa4, gfortran on macOS. Findings above are from source inspection plus the restart files produced by the run.
Hi, and thanks for maintaining HRLDAS.
While setting up a multi-stage spin-up I ran into a case where a run finishes normally but no restart file is written for the final state, because the run ended between two scheduled restart times. It took a while to track down, so I wanted to write it up in case it is worth changing.
What happens
Restart writing is keyed on the timestep counter, and there is no unconditional write after the loop finishes:
hrldas/IO_code/module_NoahMP_hrldas_driver.F#L1393-L1397The main loop is
do ITIME = 0, NTIME(hrldas/IO_code/main_hrldas_driver.F#L17-L19). So whenNTIMEis not an exact multiple of the restart interval expressed in timesteps, the last scheduled restart is the final one written, and the model state at the end of the run is not saved anywhere.Reproducer
With
KHOUR = 61368,NOAH_TIMESTEP = 3600,RESTART_FREQUENCY_HOURS = 4320,SPINUP_LOOPS = 0(SPINUP_LOOPSmatters here becauseNTIMEis multiplied byspinup_loops + 1atmodule_NoahMP_hrldas_driver.F#L758):KHOUR = 60480or64800both write a final restart, so whether the end state is saved depends on divisibility.Why it is easy to miss
The run itself completes normally, and the log does show
Write restart at ...for each restart that is written andFound restart file: ...on the next startup. What is missing is any specific indication that the final checkpoint was skipped. For a staged spin-up where the next stage is meant to continue from the end state, the second stage instead picks up from the last scheduled restart, which is up to one restart interval earlier than intended, and the run proceeds without complaint.To be fair to the current design,
<LATEST>does not fall back silently in the general case:find_restart_filesearches backward hourly fromolddate + KHOUR + 24 hand stops with an error if it reachesstartdatewithout finding anything. The quiet case is specifically the one where an earlier scheduled restart does exist and is used in place of the intended final one.Possible improvements
Either of these would have saved the debugging:
ITIME == NTIME), so the end of a run is always recoverable.mod(NTIME, restart_step) /= 0, naming the last timestep that will actually be written.I appreciate that option 2 may be the better fit if
RESTART_FREQUENCY_HOURSis meant strictly as a periodic checkpoint setting with no final-state guarantee. Happy to open a PR for whichever you prefer.Environment: offline Noah-MP single-point run, HRLDAS at commit
79f783aa2fe58bccd1c727105cecfaeaa6dff412, submodulenoahmpatb0e8fa4, gfortran on macOS. Findings above are from source inspection plus the restart files produced by the run.