Skip to content

No final restart written when a run ends between scheduled restart times #331

Description

@ktwu01

Hi, and thanks for maintaining HRLDAS.

While setting up a multi-stage spin-up I ran into a case where a run finishes normally but no restart file is written for the final state, because the run ended between two scheduled restart times. It took a while to track down, so I wanted to write it up in case it is worth changing.

What happens

Restart writing is keyed on the timestep counter, and there is no unconditional write after the loop finishes:

hrldas/IO_code/module_NoahMP_hrldas_driver.F#L1393-L1397

if (NoahmpIO%restart_frequency_hours > 0) then
   if (mod(ITIME, int(NoahmpIO%restart_frequency_hours*3600./nint(NoahmpIO%dtbl))) == 0) then
      call lsm_restart()
   endif
endif

The main loop is do ITIME = 0, NTIME (hrldas/IO_code/main_hrldas_driver.F#L17-L19). So when NTIME is not an exact multiple of the restart interval expressed in timesteps, the last scheduled restart is the final one written, and the model state at the end of the run is not saved anywhere.

Reproducer

With KHOUR = 61368, NOAH_TIMESTEP = 3600, RESTART_FREQUENCY_HOURS = 4320, SPINUP_LOOPS = 0 (SPINUP_LOOPS matters here because NTIME is multiplied by spinup_loops + 1 at module_NoahMP_hrldas_driver.F#L758):

NTIME             = 61368 timesteps
restart interval  = 4320 timesteps (4320 h at dtbl = 3600 s)
61368 = 4320*14 + 888   ->  remainder 888, so no restart at ITIME = NTIME
last restart written at ITIME 60480
final 888 h (37 days) of the run is not captured in any restart file

KHOUR = 60480 or 64800 both write a final restart, so whether the end state is saved depends on divisibility.

Why it is easy to miss

The run itself completes normally, and the log does show Write restart at ... for each restart that is written and Found restart file: ... on the next startup. What is missing is any specific indication that the final checkpoint was skipped. For a staged spin-up where the next stage is meant to continue from the end state, the second stage instead picks up from the last scheduled restart, which is up to one restart interval earlier than intended, and the run proceeds without complaint.

To be fair to the current design, <LATEST> does not fall back silently in the general case: find_restart_file searches backward hourly from olddate + KHOUR + 24 h and stops with an error if it reaches startdate without finding anything. The quiet case is specifically the one where an earlier scheduled restart does exist and is used in place of the intended final one.

Possible improvements

Either of these would have saved the debugging:

  1. Write a restart unconditionally on the final timestep (ITIME == NTIME), so the end of a run is always recoverable.
  2. Or, if periodic-only checkpointing is intended, warn at init when mod(NTIME, restart_step) /= 0, naming the last timestep that will actually be written.

I appreciate that option 2 may be the better fit if RESTART_FREQUENCY_HOURS is meant strictly as a periodic checkpoint setting with no final-state guarantee. Happy to open a PR for whichever you prefer.

Environment: offline Noah-MP single-point run, HRLDAS at commit 79f783aa2fe58bccd1c727105cecfaeaa6dff412, submodule noahmp at b0e8fa4, gfortran on macOS. Findings above are from source inspection plus the restart files produced by the run.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions