Update normalizers after optimization - #225
Merged
ClemensSchwarke merged 3 commits intoAug 26, 2026
Merged
ClemensSchwarke merged 3 commits into
ClemensSchwarke merged 3 commits into
Conversation
NeoZng
marked this pull request as ready for review
August 20, 2026 12:55
ClemensSchwarke
approved these changes
Aug 26, 2026
colson-louis
added a commit
to MontefAiDefense/rsl_rl
that referenced
this pull request
Sep 25, 2026
Restores the tree of c281c32 ("Fix term order in console logs") on top of 857de61, undoing the upstream commits pulled in by the 2026-09-24 sync. leggedrobotics#225 ("Update normalizers after optimization") moves the observation-normalizer update to after each PPO update, which slows learning on the defnder tracking task (hold_0.1mrad 0.39 vs 0.66 at 300 iterations even with a normalizer warm-up). This commit changes files only; history is kept. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011tgsHmmKs3o2w8L2qzi5WZ
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
process_env_step()Problem
process_env_step()currently updates observation normalizers after every environment step. Actions and their old log probabilities are computed before that update, whilePPO.update()recomputes log probabilities after the normalizer has changed. Consequently, the first policy ratio can differ from one even though no policy parameter update has occurred yet.The per-step update also consumes the post-step observation (
s_{t+1}), whereas the optimization batches contain the pre-step observations (s_t) stored in the rollout.Implementation
process_env_step()CONTRIBUTING.mdTests
pytest tests/algorithms/test_ppo.py tests/algorithms/test_distillation.py -q— 20 passedpre-commit run --all-files— all hooks passedThe new PPO regression test verifies that the first recomputed ratio remains one before optimization and that both normalizers are updated exactly once from stored rollout observations. The distillation regression test verifies the corresponding student-normalizer timing and data source.