Skip to content

Update normalizers after optimization - #225

Merged
ClemensSchwarke merged 3 commits into
leggedrobotics:mainfrom
NeoZng:fix/rollout-normalization-update
Aug 26, 2026
Merged

ClemensSchwarke merged 3 commits into
leggedrobotics:mainfrom
NeoZng:fix/rollout-normalization-update

Conversation

@NeoZng

@NeoZng NeoZng commented Aug 20, 2026

Copy link
Copy Markdown
Contributor

Summary

  • keep observation-normalization statistics fixed throughout PPO and distillation optimization
  • update actor, critic, student, and RND normalizers once from the flattened rollout observations after optimization
  • use the observations that were actually stored and optimized instead of the next observations passed to process_env_step()

Problem

process_env_step() currently updates observation normalizers after every environment step. Actions and their old log probabilities are computed before that update, while PPO.update() recomputes log probabilities after the normalizer has changed. Consequently, the first policy ratio can differ from one even though no policy parameter update has occurred yet.

The per-step update also consumes the post-step observation (s_{t+1}), whereas the optimization batches contain the pre-step observations (s_t) stored in the rollout.

Implementation

  • remove normalizer updates from PPO and distillation process_env_step()
  • flatten the rollout time and environment dimensions and update normalizers after all optimization epochs
  • perform the update before clearing storage so the next rollout uses the newly accumulated statistics
  • apply the same timing to PPO actor/critic/RND normalizers and the distillation student normalizer
  • add the contributor entry required by CONTRIBUTING.md

Tests

  • pytest tests/algorithms/test_ppo.py tests/algorithms/test_distillation.py -q — 20 passed
  • pre-commit run --all-files — all hooks passed

The new PPO regression test verifies that the first recomputed ratio remains one before optimization and that both normalizers are updated exactly once from stored rollout observations. The distillation regression test verifies the corresponding student-normalizer timing and data source.

@NeoZng
NeoZng marked this pull request as ready for review August 20, 2026 12:55
@ClemensSchwarke ClemensSchwarke added the enhancement New feature or request label Aug 26, 2026
@ClemensSchwarke
ClemensSchwarke merged commit 9965512 into leggedrobotics:main Aug 26, 2026
7 checks passed
colson-louis added a commit to MontefAiDefense/rsl_rl that referenced this pull request Sep 25, 2026
Restores the tree of c281c32 ("Fix term order in console logs") on top of
857de61, undoing the upstream commits pulled in by the 2026-09-24 sync. leggedrobotics#225
("Update normalizers after optimization") moves the observation-normalizer
update to after each PPO update, which slows learning on the defnder tracking
task (hold_0.1mrad 0.39 vs 0.66 at 300 iterations even with a normalizer
warm-up). This commit changes files only; history is kept.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011tgsHmmKs3o2w8L2qzi5WZ
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants