Skip to content

Feature/flashsac integration - #1157

Open
psurchit wants to merge 3 commits into
mujocolab:mainfrom
psurchit:feature/flashsac-integration
Open

psurchit wants to merge 3 commits into
mujocolab:mainfrom
psurchit:feature/flashsac-integration

Conversation

@psurchit

Copy link
Copy Markdown

Integrating the RSL-RL wrapper for FlashSAC (rsl-rl pull-request). Tested by running training with the example Mjlab-Velocity-Flat-Unitree-G1 environment and training progresses flawlessly.

Add RSL-RL config dataclasses for FlashSAC (RslRlFlashSacActorCfg,
RslRlFlashSacCriticCfg, RslRlReplayBufferCfg, RslRlFlashSacAlgorithmCfg) and
RslRlOffPolicyRunnerCfg, whose asdict() yields the exact key contract that
rsl_rl's FlashSAC.construct_algorithm expects (defaults are the single source).
Add MjlabOffPolicyRunner (persists common_step_counter, gates W&B upload, ONNX
export via dynamo=False) and a VelocityOffPolicyRunner that exports ONNX on
save. Register Mjlab-Velocity-Flat-Unitree-G1-FlashSAC (512 envs). Selectable
purely by config; train.py/play.py need no changes. Changelog updated.
FlashSAC's own many-env (IsaacLab) config uses 3-step returns for
locomotion; align the G1 velocity FlashSAC task with it. Verified the
task trains (reward trending up, critic loss decreasing monotonically)
on a live 512-env run.
Revert actor/critic hidden dims to 128/256 (paper's unified GPU-sim config)
instead of 512; the paper's proven recipe converges faster than an oversized
net under a fixed step budget. Verified the full algorithm matches the paper
method (inverted-residual + preact-BN + post-RMSNorm + distributional critic
+ adaptive reward scaling + weight norm; cross-batch value prediction;
noise-repetition; unified entropy target).

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant