Repository navigation
H031: seed </think>'s lm_head row from its first piece (experiment branch) - #1931
abhishekraok wants to merge 1 commit into
Conversation
--promoted_token_output_init_from_first_piece '</think>' copies the `</` row into `</think>`'s output row instead of the piece mean; input rows keep the mean. Both SFT paths. Raises on a tied head or a name that was not promoted. Outside TokenizerConfig, so the dataset cache key is unchanged. The olmo-core path also logs step-0 hashes per matrix (all rows but the promoted ones, each promoted row, each first-piece row), so a flag-on run can be checked against a flag-off one row by row. The KDA/hero launcher passes OUTPUT_INIT_FIRST_PIECE through to the training block only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Written by the Claude Code session Cross-family review: Codex, high effort, diff passed inline ( |
|
Written by the Claude Code session Image built from 76a0f4f: Audit 2 (dependencies), partial:
|
|
Written by the Claude Code session Audit 2 done (Beaker 01M3ZX323PRXDB4D8W7BC6V735, CPU): It comes from transformer-engine 2.16.1's unpinned resolution in the Dockerfile. onnx serializes models for ONNX export and has no role in training numerics, and this is a patch release, so I'm recording it as justified rather than rebuilding with a pin. A rebuild would add a Dockerfile change to audit 3's diff. |
Written by the Claude Code session
ledger-worker(session id72ea05cb-832c-5396-b05b-5bece635b222).Code for ledger hypothesis H031 (allenai/olmo-post-training-ledger): does seeding
</think>'s output row from</, instead of the piece mean, improve forced MATH against H029 think-emo s34521. This targets the hero-SFT experiment branchh015-think-token, not main. The image for the run is built from this commit, so it differs from H025's image (028b279) only by this diff.Changes
--promoted_token_output_init_from_first_piece <token> ...is a new flag in both SFT paths (olmo_core_finetune.SFTConfig,finetune.FlatArguments). For each named promoted token, the output (lm_head) row is copied from its first piece. Input rows still take the piece mean, and every other promoted row is unchanged.TokenizerConfig, so the cache key is unchanged andEXPECTED_NUMPY_CACHE=3f12323c3b-6068a350still asserts.Step-0 hashes, matrix iafter seeding: one hash over all rows except the promoted ones, one for each promoted row, and one for each first-piece row. This is what H031's audit 1 compares between a flag-off gate and the flag-on run.oc_sft_olmoe3_kda_think.sh:OUTPUT_INIT_FIRST_PIECE='</think>'is passed to the training block only.Audit status (H031 "Code change and audits")
git diff 028b279a 76a0f4f8touches the flag, its hash logging, tests, and two launchers.oc_sft_hero_hillclimb_1895.shis bd68b1b's change, the launcher H029 ran with. The new path is a deterministic row copy plushashlib, and draws no random numbers.1, 2, 4, 5 come after the image build. 1 is a flag-off 1x8 gate plus the flag-on run's log. 2 is a
pip freezediff against 01M38JPC. 4 is the sha256 of think-emo's savedglobal_indices_dataset_size345754_epoch1_seed34521_v1.npyagainst the new run's. 5 checks launch args against think-emo's.Tests
test_seed_embedding_rows.pyandtest_reserved_slot_tokens.py: 57 pass, 7 of them new. Run in a venv synced from this branch's lockfile; torchvision was removed locally because of a cu130/cu128 mismatch the tests don't touch.ruff formatandruff checkare clean.tyreports no new diagnostics on the changed files: 68 before and 68 after, all pre-existing.CHANGELOG=experiment branch, not main
🤖 Generated with Claude Code