Skip to content

H031: seed </think>'s lm_head row from its first piece (experiment branch) - #1931

Draft
abhishekraok wants to merge 1 commit into
h015-think-tokenfrom
h031-think-output-init
Draft

abhishekraok wants to merge 1 commit into
h015-think-tokenfrom
h031-think-output-init

Conversation

@abhishekraok

Copy link
Copy Markdown
Collaborator

Written by the Claude Code session ledger-worker (session id 72ea05cb-832c-5396-b05b-5bece635b222).

Code for ledger hypothesis H031 (allenai/olmo-post-training-ledger): does seeding </think>'s output row from </, instead of the piece mean, improve forced MATH against H029 think-emo s34521. This targets the hero-SFT experiment branch h015-think-token, not main. The image for the run is built from this commit, so it differs from H025's image (028b279) only by this diff.

Changes

  • --promoted_token_output_init_from_first_piece <token> ... is a new flag in both SFT paths (olmo_core_finetune.SFTConfig, finetune.FlatArguments). For each named promoted token, the output (lm_head) row is copied from its first piece. Input rows still take the piece mean, and every other promoted row is unchanged.
    • It raises on a tied head, and on a name that is not a promoted token, so a typo can't silently fall back to the mean.
    • The default is empty, which keeps today's behaviour.
    • The flag lives outside TokenizerConfig, so the cache key is unchanged and EXPECTED_NUMPY_CACHE=3f12323c3b-6068a350 still asserts.
  • The olmo-core path now logs Step-0 hashes, matrix i after seeding: one hash over all rows except the promoted ones, one for each promoted row, and one for each first-piece row. This is what H031's audit 1 compares between a flag-off gate and the flag-on run.
  • oc_sft_olmoe3_kda_think.sh: OUTPUT_INIT_FIRST_PIECE='</think>' is passed to the training block only.

Audit status (H031 "Code change and audits")

  1. Code: git diff 028b279a 76a0f4f8 touches the flag, its hash logging, tests, and two launchers. oc_sft_hero_hillclimb_1895.sh is bd68b1b's change, the launcher H029 ran with. The new path is a deterministic row copy plus hashlib, and draws no random numbers.
    1, 2, 4, 5 come after the image build. 1 is a flag-off 1x8 gate plus the flag-on run's log. 2 is a pip freeze diff against 01M38JPC. 4 is the sha256 of think-emo's saved global_indices_dataset_size345754_epoch1_seed34521_v1.npy against the new run's. 5 checks launch args against think-emo's.

Tests

  • test_seed_embedding_rows.py and test_reserved_slot_tokens.py: 57 pass, 7 of them new. Run in a venv synced from this branch's lockfile; torchvision was removed locally because of a cu130/cu128 mismatch the tests don't touch.
  • ruff format and ruff check are clean.
  • ty reports no new diagnostics on the changed files: 68 before and 68 after, all pre-existing.

CHANGELOG=experiment branch, not main

🤖 Generated with Claude Code

--promoted_token_output_init_from_first_piece '</think>' copies the `</` row into
`</think>`'s output row instead of the piece mean; input rows keep the mean. Both SFT
paths. Raises on a tied head or a name that was not promoted. Outside TokenizerConfig,
so the dataset cache key is unchanged.

The olmo-core path also logs step-0 hashes per matrix (all rows but the promoted ones,
each promoted row, each first-piece row), so a flag-on run can be checked against a
flag-off one row by row. The KDA/hero launcher passes OUTPUT_INIT_FIRST_PIECE through
to the training block only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@abhishekraok

Copy link
Copy Markdown
Collaborator Author

Written by the Claude Code session ledger-worker (session id 72ea05cb-832c-5396-b05b-5bece635b222).

Cross-family review: Codex, high effort, diff passed inline (codex exec review can't run git in its sandbox on Hammond). Result: no findings. The prompt asked specifically about DTensor sharding, the fp32 optimizer mains, whether the step-0 hashes are comparable between a 1x8 flag-off run and a 2x8 flag-on run, RNG consumers, and the cache key.

@abhishekraok

Copy link
Copy Markdown
Collaborator Author

Written by the Claude Code session ledger-worker (session id 72ea05cb-832c-5396-b05b-5bece635b222).

Image built from 76a0f4f: abhishekr/open-instruct-integration-test-h031-think-output-init-cuda13 = 01M3ZNSFRCA5BGK5HEC3XASNFS (CUDA 13).

Audit 2 (dependencies), partial:

  • uv.lock, pyproject.toml and Dockerfile are identical between 028b279 and 76a0f4f (0 diff lines).
  • That doesn't close it. Two Dockerfile steps install outside the lock: the unpinned build tools for transformer-engine (cmake ninja pybind11 setuptools wheel), and TE 2.16.1's own resolution (48 packages, e.g. onnx). This build installed cmake 4.4.3, pybind11 3.1.0 and wheel 0.48.0 fresh.
  • The pip freeze diff against 01M38JPC therefore still has to run inside both images, as a small CPU Beaker job. That's pending the launch go-ahead.

@abhishekraok

Copy link
Copy Markdown
Collaborator Author

Written by the Claude Code session ledger-worker (session id 72ea05cb-832c-5396-b05b-5bece635b222).

Audit 2 done (Beaker 01M3ZX323PRXDB4D8W7BC6V735, CPU): uv pip freeze of /stage/.venv in 01M38JPC (028b279) and 01M3ZNSFRCA5BGK5HEC3XASNFS (76a0f4f) gives 311 packages each. The only difference is onnx 1.23.0 → 1.23.1.

It comes from transformer-engine 2.16.1's unpinned resolution in the Dockerfile. onnx serializes models for ONNX export and has no role in training numerics, and this is a patch release, so I'm recording it as justified rather than rebuilding with a pin. A rebuild would add a Dockerfile change to audit 3's diff.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant