Skip to content

fix: make Casper training path actually learn - #5

Merged
Grar00t merged 7 commits into
mainfrom
fix/full-training-path-main-20260907
Sep 7, 2026
Merged

Grar00t merged 7 commits into
mainfrom
fix/full-training-path-main-20260907

Conversation

@Grar00t

@Grar00t Grar00t commented Sep 7, 2026

Copy link
Copy Markdown
Owner

Clean rebase onto current main 7da5605.

Repairs the C training path that previously combined zero-initialized weights with output-head-only updates and an invalid token/context check.

Changes:

  • deterministic non-zero initialization for fresh training
  • full-parameter truncated-BPTT updates over embeddings, attention/FFN projections, RMSNorm scales, and LM head
  • live tokenizer-derived vocabulary instead of hardcoded 8192
  • removes broken token-id vs context-length behavior from active trainer
  • smoke regression requires loss reduction and a real attention-backbone weight change
  • GCC/Clang release + sanitizer smoke coverage
  • preserves current Casper CLI self-check added on main
  • explicitly documents detached-KV gradient boundary; not claimed as exact full-sequence BPTT

Supersedes stale PR #4.

Copilot AI lite review requested due to automatic review settings September 7, 2026 09:37
@vercel

vercel Bot commented Sep 7, 2026

Copy link
Copy Markdown

Deployment failed for project casper with the following error:

Resource is limited - try again in 24 hours (more than 100, code: "api-deployments-free-per-day").

Learn More: https://vercel.com/gratechs-projects?upgradeToPro=build-rate-limit

@Grar00t
Grar00t merged commit 195aca0 into main Sep 7, 2026
1 of 4 checks passed

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

It introduces a large new core training implementation and optimizer/update behavior that needs careful human validation for correctness and performance impact.

Pull request overview

This PR repairs and expands the NIYAH C training path so it can learn end-to-end (beyond LM-head-only updates), and wires the new full-parameter training implementation into the trainer binary, self-checks, build script, and CI.

Changes:

  • Adds deterministic non-zero initialization plus a new full-parameter, truncated-BPTT training step implementation (detached-KV boundary).
  • Updates the trainer CLI to use live tokenizer-derived vocabulary size, run full-parameter training steps, and emit training metadata.
  • Extends the NIYAH self-check and CI smoke to include a regression that requires loss reduction and a backbone (attention) weight change, plus adds a debug sanitizer smoke job.
File summaries
File Description
scripts/build.sh Links the new full trainer implementation into niyah and trainer, and includes it in the lint gate.
README.md Updates documentation to reflect the new full-parameter training path and detached-KV training description.
Core_CPP/niyah_train.c Switches trainer to tokenizer-derived vocab size, deterministic init, and full-parameter training step.
Core_CPP/niyah_train_full.h Introduces the public API for deterministic init + full-parameter training step.
Core_CPP/niyah_train_full.c Implements deterministic initialization and full-parameter detached-KV training, including optimizer update.
Core_CPP/niyah_main.c Adds a self-check regression that enforces loss reduction and attention-backbone weight updates.
.github/workflows/ci.yml Adds a debug sanitizer smoke build/run alongside existing release smoke coverage.
Review details
  • Files reviewed: 7/7 changed files
  • Comments generated: 2
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +282 to +286
ALLOC_FLOATS(grad, nw);
ALLOC_FLOATS(x_in, (size_t)m->cfg.n_layers * d);
ALLOC_FLOATS(att_norm, (size_t)m->cfg.n_layers * d);
ALLOC_FLOATS(q, (size_t)m->cfg.n_layers * d);
ALLOC_FLOATS(k, (size_t)m->cfg.n_layers * kd);
Comment on lines +20 to +22
* Updates token embeddings, every attention/FFN projection, RMSNorm scales,
* and the LM head with AdamW. The causal KV cache is a deliberate truncated
* backpropagation boundary: a token receives gradients through its own Q/K/V
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants