Skip to content

Decode throughput: move lm_head to ANE for all LlamaPrefill models #227

Description

@axiom-of-choice

#215 / #224 folded GPT-2 into LlamaPrefill with host-side lm_head and wpe. speculative.py already has _build_lm_head that runs the head on the ANE as a tiled compile_multi, amortising cost over Kq tokens. That pattern should be generalised to LlamaPrefill.generate so every model (Llama/Qwen/GPT-2/Mistral) benefits, not just the speculative verifier.
Scope:

  • Add _build_decode_lm_head to LlamaPrefill, using limit("max_tensor_dim", detect_family()) so A16+ models with vocab < 65536 get a single dispatch.
  • Replace _logits host-matmul in generate with ANE lm_head when the model opts in (or by default if numerics stay exact).
  • Ensure release() frees the lm_head program.
    Blocked on GPT-2 decode throughput: family-aware lm_head tiling + fold wpe into the graph #226 completing first, so we have the family-aware tile size and a measured GPT-2 baseline to compare against.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions