You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
#215 / #224 folded GPT-2 into LlamaPrefill with host-side lm_head and wpe. speculative.py already has _build_lm_head that runs the head on the ANE as a tiled compile_multi, amortising cost over Kq tokens. That pattern should be generalised to LlamaPrefill.generate so every model (Llama/Qwen/GPT-2/Mistral) benefits, not just the speculative verifier.
Scope:
Add _build_decode_lm_head to LlamaPrefill, using limit("max_tensor_dim", detect_family()) so A16+ models with vocab < 65536 get a single dispatch.
Replace _logits host-matmul in generate with ANE lm_head when the model opts in (or by default if numerics stay exact).
#215 / #224 folded GPT-2 into LlamaPrefill with host-side lm_head and wpe. speculative.py already has _build_lm_head that runs the head on the ANE as a tiled compile_multi, amortising cost over Kq tokens. That pattern should be generalised to LlamaPrefill.generate so every model (Llama/Qwen/GPT-2/Mistral) benefits, not just the speculative verifier.
Scope:
Blocked on GPT-2 decode throughput: family-aware lm_head tiling + fold wpe into the graph #226 completing first, so we have the family-aware tile size and a measured GPT-2 baseline to compare against.