Skip to content

[ilu/ttx] optimize int8-KV paged prefill: dequant + reuse bf16 FA2 - #473

Draft
mojo-opset-sync[bot] wants to merge 1 commit into
masterfrom
ilu/optimize_quant_fla_attn
Draft

mojo-opset-sync[bot] wants to merge 1 commit into
masterfrom
ilu/optimize_quant_fla_attn

Conversation

@mojo-opset-sync

Copy link
Copy Markdown

Rebase of #361 onto the rebuilt master (df2f3dab).

Original PR author: AbeFei. Original commit authors, author/committer dates, and messages are preserved.

Migration status: rebased and conflicts resolved; Python syntax checked. Accelerator correctness and performance still require validation before merge.

- Add _dequant_paged_kv_block_kernel and rewrite the int8 dequant prefill
  to dequantize referenced KV blocks to bf16, then reuse the bf16 FA2
  kernel; fall back to the scalar kernel for unsupported shapes.
  ~100x faster (1199ms -> 11.6ms).
- Pass original (Hkv, D) scales (drop repeat_interleave); cache bf16
  dequant buffers.
- Make the GQA packed kernel's autotune-disabled config shared-mem-safe
  for large groups to fix OOR in CI; fix double-scaling in the non-packed
  partial-page loop.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant