-
Notifications
You must be signed in to change notification settings - Fork 73
Pull requests: meta-pytorch/MSLK
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[ROCm] Add FlyDSL paged-attention decode backend (dense + fp8) with benchmark
cla signed
module: rocm
#463
opened Jul 29, 2026 by
avbokovoy
Collaborator
Loading…
3 tasks done
Add FlyDSL grouped groupwise FP8 GEMM for ROCm
cla signed
#456
opened Jul 27, 2026 by
aryaman-gupta
Contributor
Loading…
feat: add FlyDSL batched preshuffle GEMM for FP8 rowwise scaling (WP-G2)
cla signed
#444
opened Jul 22, 2026 by
kudomcho
Collaborator
Loading…
11 tasks done
feat: add FlyDSL preshuffle GEMM for FP8 rowwise scaling (gfx950)
cla signed
#434
opened Jul 10, 2026 by
kudomcho
Collaborator
Loading…
5 tasks done
Add Triton BF16xINT4 rowwise GEMM for ROCm gfx942/gfx950
cla signed
module: rocm
#420
opened Jul 7, 2026 by
apicciau
Contributor
Loading…
Fix flash_attn_bench crash on the flash_attn.cute (FA4) backend: pass window_size=(-1, -1) instead of None
#397
opened Jun 19, 2026 by
Maurits-de-Groot
Loading…
Add Autograd Backward Support for FP4 GEMM Custom Ops
cla signed
#393
opened Jun 17, 2026 by
colalb1
Loading…
Switch to upstream Cutlass dependency
cla signed
fb-exported
meta-exported
#342
opened May 12, 2026 by
jwfromm
Contributor
Loading…
[CUDA] [PERFORMANCE] Increase speed of bf16bf16bf16_grouped_wgrad via indicating that ElementC is void / nullptr
cla signed
#329
opened Apr 19, 2026 by
benediktjohannes
Loading…
ProTip!
Follow long discussions with comments:>50.