CUDA matrix multiplication benchmark using cuBLASLt and cuSPARSELt (2:4 structured sparsity).
- Persistent FP16 buffers to remove host->device conversion overhead
- Persistent cuBLASLt context + 32MB workspace reuse
- cuSPARSELt 2:4 sparsity kernel path
- Benchmark harness: 3 warmup + 10 timed iterations
- CUDA Toolkit 12.0+
- cuSPARSELt 0.4+
No verified throughput numbers are published here. Run the included benchmark harness on your own hardware to measure FP16 dense and 2:4 sparse throughput; results depend on GPU model, driver version, and problem size.