Skip to content

Latest commit

 

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

Matrix Beast v2.1

CUDA matrix multiplication benchmark using cuBLASLt and cuSPARSELt (2:4 structured sparsity).

Changelog

  • Persistent FP16 buffers to remove host->device conversion overhead
  • Persistent cuBLASLt context + 32MB workspace reuse
  • cuSPARSELt 2:4 sparsity kernel path
  • Benchmark harness: 3 warmup + 10 timed iterations

Requirements

  • CUDA Toolkit 12.0+
  • cuSPARSELt 0.4+

Benchmarking

No verified throughput numbers are published here. Run the included benchmark harness on your own hardware to measure FP16 dense and 2:4 sparse throughput; results depend on GPU model, driver version, and problem size.

About

Honest, working, scientific matrix multiply library: C++20 + C11 + hand-written AVX2/FMA asm + OpenMP, verified against a scalar reference.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages