Conversation
…lling, and fast horizontal reduction - Replace non-fused multiply-adds with _mm256_fmadd_ps and _mm512_fmadd_ps - 4-way ILP unrolling (32 elements/iter across 4 accumulators) in AVX2 kernels to hide 4-cycle FMA latency - Dual accumulator 32-element unrolling in AVX-512 kernels - Replace slow consecutive _mm_hadd_ps in AVX2 with fast shuffle tree (_mm_movehdup_ps / _mm_movehl_ps / _mm_add_ss) - Implement AVX-512 masked tail loads and reductions, avoiding scalar fallback loops - Fixes alibaba#706
|
Hi thlurte, Thank you for contributing this optimization. We appreciate the effort to improve the FP32 distance kernels through FMA, loop unrolling, and horizontal reduction changes. We’d like to evaluate the changes more thoroughly and benchmark the performance on AVX2 and AVX512 hardware etc. If you have benchmark results, please share the CPU model, compiler and build flags, vector dimensions, and before-and-after measurements. Those would be very helpful for the evaluation. Thanks again for the contribution and your patience while it is reviewed. |
Hi @richyreachy, Thank you for the review! Here are the empirical benchmark measurements comparing the baseline kernels against the optimized kernels in this PR. Environment & Setup
1. AVX2 BenchmarksSquared Euclidean (L2) Distance
Inner Product (IP) Distance
2. AVX-512 BenchmarksSquared Euclidean (L2) & Inner Product (IP)
Key Architectural Takeaways
Please let me know if you would like me to test any additional dimensions or target architectures! |
Fixes #706.
Summary of Changes
This PR optimizes the FP32 distance kernels (Squared Euclidean and Inner Product) in
Turbofor both AVX2 and AVX-512 target architectures:acc0,acc1,acc2,acc3). This hides the 4-cycle FMA pipeline latency and saturates CPU execution ports.FMA):_mm256_add_ps(..., _mm256_mul_ps(...))and_mm512_add_ps(..., _mm512_mul_ps(...))) with single-cycle_mm256_fmadd_psand_mm512_fmadd_ps._mm_hadd_psinstructions in AVX2 with a fast shuffle tree (_mm_movehdup_ps,_mm_movehl_ps,_mm_add_ss)._mm512_maskz_loadu_ps), eliminating scalar fallback loops.Verification
clang-format(--dry-run --Werrorpasses with 0 violations)../build/bin/turbo_fp32_quantizer_test: 7/7 tests passed across all 14 test dimensions (./build/bin/flat_turbo_index_test: 26/26 tests passed (end-to-end index build, save, reload, and query across FP32, FP16, INT8, INT4).