Implement number theoretic transform for large integer multiplication - #282
Implement number theoretic transform for large integer multiplication#282byeongkeunahn wants to merge 1 commit into
Conversation
|
I want to first acknowledge this with thanks -- it's quite impressive to be near GMP performance! However to set expectations: this is also a large PR, and I will need some studying to understand what's going on, so it may take me a while to find time to review this. |
eae5b13 to
66faa24
Compare
|
Thanks. I'd like to note a few changes in the latest commit, which brings 10-15% performance gains and reduced memory footprint:
|
2639746 to
0e41192
Compare
|
Does the chart shows that current algorithm is faster than GMP? That's impressive. |
|
I ran benchmark fib_hex 100m from https://github.com/tczajka/bigint-benchmark-rs on this PR and it made num-bigint twice faster than malachite, slightly faster than gmp and 12x faster than itself. |
|
Obviously the performance here is impressive, as I said before. However, as I finally started to try reviewing this, I have a few high level objections.
(that particular one does pass under tree-borrows though) |
|
Thanks for your review and comments! I’ll try to address the issues, including removing the unsafe code and cleaning up the commit history, even if that comes at the cost of a small performance regression. Integrating the Montgomery reduction code with
|
|
We can treat the monty-consolidation as "nice to have". However, any specific tailoring needs comments, lest that work be undone by a later contributor or maintainer. I'm skeptical that those optimizations wouldn't be useful to the other monty use as well though... |
|
I’ve made the suggested improvements:
I also made a small improvement to the planner based on exhaustive Criterion.rs benchmarks covering operand sizes of up to 300 million bits each. This reduced execution time by 1.2% on average. Some of this work was completed with assistance from ChatGPT Codex. Thanks again for the detailed review. Please let me know if there are any remaining issues. |
This commit implements number theoretic transform (NTT) for large integer multiplication (issue #169).
On Ryzen 7 2700X, 64bit, it takes about 15ms for 2.7Mbits x 2.7Mbits and 170ms for 27Mbits x 27Mbits multiplication. This seems comparable to GMP 6.2.1.