Tech & AI News
Hacker News

Vectorized and performance-portable Quicksort

The open-source implementation uses SIMD “compress-store” and permute instructions to accelerate Quicksort, achieving up to ten times the speed of C++ std::sort. On an Apple M1 it sorts 32-, 64-, and 128-bit numbers at 499, 471, and 466 MB/s, while on a 3 GHz Skylake with AVX-512 it reaches 1123 MB/s, outperforming prior architecture-specific algorithms and the standard library.