Speed up randint array-bounds via chunking Lemire - #173
vlad-perevezentsev wants to merge 8 commits into
Conversation
| idx[wpos++] = j; | ||
| continue; | ||
| /* retry the chunk's rejects locally with fresh words */ | ||
| while (n_pending > 0) { |
There was a problem hiding this comment.
One idea is to move this loop out, and track n_pending outside of the chunk loop (i.e., before the for loop) then retry all of the rejects at the end. We'd have to set the idx allocation to len instead of chunk_cap too.
There was a problem hiding this comment.
@ndgrigorian I collected the performance of your suggestion and it looks a bit more expensive
int32, N=5M, high=1e9: 10.20 ms vs 11.12 ms.
I think the current approach is cheaper because the rejected values are still in cache while retrying at the end requires another pass and storing all rejected indices
| * Sped up `randint` for `bool`, `uint8`, `int8`, `uint16` and `int16`; generated values are unchanged [gh-172](https://github.com/IntelPython/mkl_random/pull/172) | ||
| * Raised the minimum build-time `Cython` requirement to `3.1.0`, the first release providing the `freethreading_compatible` directive [gh-159](https://github.com/IntelPython/mkl_random/pull/159) | ||
| * Extended the `memcpy`-based fast path of `shuffle` to multi-dimensional `ndarray` inputs whose first-axis items are contiguous, which is also much faster than the previous buffered path [gh-159](https://github.com/IntelPython/mkl_random/pull/159) | ||
| * Speed up `randint` with `array_like` bounds via word chunking and a branchless Lemire loop [gh-173](https://github.com/IntelPython/mkl_random/pull/173) |
There was a problem hiding this comment.
Actually something like "range-width-adaptive Lemire loop" matches the code, rather then "branchless Lemire loop"
| irk_uniform_bits_vec(state, chunk, words); | ||
|
|
||
| if (wide) { | ||
| WT last_s = 0, last_t = 0; /* memoized reject threshold */ |
There was a problem hiding this comment.
last_s/last_t reset per chunk lose cross-chunk caching
This PR implements the word-generation speedup suggested for
randintarray-like bounds in the #168 (comment)It changes
irk_rand_bounded_broadcastto:lo < sis free for narrow ranges but mispredicts for wide ones, so its hits are counted on the first chunk and the remaining chunks switch to the branchless test with the threshold memoized per range.Performance (Intel(R) Xeon(R) Platinum 8480+)
(
low=0,high=1e9):