perf: vectorize small u8 table take with AVX2 - #9572
Conversation
Merging this PR will regress 1 benchmark
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | WallTime | filtered_owned_i64_avx2[OneNullInEight] |
21.9 µs | 26.8 µs | -18.4% |
| ⚡ | WallTime | dict_canonicalize_gt_u8_avx2[1000000] |
419.8 µs | 58.2 µs | ×7.2 |
| ⚡ | WallTime | dict_canonicalize_gt_u8_avx512[1000000] |
419.5 µs | 58.8 µs | ×7.1 |
| ⚡ | WallTime | dict_canonicalize_gt_u8_avx2[16000000] |
6.8 ms | 1.4 ms | ×5 |
| ⚡ | WallTime | dict_canonicalize_gt_u8_avx512[16000000] |
6.8 ms | 1.4 ms | ×5 |
| ⚡ | Simulation | decode_primitives[u8, (2000, 8)] |
43.8 µs | 33.4 µs | +30.88% |
| ⚡ | Simulation | decode_primitives[u8, (2000, 4)] |
43.8 µs | 33.4 µs | +30.88% |
| ⚡ | Simulation | decode_primitives[u8, (2000, 2)] |
43.8 µs | 33.4 µs | +30.88% |
| ⚡ | Simulation | decode_primitives[u8, (1000, 8)] |
36.3 µs | 31.5 µs | +15.07% |
| ⚡ | Simulation | decode_primitives[u8, (1000, 4)] |
36.3 µs | 31.5 µs | +15.06% |
| ⚡ | Simulation | decode_primitives[u8, (1000, 2)] |
37.1 µs | 32.3 µs | +14.69% |
| ⚡ | WallTime | words_gather_scalar_avx2[65536] |
9.4 µs | 8.3 µs | +13.35% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing ji/small-u8-table-take-avx2 (134d1c5) with ji/small-u8-table-take-neon (e97db31)
Footnotes
-
218 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
|
This PR has been marked as stale because it has been open for 14 days with no activity. Please comment or remove the stale label if you wish to keep it active, otherwise it will be closed in 7 days |
Use AVX2 VPSHUFB for u8-coded tables with at most 16 one-byte values. Signed-off-by: Joseph Isaacs <joseph-isaacs@users.noreply.github.com>
0129184 to
134d1c5
Compare
Rationale
Add the x86 table-lookup implementation after the NEON implementation in #9571.
This is the upper PR in the stack: #9566 (benchmark) → #9571 (NEON) → this PR (AVX2).
Changes
VPSHUFBon x86/x86-64.u8codes, at most 16 one-byte values, and at least 64 rows.CodSpeed wall-time results
Medians over 1,000 samples on the same Sapphire Rapids metal family, compared with #9566:
The AVX-512 leg is an AVX-512-enabled whole-crate build executing this AVX2
VPSHUFBkernel; there is no separate AVX-512 table kernel.Runs: baseline, optimized.
Validation
vortex-arraytests passed; 1 skipped