-- pg_turbovec v1.22.1 -- closes the IVF build-cliff gap: the -- Lloyd-loop k-means-assignment GEMM now runs Parallelism::Rayon(0) -- instead of Parallelism::None. Scan/build-path only, no wire -- format change, no SQL surface change. -- -- This file is a *reference mirror*. The authoritative install -- script is generated by `cargo pgrx schema`. -- -- Background: v1.20.0/v1.21.0's parallel-build work row-blocked the -- normalize/rotate/assign-sweep stages of ambuild, measuring only a -- modest ~1.27x speedup on a 64-core box (see v1.20.0's "Honest -- notes"). A FLOPs analysis (prompted by a user question about a -- benchmark driver's silent-GUC bug) found the real dominant cost -- was NEVER row-blocked: gemm_lloyd_assign's cross-term GEMM runs -- over the WHOLE k-means training sample (n_sample = lists * 256, -- which equals the full corpus size at high `lists`) once per Lloyd -- iteration, up to 25 times, single-threaded (Parallelism::None, -- for determinism). At GIST-1M/960d/lists=4096 scale this GEMM is -- ~26-112x more FLOPs than the already-parallelized stages. -- -- The fix: gemm 0.18's OWN internal Parallelism::Rayon(n) tiling -- produces BIT-IDENTICAL output to Parallelism::None for every -- shape/seed/thread-count tested (verified locally and on a cloud VM, -- including the real GIST-1M-scale shape) -- because a GEMM's -- output tiles are independent dot-product reductions over the -- shared contraction dimension; unlike a cross-thread SUM (which -- does need fixed-partition-order bookkeeping, as k-means' -- centroid-update step already has), thread count can never -- perturb a GEMM's per-element result. No hand-rolled row-blocking -- needed -- one line changed (Parallelism::None -> Rayon(0)), -- which via rayon::current_num_threads() automatically and -- correctly respects turbovec.build_parallelism's bounded pool -- (train_kmeans already runs inside build_pool::install(..)). -- -- MEASURED on real hardware, real scale (not assumed): 16-core -- AVX-512 a cloud VM instance, GIST-1M corpus shape (n_sample=1,048,576, -- dim=960, lists=4096, the full 25-Lloyd-iteration k-means -- training): TODAY (Parallelism::None) = 2686.6s (~44.8 min); -- FIXED (Parallelism::Rayon(0)) = 768.4s (~12.8 min). 3.50x -- speedup, bit-identical output confirmed. This closes a real -- fraction of the build-time gap the v1.20.0/v1.21.0 fix left open; -- see CHANGELOG.md for the full investigation writeup (including -- two dead-end/retracted findings along the way -- a test-harness -- stride bug mistaken for a gemm-crate bug, and a test-harness -- thread-pool-scoping bug that made an earlier "serial baseline" -- measurement badly overstate the problem). -- -- No wire-format change (MetaPageData::version stays 5). No SQL -- surface change (no new GUC, no reloption change). No REINDEX -- -- existing IVF indexes are unaffected (this changes build wall -- clock only, not the on-disk bytes -- centroids/assignment/ -- everything downstream is byte-identical to before). -- `ALTER EXTENSION pg_turbovec UPDATE TO '1.22.1';` is sufficient. -- -- (intentionally empty beyond this note -- there is no DDL to run.)