# Changelog All notable changes to `pg_turbovec` are documented in this file. The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/) and the project adheres to [Semantic Versioning](https://semver.org/). ## [2.11.0] — 2026-10-06 MINOR: **adopt upstream turbovec 1.1.1 (staged 2/4-bit search)**, plus fixes for **two long-standing bugs found by this release's soak test: one silently corrupted index entries under concurrent writes, the other leaked memory per scan** (see Fixed; both are in v2.10.3 and earlier). No SQL surface change, no GUC change, **no wire-format change** (`MetaPageData::version` stays **8**), index bytes unchanged. `ALTER EXTENSION pg_turbovec UPDATE` plus a restart is sufficient; **no REINDEX**. Minor rather than patch for one reason: on hosts where the new search engages, the **candidate set** an index scan returns can differ slightly from 2.10.3's (scores are exact; recall@10 measured unchanged). Details below. ### Changed - **turbovec 1.0.0 → 1.1.1.** For an index of **≥ 32,768 rows** on **aarch64 (dotprod: Graviton2+, Ampere, Apple)** or **x86 with AVX-512 VBMI+VNNI (Ice Lake server+, Zen 4+)**, turbovec now keeps its in-memory search cache as separate bit planes and searches in stages: a sign-plane pass for a shortlist, ranking on the lower planes, then an exact rescore. On AVX2-only x86 (and below 32,768 rows) the scan is unchanged. Measured on **Graviton4** (c8gd.8xlarge, Debian 13, PG 16.15 non-assert), **1M × 1024-d real Cohere embeddings**, flat index, warm, whole-query `EXPLAIN ANALYZE` p50, 3 alternated A/B rounds on one index: | | search_k=32 | 100 | 256 | 1024 | |---|---|---|---|---| | 4-bit: 2.10.3 → **2.11.0** | 7.88 → **6.72 ms (1.17×)** | 10.88 → **9.52 (1.14×)** | 18.32 → **15.96 (1.15×)** | 62.17 → **54.57 (1.14×)** | | 2-bit: 2.10.3 → **2.11.0** | 7.22 → **6.86 (1.05×)** | 10.45 → **9.78 (1.07×)** | 17.81 → **16.51 (1.08×)** | 61.67 → **55.11 (1.12×)** | Recall@10 unchanged in every cell (1.000, or 0.993 at 2-bit k=32 for both). p95 improves by the same margin. **The turbovec kernel itself is 3–6.5× faster** on this host (pure kernel, same corpus: 4-bit single-query k=10 1.84 → 0.62 ms at 32 threads, 22.7 → 6.6 ms at one thread; k=1024 8.29 → 1.28 ms). End-to-end gains are smaller because the kernel is now a minor share of a flat-index query: the saving in milliseconds is the same, but the rest — the per-candidate heap fetch and exact recheck, ~48 µs per candidate — is unchanged. For flat indexes the recheck, not the scan, is now the bottleneck. - **Cold-backend latency kept (and slightly improved) via a new fork carry.** Stock 1.1.1 builds the planes cache with a serial `planes_repack`, bypassing the parallel cold-open repack that gave v2.10.3 its 3.1× cold-scan cut — it takes 279 ms (4-bit) / 1,092 ms (2-bit) at 1M × 1024-d. **Fork carry #4** parallelizes it (18–21 ms / 32 ms), byte-identical to the serial body (`parallel_planes_repack_is_byte_identical_to_serial`, x86_64 + aarch64). Cold-backend p50 on Graviton4, 1M × 1024-d: 4-bit **513.7 → 476.9 ms**, 2-bit 327.0 → 319.8 ms. ### Fixed - **Silent index-entry corruption under sustained concurrent writes + VACUUM (present since v1.29.1).** The deferred insert flush tracks which rows a transaction upserted (`touched_ids`) and splices only those onto the current on-disk index. That list was never cleared after a successful flush, and the backend's cached index outlives the transaction, so a long-lived writer connection re-spliced **every row it had ever written** on every later commit, taking their codes from its stale in-memory copy. Once VACUUM had removed such a row and PostgreSQL reused its heap slot for a different row, that re-splice either **overwrote the new row's index entry with the old row's codes** (the row is then searched by the wrong vector, so it can be missed) or **put the vacuumed entry back** (a stale pointer to a dead or reused tuple). `turbovec_check` stayed clean throughout (ids remain unique) — only a byte-level comparison against a fresh `REINDEX` shows it. Measured: after 40 min of 3 writers + VACUUM every 2 min on Graviton4, v2.10.3 had **530 live rows with another row's codes and 2,369 stale entries** in a ~660k-row index (v2.11.0 before the fix: 721 / 2,564). Fixed: `touched_ids` is emptied once flushed. Repro test `flush_does_not_resplice_previous_txn_ids` (fails before: a vacuumed id is resurrected; passes after). **Re-run A/B, same workload side by side: v2.10.3 corrupted 406 entries and left 2,073 stale; v2.11.0 matched a fresh `CREATE INDEX` of its heap byte-for-byte on all 658,221 entries** (40 min, 12,387 commits, 40 mid-flush kills, 19 VACUUMs). Details: `benches/results/tv111_arm_20261005/FINDINGS.md` §6. **Affected:** flat/IVF TurboQuant (2/3/4-bit) indexes that take writes from connections living across many transactions while VACUUM runs. Short-lived connections (one transaction per connection) were not affected. **Recovery: `REINDEX INDEX` once after upgrading** an index that has taken such writes — the bug corrupted entries already on disk, and upgrading stops new damage but does not repair old. 1-bit (`bit_width = 1`) and graph indexes use other write paths and are not affected. - **Per-scan memory leak that could OOM-kill a backend (and restart the cluster) under concurrent writes.** Present since v1.8.0; found by this release's sustained-insert soak. `amendscan` was a no-op, so every finished index scan leaked its Rust-side scan state, including a reference to the backend's cached copy of the whole in-memory index. While nothing changes that costs little, but after another session commits, the next scan replaces the cached copy and the leaked references keep every superseded copy alive: a long-lived connection (a pooled one, say) that keeps querying an index other sessions keep writing to grows by about **one in-memory index per commit it observes**. Measured: +200 MB per commit on a 200k × 1024-d 4-bit index, on v2.10.3 and v2.11.0 alike; in the soak a scanning backend reached 52 GB and the OOM killer took it, and with it the postmaster (crash recovery). Fixed: the same workload now holds at 241 → 248 MB over 30 commits. Repro test: `amendscan_releases_cached_index_handle` (fails before the fix with 6 live references after 5 scans; passes with 1). A scan that aborts with an ERROR still skips `amendscan`, so it leaks once; that is bounded and noted in the code. **If you run v2.10.x or earlier with long-lived connections and concurrent writes, this alone is a reason to upgrade.** Short-lived connections were not affected: the memory is freed when the backend exits. ### Behaviour change: approximate candidate set The staged search returns **bit-exact scores** but may miss a candidate whose sign bits alone rank it outside the shortlist. On the Cohere corpus above, staged vs whole-index scan returns the identical top-10 id set for 98.5–100% of queries (≥99.6% mean overlap), and the identical top-100 for 73.5–99% (≥99.6% mean overlap — a tail candidate or two swapped). The exact heap recheck re-ranks what turbovec returns, and recall@10 against exact ground truth is unchanged at every measured `search_k`. To restore the whole-index scan, set **`TURBOVEC_4BIT_PLANES=0`** and/or **`TURBOVEC_2BIT_PLANES=0`** in the postmaster's environment (read once per process). ### Safety The new risk is that after an `INSERT`/`DELETE`, turbovec reconstructs the packed codes we persist from the planes cache. Gates: - **Byte-identical persisted index**: the same 1M × 1024-d heap indexed by 2.10.3 and by 2.11.0 on Graviton4 has identical meta + codes/scales/ids chains (sha256). - New `#[pg_test]` **`persist_is_byte_exact_through_planes_layout`**: 40k rows (past the planes gate), real remove/add + PreCommit flush, every row's in-memory and persisted bytes checked at 2 and 4 bits. Confirmed to take the planes layout on Graviton4. - `cargo pgrx test pg16`: **448 passed / 0 failed / 8 ignored** on x86_64 (446 on aarch64 Graviton4 before the two fix tests were added). - **Sustained-insert soaks** on Graviton4 with mid-flush `pg_terminate_backend`, periodic VACUUM and UPDATE churn, then a byte-level comparison of every persisted entry against a fresh `CREATE INDEX` of the same heap. These found both bugs under Fixed; the A/B re-run against v2.10.3 is in `benches/results/tv111_arm_20261005/FINDINGS.md` §6. turbovec's own suite: 523 passed / 0 failed on Graviton4. ### Fork Pinned to `gburd/turbovec@pgtv-2.11.0-port` (`455549f`): upstream 1.1.1 + carries #1–#3 cherry-picked unchanged (`pub repack`, IdMapIndex parts API, parallel repack; upstream issues #545–#547) + new carry #4 (parallel planes repack). ### Not measured x86 AVX-512 VBMI+VNNI hosts (no measurement here; upstream reports similar kernel gains); IVF and ColBERT latency; indexes under 32,768 rows (unchanged by construction). ### Build note turbovec 1.1 needs LLVM ≥ 22 for its AVX-512 VNNI intrinsics; the nix-packaged `rustc` 1.97.0 (LLVM 21) fails with `intrinsic signature mismatch`. Any rustup toolchain ≥ 1.89 works (CI uses rustup `stable`). On aarch64 Linux, `cargo pgrx test` (debug profile) fails to assemble `gemm-common`'s fp16 inline asm without `RUSTFLAGS="-C target-feature=+fp16"`; the `--release` build (`cargo pgrx install --release`, what ships) compiled cleanly without it. This is pre-existing (`gemm` 0.18.2, unchanged by this release). ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.11.0';` and restart PostgreSQL. The format is unchanged, so no REINDEX is *required*, but **REINDEX any 2/3/4-bit flat or IVF index that has taken writes from long-lived connections while VACUUM was running** to clear entries the `touched_ids` bug (above) already corrupted. Evidence: `benches/results/tv111_arm_20261005/`. ## [2.10.3] — 2026-09-25 PATCH: **cold-scan latency**. No SQL surface change, no GUC change, **no wire-format change** (`MetaPageData::version` stays **8**), index bytes unchanged. `ALTER EXTENSION pg_turbovec UPDATE` is sufficient; **no REINDEX**. ### Changed - **Cold-scan latency cut ~3× by parallelizing the per-backend blocked-layout rebuild.** Every cold backend rebuilds the SIMD-blocked code layout from the row-major packed codes at index-open (v7+ persists only the codes, halving the on-disk footprint). That repack was single-threaded and was the dominant term in cold latency. It now runs in parallel across block-aligned ranges. Measured A/B on identical hardware/corpus/index (c7i.4xlarge, 16 vCPU, AVX-512; 1M × 1024-d 4-bit flat, 534 MB): **cold-backend p50 1766 ms → 566 ms (3.1×)**; warm-backend p50 unchanged (30.4 → 30.6 ms, within noise — a warm backend never repays the repack). See `benches/results/rebench_20260925/COLDSCAN_FINDINGS.md`. The parallel repack (turbovec fork carry #3, rev `47a26a3`) produces **byte-identical** output to the serial version, pinned by turbovec's `parallel_repack_is_byte_identical_to_serial` (bit-widths 2/3/4, sub/above the parallel threshold, tail-padding shapes). This is a speed change only; the persisted format is untouched. ### Docs - **RETRACTED the "we LOSE ~490×" latency scoreboard** in `docs/PARITY_GAPS.md`. A corrected end-to-end benchmark (top-level `EXPLAIN(ANALYZE)` Execution Time, **literal** query vectors — a query-vector subquery in the ORDER BY had added ~90 ms of InitPlan overhead to BOTH engines and produced the bogus 2552 ms figure; one warm psql session per arm) on 1M × 1024-d Cohere-wiki shows **flat-bw4 at 5.2 ms / R@10 = 1.000 BEATS pgvector HNSW at R@10 ≥ 0.95** (HNSW 8.6 ms) and is 3× faster at ≥ 0.98. IVF-bw1 is within 2.4–2.8× of HNSW at 58× smaller storage; IVF-bw4 is the weakest arm (confirms flat > IVF for `bit_width ≥ 2` at this scale). Full iso-recall table + corrected harness in `benches/results/rebench_20260925/`. - **Documented `turbovec.ivf_max_delta_pct` + "fewer, larger transactions"** guidance for continuous high-ingest into an already-large IVF index (`docs/PARITY_GAPS.md`), and recorded the sparse-ANN and bulk-INSERT design memos (`benches/results/parity_20260925/`). ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.10.3';` — that's all. No REINDEX, no downtime, index bytes unchanged. ## [2.10.2] — 2026-09-24 PATCH: documentation plus one build-time `NOTICE`. **No SQL surface change, no GUC change, no wire-format change** (`MetaPageData::version` stays **8**), index bytes unchanged. No REINDEX. ### Added - **A flat build over 100 k rows now emits a `NOTICE` naming `WITH (lists = N)`.** `lists` defaults to `0`, so a plain `CREATE INDEX ... USING turbovec` builds a **flat exact scan** — and nothing told the user that an approximate, cell-pruned IVF layer was one reloption away. An evaluator concluded from exactly that experience that pg_turbovec *"does not support ANN"* and chose a different extension. That is a discoverability failure on our side, not a misreading. Deliberately a `NOTICE`, not a `WARNING`: flat is frequently the **better** choice, so it must not read as a fault to fix. Suppressed for IVF, graph and ColBERT builds (already non-flat) and below 100 k rows, where flat is unambiguously right and the message would be noise. - **README: "Common objections, answered with measurements".** Four things evaluators say, each answered from our own benchmarks — including the two where the objection is *correct*: - *"doesn't support ANN"* — it does; the **default is exact**, which is why it looked absent. - *"HNSW has a high memory footprint and is slow to build"* — **agreed**, which is why we didn't build on it. Measured 10 M × 1536-d: our index is **4.4× smaller** (14.9 vs 65.5 GiB) and builds **2.6× faster** (1 h 24 m vs 3 h 38 m), and our own log shows HNSW slowing **super-linearly past 5 M rows**. - *"pg_turbovec's own build memory was worse than HNSW's"* — **true, and now fixed.** That benchmark measured us at **121 GiB peak + 60 GiB swap** against HNSW's 16.9 GiB. Root cause fixed in v2.10.1; measured after: 2 M × 1024-d peak **12.16 → 3.45 GiB**, and 10 M × 1024-d went from **OOM-killing a 61 GiB host** to **completing at 11.20 GiB**. Scope stated plainly in the README: the post-fix numbers are 1024-d, and 10 M × **1536-d** (the dimension the 121 GiB figure used) has **not** been re-measured. - *"HNSW is faster on query latency"* — **true, and we don't dispute it.** - **README: "Choosing `lists` — the ANN tuning knob".** A full tuning guide, because the honest answer is more subtle than "turn on ANN": - **Whether you want IVF at all.** At 1 M × 1024-d with the default `bit_width = 4`, flat is **6.08 ms at recall@10 = 1.000** while `lists = 1024` is **62 % slower and capped at recall 0.959** — a per-probe ceiling that is CPU-independent and that **no amount of tuning removes**. IVF is a trade, not a free speedup. - The measured decision table (by `bit_width`, scale, recall target, and whether the index exceeds RAM). - **`√n` is a ceiling, not a target** — at 1 M, `lists = 4096` measured *worse than 1024 on every axis*: 11× the build time and ~50 % higher latency. - Tune **`probes` at query time**, not `lists` (which is baked in at build). - Measure on your own data, at **matched recall**. ### Unchanged, deliberately - **`lists` still defaults to `0` (flat).** Changing the default to `√n` was considered and **rejected on our own measurements**: at the default `bit_width = 4` it would ship a configuration that is 62 % slower *and* recall-capped at 0.959 up to at least 1 M rows. The defect was discoverability, not the default value. ### Tests 445 passed / 0 failed / 8 ignored, uniform across pg13–19 native plus the classic lane. New: `flat_build_hints_at_the_ann_option`, which also pins the silence (IVF builds and sub-threshold indexes must not be nagged). ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.10.2';` — nothing else. No REINDEX. ## [2.10.1] — 2026-09-23 PATCH: build-time memory profile only. **No SQL surface change, no GUC change, no wire-format change** (`MetaPageData::version` stays **8**), and the on-disk index bytes are **identical** — guarded by the IVF byte-identity tests. No REINDEX. ### Fixed - **Phase Z6 — `CREATE INDEX` / `REINDEX` peak memory cut ~3.5×.** The build callback had **no memory-context management at all**: no switch, no reset, no `pfree`. `Vector` is a `PostgresType` stored as CBOR, so `FromDatum::from_datum` palloc's a decoded buffer for every row — and those buffers accumulated in the long-lived `ambuild` context for the entire scan. The `CorpusSpill` was faithfully streaming the corpus to disk while PostgreSQL held a decoded copy of **all of it** in RAM, defeating the spill's whole purpose. `build_callback` is now a thin wrapper that switches into a per-tuple context, calls the unchanged inner callback, restores, and resets. The inner function has six early-return paths, so doing this inline would have been six chances to leak the switch. One hazard came with it: `BufFileCreateTemp` palloc's in the *current* context and the spill is opened **lazily on the first row** — inside the context now being reset, which would have left a dangling `BufFile` on row two. `CorpusSpill::new_in(cxt, dim)` plus an explicit `BuildState.build_cxt` keep it in the long-lived context **by construction** at all three lazy-open sites (single-vector, BQ/graph, ColBERT). Measured on 2M × 1024-d, `lists = 1414` (`benches/results/z6_buildmem_20260922/`): | | before | after | |---|---:|---:| | at drain entry | 12.07 GiB | **2.39 GiB** (5.05× less) | | whole-build peak | 12.16 GiB | **3.45 GiB** (3.5× less) | | build wall time | 1055 s | **889 s** (16 % faster) | This is what made 10M × 1024-d builds OOM-kill a 61 GiB host. **Since measured** (`benches/results/z6_10m_20260924/`): on the same c7i.8xlarge / 61 GiB host with the identical config that died before (`mwm = 8GB`, 16 parallel workers, `lists = 3162`), the 10M × 1024-d build now **completes in 69.8 min at 11.20 GiB peak private** — 18 % of the host, zero OOM events, against `anon-rss` 58.47 GiB at the kill. The projection above was 11.2 GiB. Index verified sound (wire v8, 10M/10M slots, `is_corrupt = false`, `scan_fraction = 0.00506`, 51 ms warm). - **The Lloyd `cross` matrix was quadratic in `lists`.** `train_kmeans` allocated `n_sample × lists` where `n_sample = lists × 256` — **9.54 GiB at `lists = 3162`**. Now chunked over sample rows under a fixed ~256 MiB budget. Bit-identity is tested, not assumed: `kmeans_cross_chunking_is_bit_identical` compares the whole-sample call against chunk sizes 1/7/64/199/500. ### Changed - **`turbovec.build_parallelism` is a memory knob, and its description said the opposite** ("only build wall-clock changes"). Measured: 16 threads → 16.24 GiB peak private, 1 thread → 12.16 GiB — ≈ **0.27 GiB/thread** of thread-local GEMM packing buffers, which `maintenance_work_mem` does **not** bound. Corrected, with the advice to lower it when `CREATE INDEX` is memory-constrained. - **`trace_stage!` (under `TURBOVEC_BUILD_TRACE`) now reports private memory** per stage, plus a new `0_at_drain_entry` marker. Peak had been misattributed for three sessions because the only signal was a process-wide total — which also includes `shared_buffers`. That single marker localised this bug in one build. ### Documentation - A **constraint** is now pinned by test: `rotate_corpus_into` is **not** invariant to row-block shape (`Parallelism::Rayon(0)` makes the GEMM's reduction order depend on `m`), so the reservoir rotation cannot be chunked to save memory. A blocked-rotation optimisation was written and **CI caught it** before it could change index bytes; `rotate_corpus_is_not_row_block_shape_invariant` records why. The pre-existing `rotate_corpus_bit_identical_across_pool_sizes` does not cover this — it varies thread count at a *fixed* shape. ### Tests 444 passed / 0 failed / 8 ignored, uniform across pg13–19 native plus the classic lane. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.10.1';` — nothing else. No REINDEX. ## [2.10.0] — 2026-09-22 MINOR: one new GUC and a changed IVF insert/scan behaviour. **No SQL-surface change and no wire-format change** (`MetaPageData::version` stays **8**) — existing indexes decode byte-identically and **no REINDEX** is required. ### Changed - **Phase Z5 Route A — an IVF index no longer loses its cell layout on the first `INSERT`.** `aminsert` cannot place a row into its cell without an O(n) reshuffle, so appended rows land at the tail, outside every cell. Previously the deferred-commit flush dropped the coarse-centroid and cell-directory chains outright, so **one commit turned the index into an O(n) flat scan until `REINDEX`** — an operational cliff with no cheap recovery. The flush now writes those chains back **unchanged**, and the scan additionally sweeps the tail exhaustively. Results stay **exact**; only latency is affected. **No new meta field was needed.** The delta length is *derivable* as `n_live - cell_directory.total_vectors()`, because existing rows keep their slot (an UPDATE writes in place) and inserts append at the tail. An earlier assessment in `docs/PARITY_GAPS.md` called this "blocked on a wire-format change"; that was wrong, and the correction is recorded there. Bounded by the new **`turbovec.ivf_max_delta_pct`** (default `10`, range `0..=100`): past the bound the index degrades to flat and reports it exactly as before, so the tail cannot grow until the optimisation is an O(n) scan with extra bookkeeping. **Setting it to `0` restores pre-2.10.0 behaviour byte-for-byte.** ### Fixed - **Out-of-core path could make appended rows unreachable.** The OOC gather iterates only *probed cells*, so a tail outside every cell was never read — those rows existed on disk and could never be returned. That is silent loss, not slowness. The delta is now fed through the **same** tombstone-aware run-splitting as cells, so a tombstoned appended row cannot resurrect (the v2.7.0 class of bug). ### Documentation - **`shared_preload_libraries = 'pg_turbovec'` is now documented as required** (README + first section of `PRODUCTION.md`). Every `turbovec.*` GUC is registered by `_PG_init`, i.e. at library load; without preloading a backend has **zero** of them, `SET turbovec.probes = 16` is accepted and silently ignored, and every query runs at compiled-in defaults. Indexes still build and queries still return correct results, which is exactly why it misdiagnoses as "tuning has no effect" or "IVF pruning doesn't work" — it cost a full benchmark round before being spotted. Includes the one-line check (`SELECT count(*) FROM pg_settings WHERE name LIKE 'turbovec.%'`) and the `LOAD` / `session_preload_libraries` alternative for managed providers (verified: every GUC is `Userset`, with no Postmaster/Sighup context). ### Measured (EC2, `benches/results/z5_delta_20260922/`) c7i.4xlarge (16 vCPU, AVX-512), PostgreSQL 16.15, 1M × 256-d, `bit_width=4`, `lists=1024`, 1000 inserts, 30 warm queries per arm in one session, fresh index per arm, autovacuum disabled, index health verified after the inserts: | arm | in-memory | out-of-core | |---|---:|---:| | pre-Z5 (`pct=0`) → **degraded**, `scan_fraction` 1.0 | 4.51 ms | 4.43 ms | | **Z5 delta** (`pct=10`) → healthy, `scan_fraction` 0.0156 | **4.03 ms** | **3.46 ms** | **The win is 11 % in memory and 22 % out-of-core — not the 64× that was modelled.** The model assumed latency scales with rows scanned, but at 1M × 256-d a *full* 4-bit scan is only **1.6×** a 1-cell scan (5.97 vs 3.66 ms): the 139 MB index is RAM-resident and per-query fixed costs are the same order as the SIMD sweep. This agrees with our own published 1M bw4 result (flat **6.08 ms beats** IVF 16.04 ms; v2.8.3 "flat wins at every target"). **This release ships on the functional contract — no cliff on insert, plus the OOC correctness fix — not on the latency delta.** The >RAM regime, where a full scan is disk I/O and the win could be materially larger, is **explicitly unmeasured**. ### Corruption validation (HARD MANDATE) 6 concurrent writers + 4 readers + `VACUUM` every 30 s for 5 minutes, with autovacuum enabled: `is_corrupt = false`, no duplicate id, `n_vectors == slot_count` exactly, wire still v8, and the cell layout **survived** (`degraded = false`, `scan_fraction = 0.0156`). Two apparent discrepancies were investigated rather than assumed: the index holding 1072 more rows than the heap (lazy vacuum reclaim of 10072 dead tuples) and a 1000-row sample returning 254 (`turbovec.search_k = 32` caps candidates). ### Tests 442 passed / 0 failed / 8 ignored, uniform across pg13–19 native plus the classic lane. New: the delta-bound predicate (accept/reject, `0` disables, inconsistent layouts), plus two integration tests asserting **both** halves — cells preserved *and* every appended slot swept even when a single cell is probed, while unprobed cells stay excluded (otherwise the pruning is gone and it is a flat scan in disguise). ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.10.0';` — nothing else. No REINDEX. To keep the previous behaviour exactly: `SET turbovec.ivf_max_delta_pct = 0`. ## [2.9.0] — 2026-09-22 MINOR: adds one SQL function. **Wire format unchanged** from 2.8.x (`MetaPageData::version` stays **8**), so existing indexes decode byte-identically and **no REINDEX is required** — the upgrade is in place. ### Added - **`turbovec.index_degradation(regclass)`** — quantifies an IVF degradation instead of merely flagging it. Returns `degraded`, `lists`, `n_vectors`, `scan_fraction`, `est_slowdown` and a `recovery` string. Phase Z1 made degradation observable and Phase Z4 made the planner cost it correctly; neither told an operator the **size** of the problem, which is what decides whether to act — a degraded 10k-row index is a non-event, a degraded 10M-row index is an outage. A degraded index reports `scan_fraction = 1.0` (it reads everything) and `est_slowdown = lists / probes`, plus the `REINDEX` command naming the index. A flat index is explicitly **not** reported as degraded — it scans everything by design, and faulting it would train operators to ignore the signal. Reads only the meta page (one buffer hit), so it is safe to poll from monitoring. ### Changed - **Phase Z4 — `amcostestimate` is now probe- and filter-aware.** Three defects, all of which made the planner blind to what the AM does: - IVF was costed as a full-corpus scan although the scan clamps to `turbovec.probes` cells, so an index probing 1 of 1024 cells was costed identically to a flat scan of everything. Cost now scales with `probes / lists`, floored at one cell's worth. A **degraded** IVF index is costed as flat, since that is the path it takes. - `index_selectivity` was hardcoded to `0.0` for every query. It now derives from the planner's own `rel->rows / rel->tuples`, so we agree with the planner by construction instead of second-guessing it. - A **pre-existing unit error**: the ns→cost conversion divided *seconds* by `cpu_operator_cost`, making a 1M × 1024-d flat scan cost ~23 while PostgreSQL costs the equivalent sequential scan at ~73,000 — about 3000× too cheap, which let an ANN path beat plans that are genuinely faster. Now ~1688. The ns throughput model itself validated against our own published measurement (model 5.3 ms vs measured 6.08 ms), so only the unit was wrong. `index_pages` is likewise scoped to the pages a probed scan touches. The arithmetic lives in two pure, unit-tested functions because it is **not** observable through `EXPLAIN`: the index-scan node's cost also carries PostgreSQL's heap-fetch and qual costs, which swamp it. ### Documentation - **`docs/FILTERING.md` — the allowlist crossover is now actionable.** The measured table (2.6–14.7× faster below ~7 % selectivity, 2.6× *slower* at 100 %) was only useful if you knew your filter's selectivity. Adds a copy-pasteable way to read PostgreSQL's own estimate, a fraction→technique table, and two caveats: the crossover percentage is host-dependent (the *shape* transfers, not the number), and the estimate is only as good as your statistics. - **Phase Z3 rescoped to a non-gap, with the reasoning recorded.** "Automatically turn a `WHERE` into a kernel mask" is not implementable by a PostgreSQL AM: a scan key is `index_key operator constant` over an *index* column, so a qual on any other column becomes an executor Filter and never reaches the AM. `amgetbitmap` is not a route either — it returns an unordered bitmap, and ordering is the whole value of an ANN scan. The useful behaviour shipped in v1.8.0 as iterative scan, which is demand-driven and needs no view of the filter. - **Phase Z5 scoped, with both routes shown blocked.** A bounded mutable delta needs a wire-format change (both scan paths assert `directory.total_vectors() == n_live` and `centroids.len() == lists * dim`; a synthetic "delta cell" would need a centroid it does not have). Cheap in-place cell reassignment needs a `decode`/`reconstruct` that turbovec does not expose (`pack.rs` has only `repack`/`unblock`). The reporting half shipped instead. ### Tests 435 → **437 passed / 0 failed / 8 ignored**, uniform across pg13–19 native plus the classic lane. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.9.0';` — creates the new function. No REINDEX. ## [2.8.4] — 2026-09-21 Code-only release. Wire format unchanged from 2.8.3 (`MetaPageData::version` stays **8**); no REINDEX needed and no SQL surface change. ### Fixed - **Phase Z1 — an IVF index that takes writes now REPORTS its degradation (ordinary TurboQuant path).** An `aminsert` cannot place a row in its cell without an O(n) reshuffle, so the deferred-commit flush appends and the index falls back to a flat scan. The 1-bit BQ path always *preserved* `lists` and stamped `ivf_degraded` so `turbovec.index_is_degraded()` and the throttled `ambeginscan` WARNING fired; the TurboQuant path blanked `lists`, so `index_was_ivf()` went false and the latency cliff was **silent** — an operator got a quietly slower index with no signal and nothing to act on. Root cause: `reconcile_and_write_flush` planned its meta page via `plan_with_blocked`, which hardcodes `lists: 0`. It now captures the on-disk `lists` under the already-held exclusive rewrite lock and stamps both fields, leaving the coarse/cell-directory offsets at zero so the readers return empty and the scan takes the flat fallback *deterministically* rather than by a length coincidence. Both fields are existing v4 meta scalars — no chain is added or moved, so the chain-offset running-sum class is not implicated. - **An `INSERT` into an index built `WITH (assign_dups > 1)` no longer claims the index is corrupt.** IVF-4a soft assignment stores a boundary row in several cells on purpose, so its external id appears in several slots and `slot_to_id` is deliberately not a bijection — while the insert path loads the index into a flat `IdMapIndex`, which requires one. The rejection reported `corrupt relfile pages: duplicate ids` with a REINDEX hint; both halves were wrong (`turbovec_check` verifies such an index clean, and a rebuild reproduces the same by-design duplicates), sending operators hunting for corruption that does not exist. It now reports `FEATURE_NOT_SUPPORTED`, names `assign_dups`, states the index is effectively **READ-ONLY**, and the HINT says how to confirm it is healthy. Behaviour is otherwise unchanged: the INSERT still fails, `SELECT` still works. Zero cost on the healthy path — `from_id_map_parts` has exactly one failure mode, so reaching the handler already identifies the cause. ### Documentation - `docs/PARITY_GAPS.md` gains a **zvec source review** (local `alibaba/zvec` checkout `d88357b` vs `deac2d9`), annotated throughout as un-benchmarked: four real gaps in priority order, the one lesson worth stealing (bounded mutable delta + explicit consolidation), explicit non-gaps (hybrid fusion, ColBERT, WAL, durability, scalar filtering — PostgreSQL's or already ours), and deliberately deferred items (RaBitQ/PQ, DiskANN). Adds phases **Z1–Z5**; Z3 (automatic predicate→ANN handoff) is gated on Z4 (costing), because pushing filters while `index_selectivity` is hardcoded to `0.0` would only make bad plans confident. - `assign_dups > 1` is now documented as making the index read-only, in both `docs/BENCHMARKS.md` (which recommended it for recall without saying so) and the reloption reference in `src/index/options.rs`. - `AGENTS.md` gains four steering rules: degradation must be observable; competitor comparisons stay source reviews until measured; the BQ and TurboQuant insert paths write at **different times** (BQ synchronously in `aminsert`, TurboQuant deferred to `PreCommit`, which a `#[pg_test]` never reaches — so a plain `INSERT` in a test exercises nothing on that path); and `assign_dups > 1` indexes are read-only. ### Tests 429 → **430 passed / 0 failed / 8 ignored**, uniform across pg13–19 native plus the classic lane. New: `ivf_flush_degradation_is_reportable`, `ivf_soft_assign_index_rejects_insert_and_is_not_corrupt`, `ivf_soft_assign_insert_error_names_assign_dups`. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.8.4';` — nothing else. No REINDEX. ## [2.8.3] — 2026-09-11 **`bit_width = 4` + IVF measured at 1M — flat wins at every target.** This answers the production user's question with data rather than the forecast v2.8.2 shipped. Documentation only; the binary is byte-identical to 2.8.2. No wire change (v8), no SQL surface change, no REINDEX. ### The measurement 1M × 1024-d real Cohere corpus, AVX-512, 48 configs, **every one confirmed to run a real `Index Scan`**. Artefacts: `benches/results/bq_1m_bw4ivf_20260911/`. | target | bw4 flat | bw4 `lists = 1024` | | |---|---|---|---| | R@10 ≥ 0.90 | **6.08 ms** (w=32) | 16.04 ms (p=64) | flat by 62 % | | R@10 ≥ 0.95 | **6.08 ms** (w=32) | 16.13 ms (p=128) | flat by 62 % | | R@10 ≥ 0.98 | **6.08 ms** (w=32) | **unreachable** | flat only | | R@10 ≥ 0.99 | **6.08 ms** (w=32) | **unreachable** | flat only | **The write-up leads with the recall ceiling, not the latencies, because the conclusion does not depend on any timing.** At `probes = 128`, widening the rerank window from 32 to 2000 leaves recall at **exactly 0.959 across all 8 windows** — it does not move by a single query, because the true neighbours are not in the probed cells. Recall is CPU-independent, so **discard every latency number and 4-bit IVF still loses at the top two targets.** Flat's *cheapest* config is also its *most accurate* (R@10 = 1.000 at window 32), so IVF never gets an opening. The 62 % gap is corroboration. Published for completeness rather than left for a reader to find: the one sub-flat p50 in 48 rows is at window 2000, IVF 116.77 ms versus flat 122.19 ms (0.96×). It is **not** a win — 0.959 recall against 1.000 — so it fails a matched-recall comparison. The tightest honest framing: IVF's fastest configuration *anywhere* is 15.76 ms at R@10 = 0.875, against flat's 6.08 ms at R@10 = 1.000. ### The OOM worry is retired for this configuration Possibly more actionable for an operator than the latency result. At 1M × 1024-d with `maintenance_work_mem = '4GB'` and 4 parallel maintenance workers: | arm | bytes/vec | index | build | peak build anon-RSS | |---|---:|---:|---:|---:| | bw4 flat | 559.95 | 534 MB | **14.89 s** | **2.641 GiB** | | bw4 `lists = 1024` | 564.17 | 538 MB | **101.22 s** (6.8×) | **2.154 GiB** | **Real peaks, not lower bounds** — 0.25 s sampling, 466 in-build samples, from a pure-bash sampler that left idle loadavg at 0.00–0.01. IVF's peak is *below* flat's, and 2.6 GiB is nowhere near the 20.3 GB an **unbounded** `3GB` setting reached on the earlier 250k build. **The hazard is leaving `maintenance_work_mem` unbounded, not 4-bit IVF.** The genuine cost is the **6.8× build time**. ### Caveats, recorded rather than glossed - All 48 latency rows are contention-flagged, and the unflagged-row filter is **unavailable** here: 0 of 8 flat and 0 of 40 IVF rows survive it — exactly the trap § 0.6e documents. The bias runs *against* IVF (flat ran at 22.0 % mean `cpu_busy` versus IVF's 6.1 %, a 3.64× difference; mean loadavg 1.87×) and flat still won by 62 %, so a quiet re-time could only widen flat's margin. - **A second corpus-identity surprise.** This arm's resolvability spread was **18.6–69.1 %** against § 0.6e's **154–203 %** on nominally the same corpus and shards. Both clear the ~10 % unusable floor so each run's internal comparison stands, but the discrepancy is unexplained and the two runs must **not** be treated as same-corpus. The first such surprise forced the v2.8.1 corrections. - Ground truth took **438.4 s** against § 0.6f's ~275 s on the same parallel-CTAS path. Unexplained, flagged so a future diff does not misread it as a harness regression. - **Scope:** a ≥ 0.98 user is answered at *any* scale, since the per-probe ceiling is structural rather than a size effect. A ≲ 0.90 user well above 1M is **not** answered by this arm. Also closed a gap the previous 1M run left open: held-out queries are **value**-disjoint as well as id-disjoint (`value_overlap = 0` via md5 hash join). ### Operational: AWS burner rules `AGENTS.md` had no AWS section, which is what let a mid-run account expiry stand a benchmark instance up with no way to terminate it (account `bene` expired at 12:03 UTC, ~48 min after launch; `InvalidClientTokenId` on a previously-working key means the *account* went away, not that you broke something). Added the current burner (`lava`) and the rules that made that incident cost **zero data** — chiefly *pull artefacts as you go*. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.8.3';` — no REINDEX. ## [2.8.2] — 2026-09-10 **4-bit IVF is supported — the "1-bit-only" result was about *benefit*, not support.** Documentation only; the binary is byte-identical to 2.8.1. No wire change (v8), no SQL surface change, no REINDEX. ### The report, and the answer A production user running `bit_width = 4` read v2.8.0's "the 1M IVF+BQ crossover is 1-bit-only" as meaning IVF cannot be combined with 4-bit, and asked whether support could be added. **It already exists and always has.** `WITH (lists = N)` composes with **every** `bit_width` — 4-bit IVF is the *original* IVF path, out-of-core end-to-end since v1.13.0, and 2-bit and 3-bit work too. The only `bit_width`/kind combination the code rejects is `bit_width = 1` with `graph = true` (`src/index/options.rs`). Verified by grepping every rejection site: there is no `bit_width` gate on IVF anywhere in the tree. **Nothing to enable, nothing to wait for.** The misreading is this project's fault, not the user's — the guidance paragraph sits inside the README's 1-bit section, so a 4-bit reader lands on it naturally. Corrected in the README, and § 0.6e's heading is restated as "the crossover EXISTS, and the ***benefit*** is 1-bit-only" with a callout naming the misreading so the next reader does not repeat it. ### New § 0.6g — a 4-bit user's three questions, answered separately They were being conflated: 1. **Is it supported?** Yes. 2. **Will it help?** Probably not — **and this combination was never measured**, stated plainly rather than implied. `bit_width = 4` + `lists = N` does not appear in *any* artefact at either scale; bw4 is only ever swept flat. The mechanism predicts no win: the crossover needs **both** an expensive `O(n)` scan **and** a quantizer lossy enough to demand a wide rerank window, and 4-bit needs only a **32-wide** window versus **800** for 1-bit, so its scan is already cheap and cell-restriction mostly adds overhead. 2-bit is the direct evidence — a clean loss at 1M for exactly that reason. 3. **What could it cost?** Two things. The **per-probe recall ceiling** (0.986 at `probes = 128`; R@10 ≥ 0.99 unreachable at any setting, and widening the rerank window does *not* recover it, because the true neighbours are not in the probed cells). And **build memory** — `bw4 + lists = 512` at 1024-d reached 20.3 GB anon-RSS and was OOM-killed at `maintenance_work_mem = 3GB`. ### New: "Should you enable IVF?" in `docs/PRODUCTION.md` A decision table plus a **20-minute experiment an operator can run on their own data**: baseline at *their* recall target, build an IVF copy with `maintenance_work_mem` bounded, sweep `probes`, compare at **matched recall** (not matched settings — an iso-knob comparison flatters whichever index gets a wider effective window), and verify with `EXPLAIN` that it really is an `Index Scan` rather than a sequential-scan fallback dressed up as a latency result. The reasoning behind pushing users toward their own measurement: since v2.8.1 this project's own 1M and 250k figures come from **different corpora** (the older Cohere dataset became gated), so published cross-scale deltas are suggestive rather than measured. A user's corpus is the only authority for their workload, and a negative result from their data is worth more than a positive one from ours. A measurement of `bw4 + lists = 1024` at 1M is in flight and will replace § 0.6g's prediction with a number when it lands. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.8.2';` — no REINDEX. ## [2.8.1] — 2026-09-10 **Corrections to v2.8.0's benchmark write-up.** Documentation only — the binary is byte-identical to 2.8.0. No wire change (v8), no SQL surface change, no REINDEX. All three items came from the 1M benchmark agent's final report and were verified before being accepted. ### The 1M corpus is not the same corpus as the 250k runs v2.8.0 presented 1M-versus-250k deltas — IVF's storage overhead "halving", the per-probe ceilings "landing within a hair" of the 250k figures — as though both scales shared a corpus. **They do not.** `Cohere/wikipedia-22-12-en-embeddings` is now **gated** (confirmed HTTP 401), so the 1M arm used `CohereLabs/wikipedia-2023-11-embed-multilingual-v3`: same publisher and dimensionality, but a **different model and snapshot**. The 250k artefacts carry no corpus label at all, corroborating that they came from a different pre-existing table. Every cross-scale delta is now labelled **suggestive, not measured**. The **within-run flat-versus-IVF comparisons are unaffected** — both arms of each run share one corpus, and those are what the conclusions rest on. ### A trap in the method used to validate v2.8.0's headline The 47 % / 38 % IVF wins were validated by re-checking on contention-unflagged rows only. That is unsound whenever the **baseline** does not survive the filter — and on the `lists = 4096` arm it does not: all 8 bw1 *flat* rows are flagged (they ran first, while loadavg was still decaying from the k-means build) and **zero survive**, so a filtered comparison there would "prove" IVF wins against an empty set. Re-verified: the `lists = 1024` arm used for the headline keeps **all 8 flat rows unflagged**, so that check *was* valid — but partly by luck of execution order. The rule is now in `docs/TESTING.md` beside the existing control-arm rule: **assert the filtered baseline is non-empty before trusting a filtered comparison.** ### Restored: the ground-truth-fix documentation (now § 0.6f) An earlier rewrite of § 0.6b had overwritten it, and a later edit's assertion masked the loss — found by grepping for the measured numbers and getting nothing. **The code fix was never affected.** The restored section also records an accidental validation at 1M scale: the run's two arms straddled the fix, giving **3755.7 s** (pre-fix `INSERT` path) versus **275.4 s** (CTAS path) — **13.6×**, matching the 12× predicted from the plan shape — and the **16 flat configs shared by both arms reproduce bit-identically** across the two GT implementations, independent confirmation the fix changes no measured number. ### Also - 1M **build-memory figures relabelled as lower bounds**, not peaks: the sampler polled every 2 s and was stopped before the second arm, so the `lists = 4096` build has **no RSS measurement at all**. - The **contention cause is named**: loadavg decaying after each parallel k-means build (a monotonic decline across consecutive rows) plus an RSS sampler forking a Python interpreter every 2 s. `cpu_busy_pct` of only **3.1–3.2 %** with `cpu_steal ≤ 0.01` confirms it was never saturation. - **All 160 configs across both 1M arms ran a real `Index Scan`** — zero masked sequential scans. - Recorded honestly: held-out queries are **id-disjoint** from the corpus (join count 0), but **value-disjointness was not proven** — the 100 × 1M text comparison was abandoned as too slow. A duplicate would require the dataset itself to contain duplicate embeddings. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.8.1';` — no REINDEX. ## [2.8.0] — 2026-09-10 **The 1M IVF+BQ crossover, measured on a real corpus** — plus a parallelised ground-truth path and the v2.7.4 latency caveat resolved with data. Minor rather than patch because the harness's GT path changed shape (same output, different plan) and the operator guidance for `WITH (lists = N, bit_width = 1)` is now materially different. No index wire-format change (stays v8), no SQL surface change, no REINDEX. ### The crossover exists at 1M, and it is 1-bit-only Corpus: **CohereLabs/wikipedia-2023-11-embed-multilingual-v3 (en), 1 000 000 × 1024-d** — real, not synthetic — with 100 held-out queries, on an AVX-512 host. The § 0.6d resolvability gate passed at **154–203 %** nn1→nn100 spread (the discarded synthetic corpus was 6.6–10.4 %). | target | bw1 flat | bw1 IVF (`lists=1024`) | verdict | |---|---|---|---| | R@10 ≥ 0.90 | 25.6 ms | **13.7 ms** (p=64) | **IVF wins 47 %** | | R@10 ≥ 0.95 | 33.2 ms | **20.7 ms** (p=128) | **IVF wins 38 %** | | R@10 ≥ 0.98 | **39.3 ms** | 44.2 ms | IVF loses 12 % | | R@10 ≥ 0.99 | **56.8 ms** | unreachable | flat only | Re-computed using **only contention-unflagged rows**: identical 47 % / 38 %, so this is not a load artefact. **For `bit_width ≥ 2` it is a clean no** — flat is 4.8 ms at every target and IVF never gets under 16 ms. Mechanically the crossover needs *both* conditions: flat's `O(n)` scan grown expensive **and** a quantizer lossy enough to need a wide rerank window. 1-bit at 1M needs w=100–256 over 1M rows, so restricting to 64–128 cells of ~1000 rows is a real saving; 2-bit needs only w=32, so its scan is already cheap. **Two-axis guidance**, not one: - `bit_width = 1`, n ≳ 1M, target ≲ 0.95 → **`lists = N`** - `bit_width = 1`, target ≳ 0.98 → **flat** (IVF cannot reach it) - `bit_width ≥ 2` → **flat**, at least to 1M ### `lists = 4096` is worse than `lists = 1024` A useful negative: quadrupling the list count made **every** axis worse — 8.7 % more storage, **11×** the build (1067 s vs 94 s), and ~50 % higher latency at matched recall (21.5 vs 13.7 ms at R@10 ≥ 0.90). Each cell holds 4× fewer rows, so a given recall needs ~4× the probes (`p=256` where 1024 lists needed `p=64`). The `lists ≈ sqrt(n)` rule is load-bearing; exceeding it is a pure loss for BQ. Also confirmed: the **per-probe recall ceiling is structural**, not a small-corpus artefact (0.846/0.904/0.944/0.971/0.986 at probes 8/16/32/64/128 at 1M, within a hair of the 250k figures), and IVF's **storage overhead halves at 1M** (+3.1 % vs +6.0 %) as the fixed centroid/cell-directory cost amortises. ### Ground truth parallelised — in two steps, the second correcting the first Both verified to leave GT **row-for-row identical** (`old EXCEPT new` = 0, `new EXCEPT old` = 0 on `(qid, hit_id, rk)`, membership ignoring rank = 0), so **no published recall figure moves**: 1. The correlated `LATERAL` + `row_number() OVER (ORDER BY dist)` forced `WindowAgg → Sort → Gather`: every row shipped to the leader for a single-threaded sort. At 250k × 20 queries, `Gather (actual rows=250000, loops=20)` with a 13.9 MB leader quicksort, **148.2 s**. Replaced with a per-query InitPlan constant so each worker top-N sorts its own share: **76.8 s, 1.93×**. 2. **My first version of that fix wrapped the SELECT in `INSERT INTO ... SELECT`, which silently threw the parallelism away again.** PostgreSQL generates no parallel plan for a data-writing statement — only `CREATE TABLE AS` / `SELECT INTO` / `CREATE MATERIALIZED VIEW` are exempt. Serial `Seq Scan` at 7.49 s versus 3.99 s for CTAS; **~36 s vs ~2.9 s per query at 1M, a 12× gap**. Now CTASes each query's top-`k` into a temp table and inserts those few rows. The second defect was **caught by the 1M benchmark run's `pg_stat_activity`** (one backend at 99.9 % CPU, zero parallel workers), not by me. Root cause worth naming: I benchmarked a bare `SELECT`, measured 1.93×, then shipped an `INSERT` — a different statement with a different plan — and never re-timed. **Benchmark the statement you are going to ship, not a proxy for it.** ### The v2.7.4 contended-latency caveat is resolved with data The bench host's load problem was fixed, so the 250k sweep was re-run on the same corpus: **18 of 24 rows now clean** (was 0 of 24), **recall reproduced exactly**, the contended p50s were uniformly **14–16 % pessimistic**, and the published matched-recall ratios **held to two decimal places** (2.70× / 6.13× → 2.75× / 6.09×). The absolute milliseconds in the § 0 tables are left as published — ~15 % conservative — with the correction pointing at the clean artefact. An honest record beats a tidy one. ### Also `ivf_streaming_build_temp_file_cleanup` asserted on a **cluster-wide** `pg_ls_tmpdir()` delta, so a concurrent `#[pg_test]`'s transient spill file failed it — observed on a docs-only commit, which is definitionally not a regression. It now re-samples and fails only if the growth persists. Same shared-global-state class as the bench-harness collisions fixed in v2.7.4. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.8.0';` — no REINDEX. ## [2.7.6] — 2026-09-09 **Documentation consistency pass**, plus a narrowed root cause and salvaged numbers from the discarded 1M arm. No shippable code change — binary byte-identical to 2.7.5. No wire change (v8), no SQL surface change, no REINDEX. ### The BQ docs contradicted themselves Across the 2.7.3–2.7.5 runs the docs accumulated claims the measurements had already falsified. `docs/BQ_RECALL_BENCH.md` still *opened* by stating it "contains **no measurements**" and pointing at a README row that "currently and correctly reads **not yet published**" — both untrue since 2.7.3. `docs/ONEBIT_BQ.md` carried an orphaned fragment, stranded mid-paragraph by an earlier merge, asserting the harness "has **NOT** been run — no numbers exist yet". Every doc was swept for claims the measurements contradict; all are now consistent. The runbook sections are retained and relabelled as *how to reproduce* rather than *not yet done*. ### The discarded 1M root cause, narrowed — and a competing diagnosis ruled out A plausible alternative was raised: that the generator's uncorrelated `ARRAY(SELECT randn() ...)` subquery had been hoisted, making all 200 cluster centres identical — the same v1.24.0 test-harness bug class `AGENTS.md` warns about. It was **tested rather than accepted or dismissed**, and both halves matter: - **The hazard is real.** The subquery never references the outer `c`. Reproduced minimally: uncorrelated gives **1 distinct vector across 5 rows**; adding a correlating `WHERE c = c` gives **5**. Worth recording, because the behaviour is plan-dependent. - **It did not happen here.** Measured on the loaded corpus: **200 distinct centres**, 5000 distinct vectors per `cid`, same-cluster distance **0.108** vs cross-cluster **0.990**. `randn()`'s volatility forced per-row evaluation. So the earlier "distance concentration" diagnosis was right in substance but imprecise about *where*. Narrowed by measurement: **the clustering worked; the tie is inside each cluster.** Within one cluster a member's 100 nearest neighbours span just **7.22 %**, because 5000 iid Gaussian points at d = 768 concentrate — pairwise separation norm has relative spread ≈ `1/sqrt(2d)` ≈ 2.6 %. **Cluster separation is not sufficient for rankability**, which is the non-obvious part and the reason the original prompt's "make it clustered" instruction was not enough. ### Salvaged from that arm: valid 1M storage/build numbers Corpus geometry doesn't affect these — per-vector codes are `dim/8 * bit_width` and the flat build is a fixed `O(n·dim)` pass: | 1M × 768-d | build | bytes/vec | peak build RssAnon | |---|---:|---:|---:| | bw1 | 167.3 s | 104.87 | 8.98 GiB | | bw2 | 152.0 s | 209.72 | 5.95 GiB | | bw4 | 161.0 s | 404.76 | 6.18 GiB | `bw2/bw1 = 2.000` **exactly**, confirming the `dim/8` sign-code stride holds at 1M rows. Note bw1 has the **highest** peak build RSS despite the smallest output: BQ reads the corpus back resident to compute the corpus mean before it can take signs, so its build memory tracks `n · dim · 4` regardless of bit width. Worth knowing before building a 1-bit index on a RAM-constrained host. ### Harness defect: the ground-truth query serialises on the leader The GT query puts `row_number() OVER (ORDER BY )` in a Sort **above** the Gather, so parallel workers ship raw vectors to the leader and the leader recomputes every distance single-threaded. That is why ground truth took **23 084 s (6.4 h)** for 100 queries at 1M rows — a plan pathology in the harness, not a turbovec cost, and the main thing making a real 1M arm expensive. ### New: mandatory pre-flight for synthetic corpora `docs/BQ_RECALL_BENCH.md` § 0.6d. Before trusting any recall number from generated data, run the resolvability probe and require the nn1→nn100 spread to be comparable to a real corpus: | corpus | spread | verdict | |---|---:|---| | real Cohere-wiki 1024-d | **37–268 %** | rankable | | the discarded synthetic 768-d | **6.6–10.4 %** | unusable | With an explicit warning that `is_degenerate()` and rankability are **different checks** — a corpus can pass the first and still be statistically unrankable. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.7.6';` — no REINDEX. ## [2.7.5] — 2026-09-09 **The 1-bit dimension sweep is measured, and it produces the sharpest practical guidance the BQ work has yielded: 1-bit is a high-dimension technique.** Also corrects a `hi_dim_rerank` documentation error. No shippable code change — binary byte-identical to 2.7.4. No wire change (v8), no SQL surface change, no REINDEX. ### The penalty collapses as dimension rises 256/512/1024-d, 250 000 rows and 100 held-out queries per dim, matching § 0's setup. Artefacts: `benches/results/bq_dimsweep_20260909/`. 1-bit R@10 at a **fixed** rerank window rises with dim at **all seven** swept windows, no exception: | window | 256-d | 512-d | 1024-d | |---:|---:|---:|---:| | 32 | 0.394 | 0.581 | 0.744 | | 256 | 0.729 | 0.888 | 0.967 | | 800 | 0.866 | 0.966 | 0.994 | | 2000 | 0.933 | 0.990 | 1.000 | At **matched recall** the window penalty versus 2-bit collapses: | dim | 1-bit window for R@10 ≥ 0.95 | 2-bit | penalty | |---:|---:|---:|---:| | 256 | **4000** | 100 | **125×** | | 512 | 800 | 32 | 25× | | 1024 | 256 | 32 | **8×** | At 256-d, reaching R@10 ≥ 0.99 requires a window of **16 000** — reranking 6.4 % of the whole corpus. **1-bit is effectively unusable at 256-d and below; prefer it at 768-d and up.** Storage moves the same way (1.902× → 1.946× → 1.971× versus 2-bit) because 1-bit's fixed per-index overhead amortises away as dim grows: 30 % overhead over the `dim/8` ideal at 256-d, 11 % at 1024-d. Both axes favour 1-bit more strongly at higher dimension. Depth confirms it: R@100 at the `auto` default is 0.455 (256-d), 0.737 (512-d), 0.948 (1024-d) — at 256-d the default loses more than half the true top-100. **Validation.** The 1024-d arm was re-run from a fresh database on a re-sliced corpus and reproduced § 0's published recall **bit-identically at all seven windows**, p50s within 1–3 %. Independently re-verified against the published artefact. That validates the published figures *and* the rebuilt, isolated harness. > **Caveat that bounds this.** The 256-d and 512-d corpora are **prefix > slices** of the 1024-d Cohere-wiki embedding, not natively-trained > embeddings of those dimensions. A native 256-d model concentrates its > information into 256 coordinates; a truncated 1024-d vector keeps only the > first quarter of a representation spread across all of them. That likely > makes the sliced low dims look worse than a native model would, so the trend > is an **upper bound** on dim-sensitivity — directionally sound, magnitude not > transferable. Nothing was padded to fabricate a higher dim. Same contention caveat as § 0: recall and storage stand; absolute p50s are indicative. ### Correction: the 1-bit `hi_dim_rerank` special case is a no-op at dim ≥ 256 `docs/ONEBIT_BQ.md` and `docs/BQ_RECALL_BENCH.md` said `hi_dim_rerank` treats a 1-bit index as high-dim "at **any** dim", implying it widens BQ's rerank window generally. That is literally true of the code and **misleading about the effect.** The `auto` window is `clamp(effective_dim, 256..=1024)`, and since `effective_dim = max(dim, 256)` for 1-bit but `dim` for 2/3/4-bit, the two are **identical for every `dim >= 256`**. The special case only widens the window below 256-d. This *strengthens* the published results rather than weakening them: a 1-bit-vs-2-bit comparison at `dim >= 256` with default settings compares **equal** windows, so § 0 and § 0.6c are quantizer-vs-quantizer, not knob-vs-knob. Both docs now carry a table showing exactly where the special case bites. Found by re-deriving the clamp independently against the Rust source while chasing down why the sweep's pre-registered prediction P-D3 had named the wrong mechanism. ### On pre-registration Three of the sweep's four predictions were wrong in some respect — the storage direction was backwards, the `hi_dim_rerank` mechanism was misnamed (which is what surfaced the doc error above), and the latency-vs-dim shape was non-linear. Only the central hypothesis was confirmed, and it was understated. Recording predictions before the run is what made all of that visible instead of being quietly rationalised afterwards. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.7.5';` — no REINDEX. ## [2.7.4] — 2026-09-09 **Documentation accuracy and benchmark-harness isolation.** No shippable code change — the binary is byte-identical to 2.7.3. No wire change (v8), no SQL surface change, no REINDEX. ### Correction to v2.7.3's latency figures v2.7.3 published the 1-bit BQ p50s **without disclosing that the harness had flagged every one of them.** All 24 rows carry `latency.contention.contended_flag = true` — 1-minute loadavg 3.16–4.64 against the harness's gate of 1.5. The artefact recorded this correctly; the write-up did not surface it. That was my error and this is the correction. The gate is **unreachable on that host** and not because of the benchmark: a stuck `systemd --user` spinning at 77–86 % CPU for 31 days pins its idle loadavg near 2.0, so *any* run there is flagged. Mitigating measurement, taken rather than assumed: `cpu_busy_pct` on the pinned cores was only **16–26 %** — the load is runnable-elsewhere processes, not saturation of the bench cores. So the **ratios and the shape of the window-vs-recall curve stand**, and the absolute milliseconds are indicative. Recall, storage, bytes/vector and build time are CPU-independent and unaffected. `README.md`, `docs/RECALL.md`, `docs/ONEBIT_BQ.md`, `docs/UPGRADING.md` and `docs/BQ_RECALL_BENCH.md` now all carry the caveat. Found by a sub-agent checking its own contention flags; I had not been checking mine. ### Measured: IVF + 1-bit BQ `docs/BQ_RECALL_BENCH.md` § 0.6a, artefact `benches/results/bq_ivf_20260909/`. `WITH (lists = N, bit_width = 1)` builds and scans correctly. Storage overhead over flat BQ is **+6.0 %** — a fixed ~8.5 B/vector of coarse centroids and cell directory, so proportionally worst for the smallest codes. **The finding: IVF imposes a per-probe-count recall ceiling that a wider rerank window cannot break.** At `probes = 8`, 1-bit saturates at R@10 = 0.846 and stays there from window 256 through 2000 — the true neighbours are not in the probed cells, and exact re-ranking cannot invent them. Flat BQ reaches 0.994 with no probe tuning. This is the **mirror image of Gap-B (v1.25.0)**, and the distinction is the useful part: there, high-dim recall loss was *not* retrieval-bound (cell recall 0.98–0.996) and a wider window fixed it; here it *is* retrieval-bound and the window is irrelevant. Same symptom, opposite cause — diagnose which one you have before reaching for a knob. **At 250k, flat BQ dominates IVF+BQ.** That is a *scale-dependent* boundary, deliberately not a verdict: IVF's whole value is bounding scan cost as `n` grows, and 250k is below the crossover. Recording it as a boundary is what the graph kind's early iso-beam numbers should have done. ### A 1M-scale arm was discarded, not published A synthetic 1 M × 768-d run produced R@10 = 0.031 at window 32, which looks like catastrophic scale collapse. It isn't — **the generated corpus was statistically unrankable.** Resolvability probe: 1st vs 100th nearest neighbour differed by only **6.6–10.4 %** in cosine distance, versus **37–268 %** on the real Cohere-wiki corpus. When the top-100 is effectively a tie, no quantizer can rank it and "recall" measures the tie-break order. Cause, narrowed on 2026-09-09 with direct measurement: the **clustering worked** (200 distinct centres; same-cluster distance 0.108 vs cross-cluster 0.990, a 9× separation) — the tie is **inside** each cluster. Within one cluster, a member's 100 nearest neighbours span only 7.22 %, because 5000 iid Gaussian points at d = 768 concentrate: the pairwise separation norm has a relative spread of ≈ `1/sqrt(2d)` ≈ 2.6 %. A plausible competing diagnosis — that the generator's uncorrelated `ARRAY(SELECT randn() ...)` subquery had been hoisted, making all 200 centres identical (the v1.24.0 harness-bug class) — was **tested and ruled out**: the SQL genuinely invites that hoist (reproduced minimally: 1 distinct vector across 5 rows uncorrelated, 5 when correlated), but `randn()`'s volatility forced per-row evaluation here and the loaded centres are distinct. Discarded with a full post-mortem, including the probe to run before trusting any generated corpus, in `benches/results/bq_scale_20260909/DISCARDED.md`. Also surfaced by that arm: the harness's ground-truth query wraps `row_number() OVER (ORDER BY )` around the per-query scan, forcing a `WindowAgg → Sort → Gather` shape where workers ship rows to the leader and the leader sorts the whole corpus (confirmed by `EXPLAIN ANALYZE`: `Gather (actual rows=200000, loops=5)` with a 4.3 MB external disk merge, versus per-worker top-N heapsort in 31 kB with a constant query vector). A real defect — but **not** the reason GT took 23 084 s, which the v2.7.6 notes over-credited it for. Restructuring measures **1.00× narrow / 1.25× on 3 KB-wide rows**, and 6.4 h over 100M comparisons is 231 µs each against a 154 GFLOP job: the cost is per-row overhead across five full sorts of a 1 M-row wide table on a pre-AVX2 host, not the plan. Left unapplied pending validation at 1 M scale on an AVX2 host. Storage *did* confirm the `dim/8` stride and an exact 2.00× 1-bit:2-bit ratio at 1 M rows. Note a synthetic corpus can pass `is_degenerate()` and still be unrankable — different checks, and only the second predicts whether recall means anything. ### Harness: `--run-id` so concurrent arms cannot corrupt each other Two arms running in one database silently corrupted each other three ways: a shared `bq_query_set` (one arm's 256-d query set replaced another's 1024-d one mid-sweep), `DROP TABLE IF EXISTS bq_gt` destroying 860 s of ground truth, and a hardcoded `bqbench_` index prefix — index names are **schema**-scoped, so one arm's `CREATE INDEX` failed against a sibling's index on a *different table*, silently costing it an entire `bit_width` leg. All three are now namespaced by `--run-id` / `BQ_RUN_ID`, validated by the existing identifier guard. Default behaviour without `--run-id` is byte-identical. ### Stale documentation corrected - Test counts: `334/334` (AGENTS.md) and `341/346` in six rows of `PG_VERSION_SUPPORT.md` → the actual **427 passed / 8 ignored**, uniform across all seven CI legs pg13–19. - `AGENTS.md`: migration matrix was stuck at **v1.27.1** and the wire version said **7** when it has been **8** since v2.0.0. Both corrected, and ~160 lines of v1.x release prose that duplicated `CHANGELOG.md` replaced with current state (kinds, which to recommend, the corruption-history warning for anyone touching persist/scan code, and the upstream ctid bug). - `docs/PRODUCTION.md`: new section on bounding `maintenance_work_mem` for high-dim IVF builds. Measured: `bit_width = 4, lists = 512` at 1024-d reached **20.3 GB anon-RSS** and was **OOM-killed** at `maintenance_work_mem = 3GB`, then completed in 50.6 s at 1 GB with two parallel workers. The effective ceiling is `maintenance_work_mem × (1 + max_parallel_maintenance_workers)`, and an operator sees this as an unexplained backend termination. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.7.4';` — no REINDEX. ## [2.7.3] — 2026-09-08 **The 1-bit sign-BQ frontier is measured and published**, and a real insert-path bug found while running it is fixed. No wire-format change (stays v8), no SQL surface change, no REINDEX. ### Fixed: a 1-bit index built on an EMPTY table rejected its first insert `CREATE TABLE` → `CREATE INDEX ... WITH (bit_width = 1)` (while empty) → `INSERT` failed with `dim mismatch — index expects 0, row has 1024`. The empty build stamps `dim = 0` (no rows, no reloption-pinned dim) and `insert_bq_row` read the dim from the meta page. The flat path never had this bug because it takes the dim from the **incoming row**; the BQ path now does the same, seeding the mean as zeros for the 0-row case (a one-row corpus mean *is* that row, so every centred coordinate is 0 — exactly what a rebuild would produce). Found on a real 1024-d corpus host while setting up the benchmark, **not by a test**: every in-tree BQ fixture happened to index an already-populated table, so all of them missed it. The regression test covers the exact failing order and also asserts a wrong-dim row is *still* rejected once the dim is pinned, so the fix does not paper over dim checking. ### Measured: the 1-bit recall / storage / latency frontier Closes the last "not yet published" claim in the README. `arnold` (i9-12900H, **AVX2** — latency is only publishable on an AVX2 host per `AGENTS.md`), PostgreSQL 17.9, pg_turbovec 2.7.2, 250 000 × 1024-d Cohere-wiki as a native `turbovec.vector` column, **100 held-out queries** (zero corpus overlap, verified), exact top-100 in-DB ground truth using the *same* operator the index serves, postmaster and driver pinned to P-cores 2-5, and all 24 configs confirmed to run a real `Index Scan`. | `bit_width` | bytes/vector | vs 4-bit | R@10 ≥ 0.99 needs | p50 there | |---|---:|---:|---:|---:| | **1** | **142.3** | **3.98× smaller** | window **800** | 36.7 ms | | 2 | 280.5 | 2.02× smaller | window 32 | 6.0 ms | | 4 | 565.7 | 1.00× | window 32 | 9.2 ms | **1-bit trades latency for storage, steeply**: twice the storage saving of 2-bit, for 2.7–6.1× the latency at matched recall and a **25× wider** exact-rerank window to reach R@10 ≥ 0.99. That is a storage-constrained-workload option, not a default — which is how the reloption was already documented, now with numbers behind it. All four predictions registered *before* the run held. P2's falsification condition was "1-bit crosses at a comparable window, which would mean the `hi_dim_rerank` 1-bit special-case is unnecessary" — it did not, so the data **justifies** that special-case rather than merely tolerating it. Two further findings: - **1-bit degrades faster at depth than at k=10.** R@100 tops out at 0.981 (window 2000) and is 0.948 at the `auto` default, while 2-bit reaches 1.000 by window 800. Budget a wider window if you paginate or re-rank past the top 10. - **Harness cost trap, documented.** `--vec-expr` accepts any SQL expression, and a cast like `(emb::real[]::turbovec.vector)` is evaluated **per row per query** — 200 M casts for a 1M × 200-query ground-truth build, ~53 s *per query*. Materialising a native `turbovec.vector` column first took one query from **53 s → 2.7 s (20×)**. The 1M arm was abandoned for this reason, not for anything about turbovec. Caveats are stated in `docs/BQ_RECALL_BENCH.md` § 0.5 and not glossed: one corpus, one dim, one shared host (load ≈ 2 at start, so treat absolute milliseconds as indicative and the *ratios* as the result), and no IVF arm. This corpus also cannot separate 2-bit from 4-bit on recall — both sit at ≈1.000 nearly everywhere; it separates 1-bit from both, which is what it was run for. Artefact: `benches/results/bq_frontier_20260908/` (24 configs + sweep log). ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.7.3';` — no REINDEX. ## [2.7.2] — 2026-09-08 **BUG#6 is now reported upstream.** Documentation-only: the binary is byte-identical to 2.7.1 (the sole `src/` edit is a doc comment). No wire change, no SQL surface change, no REINDEX. The root-cause analysis and one-line core fix that v2.7.1 verified have been filed on pgsql-hackers: > **[ExecForceStoreHeapTuple() loses tts_tid, so ORDER BY-op index scans > project an invalid ctid](https://www.postgresql.org/message-id/0498c10f-839b-4f68-9994-c29b454e55a4%40app.fastmail.com)** > — 2026-09-08, with the patch attached as `v1-0001-...`. The filed patch's `execTuples.c` hunk is **identical** to the one verified here by A/B build. The filed version additionally adds a **core regression test** to `src/test/regress/{sql,expected}/gist.*`, and that test was itself checked against both builds: it reports `ctid_matches = 1` on unpatched 18.4 and `5` on the patched build, so it genuinely gates the fix rather than passing vacuously. `docs/FILTERING.md` and the `knn_scan_ctid_projection_upstream_limitation` test now carry the thread link. That test still asserts the **current (broken)** behaviour, so it will **fail once a fixed PostgreSQL reaches CI** — which is the intended signal to flip the assertion (gated on the fixing version) and relax the docs, not a regression in this AM. Until a fixed PostgreSQL ships nothing changes for users, and the workaround remains a proven necessity rather than a preference: `MATERIALIZED` CTEs, text casts inside a subquery, and extra subquery nesting were all tested and all still yield the sentinel. Chain on your own key column, or use `turbovec.knn()`. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.7.2';` — no REINDEX. ## [2.7.1] — 2026-09-08 **BUG#6 root cause proven against stock PostgreSQL, and the one-line core fix verified.** Documentation + upstream-patch release: the binary is byte-identical to 2.7.0 (no wire change, no SQL surface change, no REINDEX). The only `src/` edit is an expanded doc comment on the existing tripwire test. BUG#6 is the one where `SELECT ctid ... ORDER BY emb <=> q LIMIT k` projects the invalid-item-pointer sentinel `(4294967295,0)`. It was previously *argued* to be a core bug; it is now **demonstrated**: - **Reproduced on stock PostgreSQL 18.4 using only core GiST, with zero turbovec loaded** — thin diagonal polygons, where the bounding-box distance under-estimates the true polygon distance, so `was_exact` comes out false and tuples are routed through `nodeIndexscan.c`'s reorder queue. - **Traced line by line in stock source.** `indexam.c:983` sets `xs_recheckorderby` for any AM that asks; `nodeIndexscan.c:290` queues the tuple; `reorderqueue_pop` calls `ExecForceStoreHeapTuple`, whose `TTS_IS_BUFFERTUPLE` branch calls `ExecClearTuple` (which does `ItemPointerSetInvalid(&slot->tts_tid)`) and never restores `tts_tid` from `tuple->t_self`; `tuptable.h:420`'s `slot_getsysattr` returns `&slot->tts_tid` for the ctid system column. The sibling `tts_heap_store_tuple` **does** set `tts_tid`, so it is an asymmetry, not a design choice. - **The fix is proven, not proposed.** PostgreSQL built both ways on one machine, one script: | | unpatched | patched | |---|---:|---:| | ctid self-join, expect 5 | **1** | **5** | | `UPDATE ... WHERE ctid`, expect 5 rows | **1** | **5** | | sentinel ctids at `LIMIT 50` | **49/50** | **0/50** | The `UPDATE` line is the dangerous one: no error, it just affects one row instead of five. - **Byte-identical across branches**: the affected function body hashes the same in the 13.23, 14.22, 15.17, 16.14, 17.9 and 18.3 trees, so the one-line fix applies unchanged to every supported major. - **No query-level workaround exists** (verified, not assumed): `WITH ... AS MATERIALIZED`, casting to `text` inside a subquery, and extra subquery nesting *all* still return the sentinel, because it is already in the slot before any of them run. Only abandoning the index scan avoids it. So the documented guidance — chain on your own key column, or use `turbovec.knn()` — is the only real answer, and now it is known to be rather than assumed to be. **New finding.** turbovec sees this on **100%** of rows while core GiST loses only some. `IndexNextWithReorder` sets `was_exact = (cmp == 0)`, so a tuple *skips* the queue when the AM's advertised ORDER BY value compares exactly equal to the recomputed one. turbovec advertises `f64::NEG_INFINITY`, which never compares equal to a real distance, so every one of its tuples is queued. That is a difference in **exposure, not in cause** — advertising a real lower bound would only make the fault intermittent, which is harder to diagnose, not safer, so the `-inf` choice stands. Ships `docs/upstream/bug6-pgsql-hackers-DRAFT.md`, a submission prepared for review and **not sent**, alongside the existing patch file (now carrying the A/B table). ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.7.1';` — no REINDEX. Nothing in the extension's behaviour changes; the fix this release documents is a PostgreSQL core patch, not something pg_turbovec can apply. ## [2.7.0] — 2026-09-08 **IVF + 1-bit sign-BQ compose**, the Hamming kernel gets ~4.4× faster, and **two corruption-class bugs shipped by v2.6.0 are fixed**. No wire-format change (stays v8), no REINDEX, no SQL surface change. ### Two bugs in v2.6.0's BQ code — fix these by upgrading Found by an audit while composing IVF with BQ, not by a field report: - **`MetaPageData::set_ivf_chains` omitted `bq_mean_count`** from its running chain-offset sum. An IVF+BQ build would have written the coarse-centroid chain **on top of the corpus mean** — destroying the centring vector that every sign code and every query depends on. This is the **fourth** occurrence of this bug class (v1.24.0 omitted `graph_count`; v2.6.0 found and fixed three sites). `set_graph_chain` had the identical omission (unreachable today, since graph + 1-bit is rejected) and is fixed too. Now guarded by an exhaustive pairwise no-overlap test, verified to fail when the omission is reintroduced. - **Flat-BQ `aminsert` did not re-persist the tombstone bitmap**, so every `INSERT` after a `VACUUM` silently **resurrected every deleted row** — the same M2 bug the graph kind fixed in v2.1.0, reintroduced because the BQ write path had no tombstone parameter. It also **appended unconditionally**, so re-inserting an existing heap TID added a second slot for the same row (unbounded growth under upserts, plus a duplicate id — the shape the flat bijection guard calls corruption). Both fixed; the write path now carries tombstones through a single meta write. Only `bit_width = 1` indexes were affected, and only on insert-after-VACUUM (resurrection) or re-insert (duplicate slots). 2/3/4-bit indexes were never affected. There is no on-disk format change, so upgrading is sufficient — but an existing 1-bit index that has taken inserts after a VACUUM should be `REINDEX`ed, since it may hold resurrected or duplicated slots. ### `WITH (lists = N, bit_width = 1)` Previously rejected with a clear ERROR. Now builds and scans: sign codes are stored **cell-contiguous** (the permutation IVF already applies) alongside the coarse centroids and cell directory, so a scan probes only the nearest `turbovec.probes` cells and runs Hamming within them — the `dim/8` storage win combined with the probe-a-fraction scan win. Wire version stays 8: the shape is `kind = KIND_BQ` plus the existing v4 IVF chain fields. **BQ cells live in the raw L2-normalised space, not the rotated space TurboQuant IVF uses.** TurboQuant trains cells in the rotated space because that is where its fine quantizer encodes; a BQ code is the sign of the *centred raw* coordinate, so there is nothing to align to. Getting this asymmetric is the sharpest available silent failure — a rotated query against un-rotated centroids probes the wrong cells and collapses recall with no error — so the scan skips the rotation explicitly and the tests assert per-id self-neighbour recovery. An IVF+BQ `aminsert` **degrades the index to a flat BQ scan** (it appends and drops cell metadata). This is *observable*: `lists` is preserved and `ivf_degraded` is stamped, so `turbovec.index_is_degraded()` reports it — strictly better than the TurboQuant IVF insert path, which blanks `lists` and cannot be reported. Rebuild with `REINDEX` to restore cell-scoped scanning. ### Hamming kernel: ~4.4× faster, and AVX2 measured then declined The kernel now counts 8 bytes at a time instead of one. **Measured 4.4–4.8× at embedding dims** (100k×768d: 6.13 ms → 1.28 ms; 1M×768d: 59.0 ms → 13.7 ms; independently reproduced at 4.5–5.4× on a second machine). Stable from L3-resident to firmly-DRAM sizes, so the scan is not bandwidth-bound at these sizes. **No `unsafe`, no runtime CPU dispatch** — every machine and architecture runs the identical instruction sequence. An AVX2 intrinsics kernel *was* written and **proven bit-identical**, then **declined on measurements**: it buys only ~1.8× more above `dim = 512` and is a net **loss** below it (the 32-byte loop never runs at `dim ≤ 256`), while costing a second `unsafe` block on the scan hot path, a dim-dependent dispatch threshold, and a code path CI cannot exercise both sides of — the exact blind spot that let the v1.7.3 wrong-results bug ship. `u64x4` was also measured and rejected (LLVM already extracts the ILP). Full numbers in `docs/ONEBIT_BQ.md` §7. Bit-identity is proven in-tree, not assumed: **5720 random code pairs across 143 dims** against a bit-by-bit reference sharing no code with the kernel, **825 top-k cases** asserting the full sequence *including tie order*, a tie-density guard so that is not vacuous, and mutation testing (dropped tail, byte-swapped operand, `AND` for `XOR`, skipped last word, `count_zeros`) where each mutation fails 3–6 tests. ### Benchmark harness for the unpublished BQ frontier `benches/scripts/bq/` plus `docs/BQ_RECALL_BENCH.md`: a sweep over `bit_width` ∈ {1, 2, 4} measuring R@10 against exact in-DB ground truth, storage, build time and — **only on an AVX2 host** — latency. **It has not been run; no numbers exist yet.** The AVX2 gate is structural: on a scalar host the driver does not time queries at all, rather than emitting a number that reads as fast. Ground truth uses the same operator the index serves (avoiding the cosine-vs-L2 trap), the re-rank window is recorded per row against a mirror of the Rust clamp, and iso-recall rows emit `null` when a `bit_width` never clears the target — that absence being the result. Predictions are written down in advance so a real run can falsify them. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.7.0';` — no REINDEX for 2/3/4-bit indexes. **A 1-bit index that has taken inserts after a VACUUM should be `REINDEX`ed** (see the bug notes above). ## [2.6.0] — 2026-09-07 **1-bit sign binary quantization (`WITH (bit_width = 1)`) works end to end** — build, scan, `aminsert`, VACUUM. **No wire-format change (stays v8), no REINDEX, no SQL surface change.** Previously the reloption was accepted for forward compatibility but the build raised a clear "not yet implemented" ERROR. 1-bit is **not** TurboQuant-at-1-bit: the turbovec crate hard-rejects `bit_width < 2` in both constructors, so this is a distinct scheme — sign binary quantization, the DiskANN/pgvector/Qdrant coarse code. Per-vector storage is `dim/8` with **no** per-vector scale (half the 2-bit stride, and it also drops the codebook, the rotation/TQ+ chain and the blocked chain), scored by Hamming (`popcount(XOR)`) and then **exactly reranked** by the AM's existing `xs_recheckorderby` machinery. ### Why there is no wire bump A new **`kind` byte** (`KIND_BQ = 3`), not a version bump. Every existing index keeps `kind = SINGLE/COLBERT/GRAPH` and decodes byte-identically — the same additive per-kind path v4→v5→v6 used. The three `bq_mean_*` meta fields occupy page offset 316, reserved-and-zero on every prior version, so an old meta page reads them as "absent". (`docs/ONEBIT_BQ.md` §4 originally specified a 7→8 bump; that was written when v7 was current.) ### Mean-centering is load-bearing The naive sign-at-zero rule sets every bit on dense-positive data (measured R@10 = 0.0 on GIST), so the per-dim corpus mean is subtracted before taking signs, **persisted**, and applied to queries too. A corpus still collapsed *after* centering (constant / near-constant) is rejected at build rather than shipping a signal-free index. Scanning an index whose mean is missing ERRORs with a REINDEX hint rather than returning garbage. ### Three instances of the v1.24.0 corruption class, found and fixed `write_tombstones_and_meta`, the tombstone placement inside the rewrite path, and `MetaPageData::total_blocks()` each summed chain page counts **without** `bq_mean_count`. On a BQ index that would have placed the tombstone chain **on top of the mean vector** and under-sized the relation — the identical shape of the v1.24.0 graph bug (which omitted `graph_count`). Found by auditing every chain-offset sum in the tree, not only the path being added. ### Notes - `aminsert` does **not** recompute the mean: that would invalidate every sign code already packed against the old mean, silently degrading the whole index's ranking from one insert. Build-time mean is fixed; drift is a REINDEX concern. The insert is one exclusive lock across read-modify-write — an unlocked RMW is exactly what silently lost graph rows before v2.1.0. - VACUUM is tombstone-only, sharing the graph kind's path (compacting would mean renumbering every slot). - `turbovec_check` reports `kind = 'bq'` and skips the v2.2.2 scales validation for that kind only — BQ has no scales chain, so validating one would report every BQ index corrupt. - The Hamming kernel is deliberately **scalar** (`count_ones` lowers to `POPCNT`); any hand-vectorised version must first be proven bit-identical against it. This is the v1.7.3 lesson — a mis-specialised kernel returned *wrong* ANN results on pre-AVX2 CPUs — applied pre-emptively. ### Not yet supported `bit_width = 1` with `lists > 0` (IVF) is rejected with a clear ERROR; the cell-contiguous sign layout and per-cell Hamming scan are unwired. `bit_width = 1` with `graph = true` stays rejected in the reloption validator (and the graph kind is deprecated as of 2.5.0). BQ's recall/latency frontier on real corpora still needs an AVX2-host run; the tests prove correctness, storage and scan behaviour, not a published frontier. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.6.0';` — no REINDEX. ## [2.5.0] — 2026-09-07 Two changes: the graph kind is **deprecated**, and **Phase S-1 partition pruning** lands (additive SQL). No index wire-format change (stays v8), no REINDEX. ### `WITH (graph = true)` is deprecated It now emits a deprecation `WARNING`; the build path is scheduled for removal, with decode retained one further release so a stale graph index fails loudly with a REINDEX hint rather than silently. The kind was added in v1.23.0 to chase HNSW's query latency while keeping TurboQuant's storage compression. Measured at **matched recall** it never delivers. Its apparent sublinearity holds only at iso-**beam** (p50 1.11× for a 10× corpus — but recall falls 0.605 → 0.472); once recall is held equal the curves **diverge and never cross**: | corpus / target | flat | IVF | graph | |----------------------------|---------------------|----------------------|--------------| | SIFT-1M/128d, R@10 ≥0.95 | 0.98 ms, qps@8 1380 | 1.8 ms, qps@8 2039 | 26.2 ms, qps@8 299 | | GIST-1M/960d, R@10 ≥0.95 | 5.88 ms, qps@8 279 | 11.3 ms, qps@8 480 | **unreachable** | | GIST-10M/960d, R@10 ≥0.98 | 34.2 ms, qps@8 31 | 28.4 ms, qps@8 161 | **unreachable** | It also loses on build time (57–90×), storage, and has no out-of-core path. **It is IVF, not the graph, that beats flat's O(n) wall.** Use the default flat index below ~1M vectors and `WITH (lists = N)` at scale. Retained deliberately: `turbovec.graph_ef`, `pack::repack`, and the `coarse_graph` work (Phase G-1) — that one navigates *centroids*, and IVF's win partly rests on it. Deprecating the graph *kind* is not deprecating graph *techniques*. Also recorded: the "60× parallel build speedup" was an artefact. `graph_build_partitions_decide` coupled shard count to thread count, and shards cost recall (GIST-1M R@10 0.920 at P=4 → 0.605 at P=83); threads at a recall-preserving P buy <5×. The partitioned-build parity test is annotated with the blind spot that hid this (2.5k rows/shard at dim 64, versus ~12k at 960d in reality) rather than re-tuned, since the kind is on its way out. ### Phase S-1: partition-level coarse quantizer At 1T scale the design is a partitioned parent with one turbovec index per partition and native `Merge Append` doing scatter → gather. That is correct for any N but the per-query fan-out is `O(N)` in the partition count: at ~100k partitions, opening an `Index Scan` per partition dominates. S-1 lifts IVF one tier up — each partition gets a **summary** (its mean in the **original**, un-rotated space, so summaries are comparable across partitions; each partition trains its own rotation, so the persisted coarse centroids are *not*) — and `nearest_partitions()` returns the `Kp` nearest so the caller fans out to `Kp` instead of `N`. New SQL (additive): table `turbovec.partition_summary`, plus `turbovec.refresh_partition_summary(parent, vec_col)` and `turbovec.nearest_partitions(parent, query, k_partitions, metric)`. S-1 was first attempted in v1.29.0 and reverted when its `#[pg_test]` failed CI. **The assertion was wrong, not the code**: it demanded the pruned top-k be identical to the full-fan-out top-k, but pruning is an approximation by construction — scoring a query against partition *means* cannot be equivalent to scanning every partition, so on unclustered data a `Kp < N` probe can legitimately miss a true top-k member. The revived test asserts what pruning actually guarantees: exactly `Kp` partitions returned, the query's own cluster ranked first on content-clustered data, and a **recall floor** against full fan-out. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.5.0';` — no REINDEX. Existing graph indexes keep working and will warn on rebuild. ## [2.4.0] — 2026-09-07 **WAL amplification follow-up: stable chain starts.** No wire-format change (stays v8), no SQL surface change, no REINDEX. Minor rather than patch because the on-disk *allocation* layout of new/rewritten indexes changes (chain contents and decoding do not). v2.3.0 stopped WAL-logging pages whose contents hadn't changed, which cut the reported ~500 MB-per-commit to ~28 MB. This release addresses what was left. Chain starts were packed back-to-back — `scales_first = codes_first + codes_count`, `ids_first = scales_first + scales_count` — so **any** growth in the codes chain moved every scales and ids page to a new block number, and a relocated page is genuinely different, so it had to be rewritten. On the reporter's 768d/4-bit index a codes page holds only **21 rows**, so essentially every flush crossed an allocation boundary and paid ~27.6 MiB to move the scales+ids chains. - The three growing chains' allocations are now rounded up to a multiple of `MetaPageData::PAD_PAGES` (256 pages), so a shift happens once per ~5400 rows instead of once per 21. Modelled on that index, WAL per row falls from ~55 → ~5.7 KiB at batch=512, and ~1375 → ~37 KiB at batch=1, on top of v2.3.0's own ~33×. - Cost is **bounded slack**: at most `3 * PAD_PAGES` pages (6 MiB) per index, independent of index size. Chains smaller than one padding unit are **not** padded, so small indexes keep the previous layout byte-for-byte — without that gate a 1000-row 768d index would have gone 0.4 → 6.0 MiB (15×) to buy an amortisation it never needs. - `padded_pages_needed()` is the single definition, used by the planner *and* the incremental grow/shrink paths in `relfile.rs`. If one site padded and another didn't they would disagree about where the following chains start — precisely the class of block-offset bug that silently corrupted a graph index in v1.24.0. Why this is safe without a REINDEX: `*_count` keeps its existing meaning of "blocks **allocated** to this chain", which is what every consumer already uses it for (sizing the relation, and locating the next chain via running sums). A chain's **contents** are located by `*_first` plus `n_vectors`/`rows_per_page` and never by `*_count` (see `read_chain`), so a padded index decodes identically. Existing unpadded indexes keep working as-is; their chains relocate on the first full rewrite, which is safe by the v1.29.4 invariant (all chains are written **before** the meta page, so an interrupted rewrite leaves the old meta pointing at the old, intact chains). Four new layout tests cover the invariants directly: chain starts don't move when up to 10k rows are added; padding never changes content addressing; small and empty indexes aren't padded; and across a dim × bit_width × n sweep every padded extent is large enough for its rows with overhead bounded by one padding unit. `docs/TESTING.md` gains a section on measuring global counters, distilled from the three self-inflicted CI failures in v2.3.0 — concurrent `#[pg_test]`s share one cluster, so `pg_current_wal_lsn()` deltas capture other tests' WAL (the same operation measured 32 KB / 1097 KB / 1425 KB across runs and once ranked the arms inverted). Count a backend-local value you control, reset each arm to an identical starting state, and assert the control arm is non-trivial so "0 vs 0" can't pass. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.4.0';` — no REINDEX. ## [2.3.0] — 2026-09-07 **WAL amplification fix: a flush now only WAL-logs the index pages that actually changed.** No wire-format change (stays v8), no SQL surface change, no REINDEX. Minor rather than patch because insert-path write behaviour changes materially. A field report (2026-09-08, pg.ddx.io) measured pg_turbovec inserts accounting for **~100 % of all WAL** on the host: **1322 MB / 25 s** with an embedding backfill running versus **31 kB / 25 s** with that one unit stopped — a **~42,000×** difference from a single writer doing ~100–1400 rows/minute. The decisive detail: `pg_stat_statements` attributed only 434 kB of a ~3 GB/60 s window to SQL, so **>99.98 % of the WAL was generated outside any statement** — index maintenance, not the `INSERT`. WAL per *commit* was roughly **constant and close to the 882 MB index size** (~500 MB at 16 rows/txn, ~754 MB at 128, ~670 MB at 512) and WAL per *row* a clean `1/batch` curve with no knee — the signature of the whole relfile being rewritten and fully WAL-logged on every flush. A scratch-index A/B isolated it to commit boundaries alone: 32 × 16 rows = **643 MB** WAL versus 1 × 512 rows = **2.6 MB** (245×), both *after* `wal_compression=zstd`. Cost: **~4.3 TB WAL/day**, archived off-host so paid for twice, and the dominant consumer of the host NVMe's endurance (49 % used, 172 TB written in 5306 power-on hours). It also kept `num_requested` checkpoints at ~2× `num_timed` with a 220–390 % FPI ratio, which sent the operator chasing a checkpoint misconfiguration. Cause: `write_chain_at` registered **every** page of **every** chain with `GENERIC_XLOG_FULL_IMAGE`. But `reconcile_flush_image` appends new slots at the **end** and updates touched slots in place, so every other page was already byte-identical on disk — we were paying a full-page image to rewrite pages with their own contents. - Each full page is now compared against the bytes about to be written (under a shared buffer lock) and skipped when identical, so WAL scales with bytes **changed** rather than index size. Only pages already carrying our own no-hole header are eligible, so a page whose header `GenericXLogFinish` would treat as a hole is never inherited; a skipped page is by definition already correct, so crash recovery is unaffected. - The `GenericXLog` state is started **lazily**, so a batch in which every page is skipped emits no WAL record at all instead of an empty one. - Removed the dead, never-called `write_chain()` helper. It wrote pages with `MarkBufferDirty` and **no WAL** — a latent footgun that would have made a skipped page's missing WAL permanent. Every page mutation now provably goes through `GenericXLog` (zero live `MarkBufferDirty` sites). - Test `insert_wal_scales_with_change_not_index_size` drives a real flush (via the existing `flush_to_relfile_for_test` hook — a `#[pg_test]`'s outer transaction always rolls back before PreCommit fires, so the aminsert path can't be observed in-band) on a ~109-chain-page index and asserts a no-change flush WAL-logs <1/10th the pages of a full rewrite. Fail-before/pass-after via `SKIP_UNCHANGED_PAGES`. It counts pages registered for WAL rather than diffing `pg_current_wal_lsn()`, which is load-bearing: `#[pg_test]`s run concurrently against one cluster, so a global LSN delta also captures every other test's WAL. Three CI runs reported the same operation as 32 KB, then 1097 KB, then 1425 KB, and once ranked the two arms *inverted* — the LSN measurement was never valid in either direction. Registered pages are per-backend, deterministic, and are the direct driver of WAL volume (one full-page image each). Note the residual cost on an *append* (as opposed to a no-change flush): chain starts are packed back-to-back (`scales_first = codes_first + codes_count`), so when the codes chain grows by one page every scales and ids page lands at a new block number and genuinely must be rewritten. Measured 392 KB vs 1416 KB always-rewriting on a 20k-vector/64d index. Making chain starts stable across growth would shrink that further and is left as follow-up. **Batching still matters** and the docs now say so: WAL scales with commits, so a flush's remaining cost (tail pages, meta page, IVF cell directory) is paid per transaction. `docs/PRODUCTION.md` gains a "WAL cost of inserts — batch your writes" section including how to measure it with LSN deltas, since `pg_stat_statements` will not show it. Also documents the **partial-index verification gotcha** (Ask 3 of the 2026-09-05 report): a KNN query that omits a partial index's predicate silently gets a sequential scan, which masks index faults as latency problems — always confirm `EXPLAIN` shows `Index Scan using `. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.3.0';` — no REINDEX. ## [2.2.2] — 2026-09-05 **`turbovec_check()` no longer has a scan-fatal blind spot.** Code-only — no wire-format change (stays v8), no SQL surface change (the `reason` column already exists as of 2.1.0), no REINDEX required to upgrade. A field report (2026-09-05, pg.ddx.io) found an IVF index — ~2.18 M vectors, 768 d, `bit_width = 4`, `lists = 1400` — on which **every** KNN scan failed instantly with turbovec's `InvalidScaleValue { slot: 1, value: -2.559434e22 }`, while `turbovec_check()` reported `is_corrupt = false`, `count_matches = t`, `duplicate_id = NULL`. Because the operator's automated self-heal trusts that check, the index degraded **silently** until a manual `REINDEX` (~12 min) cleared it. The reporter correctly identified this as the more important of the two issues they filed: a checker that cannot detect an index failing 100 % of scans is a false-confidence generator. Root cause of the blind spot (a deliberate, documented decision that this incident falsifies): the check read only the meta page and the ids chain, skipping "the much larger codes/scales chains". But the **scales chain is the cheap one** — one `f32` per vector, **8.7 MB** at 2.18 M vectors, versus the **17.4 MB** ids chain it *already* read, versus **0.84 GB** for the codes chain it rightly still skips — and it is load-bearing for `TurboQuantIndex::from_parts`, i.e. exactly the gate a scan hits. - `turbovec_check()` now validates every persisted scale (finite, non-negative, sane magnitude — the same invariants `from_parts` enforces) for **every** index kind, inside the *same* shared rewrite-lock bracket as the meta/ids read so all three are one consistent observation, and names the failing slot in `reason`. Cost is +50 % on an already-cheap query; no opt-in `deep` flag, because a default that cannot see a scan-fatal fault is the bug. - The scan-path rejection is now a proper PostgreSQL `ERROR` naming the fault and the **`REINDEX INDEX`** recovery, instead of the bare Rust `.expect("...from_parts rejected raw parts")` string an operator used to get with no hint that REINDEX fixes it. - Regression test `turbovec_check_detects_invalid_scale` writes the reported value (`-2.559434e22`) into the *persisted* scales chain of an IVF index via a page-straddle-safe test-only helper and asserts the checker flags it; fail-before/pass-after. Note this is the same failure *shape* as the graph lost-update called out in the 2.1.0 notes: **internal self-consistency of meta+ids is not evidence of scannability.** The originating corruption itself (a damaged scales chain on an index that had survived several `pg_cancel_backend`/`pg_terminate_backend` events, a `kill -QUIT` immediate shutdown and a restart that aborted an in-flight backfill) is not yet reproduced and remains under investigation; this release makes it *detectable* and *actionable* rather than silent. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.2.2';` — no REINDEX to upgrade. (An index already damaged needs a one-time `REINDEX INDEX ;`, which this release will now actually tell you about.) ## [2.2.1] — 2026-09-05 **Safety patch: the PARALLEL graph build was effectively uncancellable.** Code-only — no wire-format change (stays v8), no SQL surface change, no REINDEX. v2.1.0 made the graph build interruptible, but its hook lives in thread-local storage and is therefore invisible to rayon workers *by design* (a worker must never `longjmp` out of the pool). That left a hole nobody had measured: during the partitioned build the driver parks in a futex inside rayon's `join` for the entire parallel phase, so it cannot reach a poll either. Measured on a 10M-node build: `pg_cancel_backend()` **and** a direct `SIGINT` were both ignored for **over 13 minutes** while 32 threads ran at 100% CPU — and because no backend code was executing, `pg_stat_activity` reported `wait_event = NULL`, so the runaway build looked *idle*. An operator had no way to stop it. Workers now consult a cheap, thread-safe **abort predicate** at their safe points and stop producing work; the driver raises PostgreSQL's real cancel/terminate error at its own poll once the phase collapses, so all error *raising* still happens only on the backend thread. The predicate is **injected** by the PG-side caller, so `src/index/graph.rs` stays Postgres-free and the mechanism is unit-testable without a server (`partitioned_build_bails_out_when_abort_is_requested`). Found while measuring whether the graph kind could be made production-grade at scale (see the 2.2.1 note in `docs/UPGRADING.md` and the deprecation discussion for v2.3.0). ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.2.1';` — no REINDEX. ## [2.2.0] — 2026-09-04 **Graph-kind scan-beam retune.** The graph kind's scan-time beam width is now its own knob (`turbovec.graph_ef`, default auto = 512) instead of a side effect of the candidate count, and the auto default is set from a measured 1M-scale recall-vs-latency frontier on SIFT-128 and GIST-960. **No wire-format change (stays v8), NO REINDEX** — but a new GUC is a SQL-surface addition, so this is a MINOR. ### Fixed — the graph "high-dim recall cliff" was a BEAM bug, not a scorer bug The tracked "GraphScorer diverges from the kernel at `dim >= 512`" is **disproved**. An exact-f32-oracle control gives recall identical to the LUT scorer (within 0.015), and recall recovers monotonically with the BEAM on the same index and the same scorer (`search_k` 32 → 128 → 512 gives R@10 0.750 → 0.990 → 1.000). The real mechanism was the scan-time beam width, and — worse — **which knob was setting it**: - Through v2.1.0, `graph_search` computed `ef = (k * 4).max(64)` where `k` was the candidate count `scan.rs` had already widened. Since `turbovec.hi_dim_rerank = auto` raises that count to `clamp(dim, 256..=1024)` for `dim >= 256` — a knob whose actual job is the **flat/IVF exact-rerank window** over a cell scan's quantized ranking — the graph's beam was being set as a side effect of an unrelated feature. - The visible symptom was an **inverted** cliff. With `hi_dim_rerank = auto`, 128d was the WORST (R@10 0.720: 128d is below the rerank threshold, so the beam stayed at the bare 64 floor) while 384/512d looked perfect (0.99/1.00: the inflated candidate count multiplied the beam up). With `hi_dim_rerank = off` recall collapsed at EVERY dim (~0.40/0.38). Recall was never dim-dependent; it was beam-dependent, and the beam was accidental. Fixed by splitting `graph_search` into `graph_search_with_ef(.., ef, ..)` (the live scan path, beam resolved from the new GUC) and the original `graph_search` (the pure-`k` fallback for callers with no live GUC — unit tests, the `aminsert` findability probe). `scan.rs`'s graph arm now passes the user's OWN candidate count (`search_k * oversample`), not `hi_dim_rerank`'s floor. `hi_dim_rerank` keeps doing its real job for flat/IVF, unchanged. Verified on 42 paired `(k, ef)` configs across SIFT-200k, SIFT-1M and GIST-1M: `off` and `auto` now give recall within **0.0000** of each other at every beam, at 128d AND at 960d. ### Added - **`turbovec.graph_ef`** (int, `Userset`, default `0` = auto = 512, range `0..=1000000`) — the graph kind's recall/latency dial, the direct analogue of `hnsw.ef_search` (and of `turbovec.probes` for IVF). `0` = auto; a positive `N` pins the beam. Always clamped up to the query's candidate count (a beam narrower than it cannot fill it) and down to the live corpus size. Pure scan-time knob: no wire change, no REINDEX, honoured immediately by any graph index built by any 2.x binary. ### Changed — defaults - **Graph scan beam default 64 → 512** (`GRAPH_SCAN_EF_DEFAULT`), and it is now ABSOLUTE rather than `(k * 4).max(64)`. Measured frontier (32 vCPU AVX-512, 1M rows, R@10 vs exact cosine GT, warm p50, `search_k` pinned to `max(32, k)`): | corpus | beam 64 | beam 512 | beam 2048 | |---|---|---|---| | SIFT-1M (128d) | R@10 0.943 / 1.25 ms | **R@10 0.990 / 5.19 ms** | R@10 0.991 / 14.13 ms | | GIST-1M (960d) | R@10 0.760 / 8.65 ms | **R@10 0.920 / 34.55 ms** | R@10 0.966 / 93.44 ms | At 128d 512 is the unambiguous knee: it is within 0.001 R@10 of the 2048-beam ceiling at 0.37× its p50, and doubling past it buys +0.001 for 1.7× the latency. At 960d the curve has NOT plateaued by 2048, so 512 is a deliberate latency cap — the widest beam that keeps 960d p50 under ~35 ms — not a knee; `SET turbovec.graph_ef = 2048` reaches 0.966 for ~95 ms if that is the trade you want. ⚠️ **One configuration REGRESSES on recall, deliberately.** A 960d graph index queried with `hi_dim_rerank = auto` (the default) goes R@10 0.971 → 0.920 and R@100 0.958 → 0.826, in exchange for p50 194 ms → 34 ms (5.6×) and 196 ms → 42 ms (4.7×). The old numbers came from the side-effect beam of **3840** that `auto` was silently buying at 960d — never a documented beam, and it evaporated the moment a user set `hi_dim_rerank = off` (0.840, the "collapse"). A ~200 ms p50 is not a defensible default, and the old recall is now REACHABLE and documented for the first time: `SET turbovec.graph_ef = 3840` reproduces the pre-patch `auto` beam exactly. 128d indexes improve in BOTH modes (R@10 0.976 → 0.990 at 1M) and lose nothing — 128d is below `hi_dim_rerank`'s `dim >= 256` threshold, so 128d `auto` was never above `off`. ### Known limitation — the graph kind is NOT the fastest kind at 1M Measured on the same box, same corpora, same GT, at k=10: the **flat** kind beats the graph on BOTH recall and latency on BOTH corpora. SIFT-1M flat R@10 0.993 / 0.96 ms vs the graph's best-defensible 0.990 / 5.19 ms (5.4× lower latency, higher recall). GIST-1M flat R@10 0.997 / 5.91 ms vs the graph's 0.920 / 34.55 ms (5.8× lower latency, +0.077 recall) — and even beam 2048 (0.966 / 93 ms) does not catch flat. At 1M rows on a 32-core AVX-512 host, turbovec's SIMD full scan (one linear sweep of a compact quantized buffer at memory bandwidth) is simply faster than navigating a graph (~`ef` scattered gathers with a serial dependency between hops). The graph's advantage is asymptotic and 1M is below the crossover on this hardware. **Guidance: use flat (or IVF for a smaller RAM footprint) at this scale.** The graph's one measured win is CONCURRENCY scaling — the flat scan saturates memory bandwidth and goes 713 → 1313 qps from 1 to 8 clients (1.8×) while the graph goes 688 → 4906 (7.1×) — so it is the better throughput engine under load even where it loses isolated p50. Full curves and the reasoning in `docs/GRAPH_EF_BENCH.md`. ### Docs - `docs/GRAPH_EF_BENCH.md` — the full measured frontier (both corpora, both `hi_dim_rerank` settings, the whole `ef` grid, at `k` = 10 and 100), the flat/IVF comparators at the same `k`, the before/after A/B on byte-identical index files, and the reasoning behind 512. - `docs/PRODUCTION.md` — `turbovec.graph_ef` reference section. - `README.md` — GUC table + count (18 → 19). ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '2.2.0';` and restart the backend (the new GUC is registered in `_PG_init`). **No REINDEX**: the wire format is unchanged (v8) and the beam is resolved at scan time, so existing graph indexes get the new default immediately. Flat and IVF indexes are unaffected — the beam does not exist on those paths. To restore a pre-v2.2.0 effective beam exactly: `SET turbovec.graph_ef = 128` for what `hi_dim_rerank = off` gave at any dim (and what `auto` gave below 256d), or `= 3840` for what `auto` gave at 960d. ## [2.1.0] — 2026-09-04 **Graph-kind correctness release.** The `WITH (graph = true)` kind had a critical concurrent-INSERT data-loss bug, could return fewer than `k` rows, and its build could not be cancelled. All three are fixed here, plus graph health monitoring and an upstream-bug writeup. **No wire-format change (stays v8), NO REINDEX** — but `turbovec_check()`'s return type changes, so this is a MINOR (see the upgrade note). ### Fixed — graph kind - **C1 (CRITICAL): concurrent `INSERT` into a graph index lost data.** `insert_graph_row` did an *unlocked* whole-index read, then a *blind* whole-relfile rewrite. Two concurrent graph inserters produced either (a) a torn adjacency chain — `corrupt graph adjacency chain: graph offsets[n]=54272 != neighbors.len()=54240`, a loud ERROR — or, worse, (b) a **silent lost update**: the losing inserter's rows were simply gone (heap rows committed with no index entry) while `turbovec_check()` still reported `is_corrupt = false`, because the surviving relfile was internally self-consistent. Measured fail-before: 8 of 100 concurrent inserts lost, 8 heap rows unindexed. Fixed by holding ONE `lock_relfile_write` across the entire read-modify-write, re-reading the meta page under it, and adding guards that ABORT rather than reconcile onto torn state. Also folded in the tombstone two-write gap: the bitmap is now planned into the **single** meta write (previously a second write re-attached it, and a crash in that window permanently un-tombstoned every VACUUM delete). Validated: 320/320 concurrent inserts clean, and a sustained 240 s run of 6 inserters against a concurrent DELETE/VACUUM loop stayed `is_corrupt = false` with zero SIGABRT. - **BUG#2: `graph_search` could return fewer than `k` rows.** Two real causes: tombstoned nodes are never routing hops, so a large dead set disconnects the live remnant (n=2000 at 99% dead returned 2 rows for `k=10`); and a degenerate corpus of tied distances collapses `RobustPrune`'s out-lists (500 identical rows returned 8 for `k=10`). Fixed with a bounded backfill from unvisited live slots, gated behind the existing `out.len() >= k` guard so the normal path is byte-identical. - **FINDING#2: a graph `CREATE INDEX` could not be cancelled** and hung `pg_ctl stop -m fast` (the rayon build had zero interrupt polling). Added driver-thread stage polls mirroring `ivf_build_and_write`, plus a TLS-scoped hook the Postgres-free graph module calls at its own stage boundaries and every 256 nodes — a rayon worker cannot see the TLS slot, so it can never `longjmp` out of the pool. ### Documented — not a pg_turbovec bug - **BUG#6: a kNN index scan projects a sentinel `ctid`** `(4294967295,0)` instead of the real heap tid. Root-caused to **PostgreSQL core**, not this extension: `ExecForceStoreHeapTuple()`'s buffer-slot branch calls `ExecClearTuple()` (which invalidates `slot->tts_tid`) and never restores `tts_tid` from `tuple->t_self`, while its sibling `ExecStoreHeapTuple()` does — a plain asymmetry. `slot_getsysattr()` reads the projected `ctid` straight out of `tts_tid`. Any AM that sets `xs_recheckorderby = true` is affected; **core GiST reproduces it with no turbovec loaded**, and a one-line core patch fixes both. Row *data* is unaffected (`xs_heaptid` is correct), and the documented `turbovec.allowlist`-from-`ctid` recipe still works (it harvests from a *filter* scan). Only chaining off a *kNN scan's* `ctid` breaks. The proposed core patch is in `docs/upstream/`, the limitation and the supported alternatives are in `docs/FILTERING.md`, and a tripwire test fails the moment upstream fixes core so this note gets retired. ### Fixed — BUG#5 (monitoring, details) - **BUG#5: `turbovec_check()` was graph-adjacency-blind.** It validated the flat/IVF ids bijection + tombstones but never looked at a graph index's CSR adjacency chain, so a corrupt adjacency (a torn write, or the concurrent-insert corruption) reported `is_corrupt = false` — operators had no health signal for the graph kind at all. For a `kind = graph` index the check now decodes the adjacency chain and validates: offsets monotonically non-decreasing and consistent with the neighbor-array length, every neighbor id `< n_vectors`, the entry point in range AND having at least one out-neighbor, no self-loops, strictly-ascending (deduplicated) neighbor lists, the adjacency's node count matching `meta.n_vectors`, and at least one edge when `n_vectors > 1` (not trivially disconnected). Cost is `O(n + edges)` over the adjacency chain only — no vector data is read, so it stays cheap enough to poll. Flat/IVF/ColBERT behaviour is byte-identical (the new code is gated on `meta.is_graph()`). - The chain decode used by the monitoring path is now non-fatal (`relfile::try_read_graph_adjacency`), so a malformed chain is REPORTED rather than ERRORing out of the operator's health query. It also bounds the chain against the physical relation length before reading, so a truncated relfile is reported instead of raising `read_chain`'s hard past-EOF ERROR. The scan / insert / VACUUM paths keep the ERROR policy (they cannot proceed on a broken graph) and share the same decode implementation. ### Changed (SQL surface — additive column, but see the upgrade note) - `turbovec.turbovec_check(regclass)` gains a trailing **`reason text`** column: NULL when healthy, otherwise a human-readable description of the first problem found (duplicate id, count drift, or the specific graph-adjacency invariant violated). **Adding an OUT column changes the function's return type**, which `CREATE OR REPLACE FUNCTION` cannot do — the release that ships this needs a `DROP FUNCTION turbovec_check(oid); CREATE FUNCTION ...` pair in its `sql/pg_turbovec----.sql` upgrade script, and it is a MINOR bump at minimum (never a patch). No wire-format change (still v8); no REINDEX. ## [2.0.0] — 2026-08-26 **MAJOR: pg_turbovec now runs on upstream turbovec 1.0.0.** This adopts the stable turbovec 1.0.0 crate (Ryan Codrai's TurboQuant implementation, first stable release) in place of the long-lived 0.9.0 fork. **On-disk wire format v7 → v8** — a breaking change, with a documented REINDEX-from-heap migration (below). ### What changed - **turbovec 0.9.0 → 1.0.0.** The fork's hand-rolled centroids/boundaries codebook + QR rotation are replaced by turbovec 1.0.0's **TQ+ per-coordinate calibration** and **v5 block-Hadamard rotation** (which also drops the OpenBLAS build dependency). Every encoded byte differs from v7, hence the wire bump. pg_turbovec carries two small additive fork patches on top of stock 1.0.0 (`pub fn repack` and a re-exposed `IdMapIndex` parts API used by the buffer-manager cache-fill path), both tracked to be offered upstream. - **Wire format v7 → v8.** `MetaPageData::version = 8`, `EXPECTED_WIRE_FORMAT_VERSION = 8`. - **Materially faster, same storage.** On EC2 i4i.8xlarge (AVX-512), head-to-head vs v1.29.7: SIFT-1M flat p50 **12.9 ms → 1.3 ms (9.6×)**; GIST-1M IVF p50 @ R@10 0.95 **512 ms → 92 ms (5.6×)**; cold-scan (SIFT IVF) **3139 ms → 399 ms (7.9×)**; build 1.2–2.1× faster. **Index size byte-for-byte unchanged** (77 MB SIFT-1M, ~4.96 GB @ 10M). Recall matched-or-better at matched config. Determinism intact (the build-parallelism byte-identity gates all pass; v5 rotation removed the old OpenBLAS nondeterminism). Two minor warm-flat regressions (~11–15% at already-saturated recall), no correctness impact. - All v1.29.x corruption fixes (meta-LAST torn-write ordering, reconcile-on-flush, pre-flush validate-all, IVF `lists==0` dup-gate, VACUUM shrink guard, read_chain bounds, SubXact rollback) are re-proven to fire on the v8 persist path. The 90-minute no-VACUUM upsert+writer-restart corruption A/B ran clean on v8 (92 checks `is_corrupt=f`, 0 dup-id, 0 SIGABRT, 26 restarts). ### Migration — REQUIRES a one-time REINDEX per index (wire v7 → v8) `ALTER EXTENSION pg_turbovec UPDATE TO '2.0.0';` + restart the backend, **then `REINDEX INDEX ;` once per turbovec index.** A pre-v8 index opened under 2.0.0 is detected (`MetaPageData::is_legacy_v7()`) and `ambeginscan` ERRORs at first scan with a `REINDEX INDEX ;` hint — never a silent misread (validated end-to-end on EC2: build v7 → open on 2.0.0 → ERROR+hint → REINDEX → serves correctly on v8, same neighbors). An in-place page converter was investigated and **rejected**: measured recall loss of −20.7 pp @ R@10 (SIFT-1M 4-bit) and catastrophic at 2-bit, from double quantization at re-encode. REINDEX-from-heap re-encodes the heap's source vectors directly (full recall) — the heap is the corpus, so this is not a rebuild-from-external-corpus. See [`docs/UPGRADING.md`](docs/UPGRADING.md). ## [1.29.7] — 2026-08-25 **Numerical-robustness patch** — `normalise_into` (run on every indexed row via `normalize_on_insert`) computed the reciprocal norm as `(1.0_f64 / norm) as f32`, which **overflows to `+inf`** when `norm` is a tiny-but-nonzero f64 (a vector whose elements are near f32 underflow, e.g. a single `~2e-39` coordinate). The `+inf` reciprocal then poisoned every element (`x * inf = inf`), so the "normalised" vector had `inf` coordinates and infinite norm — feeding garbage into the quantizer for that row. Now divides **per-element in f64** and casts each result to f32 (`(f64::from(x) / norm) as f32`), which stays finite because `|x/norm| <= |x|` for a real vector. Patch bump — **no wire change (v7), no SQL surface change, no REINDEX.** Found during the pg_turbovec 2.0.0 (turbovec 1.0.0) port's full test re-run: the pre-existing Hegel property test `prop_normalise_is_unit_norm_and_idempotent` flaked (~1 in N seeds) on the `norm inf` assertion. Added a deterministic regression test `normalise_tiny_norm_stays_finite` (fail-before/pass-after proven). ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '1.29.7';` — no REINDEX. A row inserted under an older binary whose vector hit this edge would have stored a garbage (inf) code; such a row (if any) is corrected by re-inserting it. In practice the trigger requires near-underflow input magnitudes, which real embeddings do not produce. ## [1.29.6] — 2026-08-15 **Dependency-hygiene patch — clears all outstanding RustSec advisories in the transitive tree.** Patch bump, **`Cargo.lock`-only** (no source change, no wire change (stays v7), no SQL surface change, no REINDEX, byte-identical build output — the `gemm`, `turbovec`, and `pgrx = 0.19.1` pins are unchanged, so IVF build determinism is preserved). `cargo update` pulled semver-compatible upstream fixes: - **`crossbeam-epoch` 0.9.18 → 0.9.20** (RUSTSEC-2026-0204, invalid pointer dereference in the `fmt::Pointer` impl) — reached via `rayon`, which pg_turbovec uses for build/scan parallelism. This is the one advisory in a shipped runtime path; pg_turbovec never formats those pointers, so exposure was nil, but the dependency is now patched. - **`tokio-postgres` 0.7.17 → ≥ 0.7.18** (RUSTSEC-2026-0178), **`postgres-protocol` 0.6.11 → ≥ 0.6.12** (RUSTSEC-2026-0179 / 0180) — these reach the tree ONLY through `pgrx-tests`, a **dev-dependency** (the test harness's PostgreSQL client). They are **not present in the shipped `pg_turbovec.so`** and were never a production exposure; patched for a clean `cargo audit`. Remaining `cargo audit` output is two "unmaintained" *warnings* (`serde_cbor` via `pgrx`, `paste` via `turbovec`/`statrs` and `hegeltest`) — not vulnerabilities, and both are upstream-controlled (pgrx's `PostgresType` serialization and a transitive math dep), not fixable from this crate. **Upstream `turbovec`:** intentionally NOT advanced. Our pinned rev (`befc4cbf`, the tip of the `feat/unblock-inverse-repack` branch, 2026-07-06) carries the `pub from_parts` + `pack::unblock` surface pg_turbovec depends on and is NEWER than upstream `main`'s last release (0.6.0, 2026-05-25). `main` diverged in a slim-down direction that removes that surface, so advancing the pin would break the build; rebasing our branch onto main's encode-vectorization / codebook-caching improvements is a deliberate, separately-validated kernel change (wire / determinism-sensitive), not a routine bump. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '1.29.6';` — no REINDEX, no behavior change (dependency hygiene only). ## [1.29.5] — 2026-08-15 **Production-hardening patch from a deep code re-audit + an at-scale feature stress test.** Patch bump — **no wire change (stays v7), no SQL surface change, no REINDEX** (`ALTER EXTENSION pg_turbovec UPDATE` is sufficient). Every fix is a guard or ordering correction; none change results on valid input. - **IVF incremental `INSERT` regression fix (introduced in v1.29.4).** v1.29.4 added an on-disk duplicate-id guard to the deferred-flush reconcile path, but it ran **ungated** — and an IVF index legitimately stores the same external id in multiple cells (soft-assignment to the 2nd..Mth nearest cell). So the **first `INSERT` into any soft-assigned IVF index** tripped `refusing to reconcile onto a corrupt .tvim id table (id N appears in more than one slot on disk)` and aborted — breaking IVF incremental insert. The guard is now gated to the bijective flat/single kind (`lists == 0`), matching the insert/read paths. Regression test `ivf_insert_after_soft_assign_does_not_falsely_abort`. - **Multi-index partial-flush corruption-spreader (C-NEW-1).** On a table with 2+ turbovec indexes, if the PreCommit flush of the second index tripped a persist guard (`ERROR`/longjmp), the first index's relfile pages were already physically written (GenericXLog is not rolled back on abort) — leaving one index with phantom CTIDs and a sibling missing rows. The PreCommit flush now does a **pure validate-all pass (no I/O) before writing any index**, so a guard trip aborts with zero physical writes. - **Corrupt/torn-meta unbounded read + palloc guard (H-NEW-2/3).** `read_chain` now bounds the chain against the physical relation length (`RelationGetNumberOfBlocksInFork`) and rejects `rows_per_page == 0` before allocating or walking blocks — a bit-flipped/truncated meta no longer over-reads past EOF or allocates a `Vec` sized by a corrupt `n_vectors`; it raises `ERRCODE_DATA_CORRUPTED` + a REINDEX hint. - **`SAVEPOINT` / subtransaction rollback (M-NEW-4).** Registered a `SubXactCallback`; a `ROLLBACK TO SAVEPOINT` now conservatively invalidates the dirty cache set, so a rolled-back insert is no longer persisted onto disk at top-level commit (was index bloat + silent recall loss, masked by recheck). - **VACUUM shrink guard gap (M-NEW-5).** The flat swap-remove VACUUM path — the one whole-relfile mutation that bypassed `write_full_inner`'s id-0/dup guard — now re-checks the surviving ids' bijection before committing the shrink (flat kind only). - **Empty-index KNN query fix (NEW BUG #1).** `ORDER BY emb <-> q LIMIT k` on a freshly-created, not-yet-populated index returned 0 rows instead of `ERROR: query dim N != index dim 0` (the dim check now runs after the empty-index early return) — fixes the "create index, backfill later" pattern. - **Overflow hardening (L-NEW-7):** `MetaPageData::total_blocks()` uses `saturating_add` so a corrupt meta can't wrap a large block count to a small total. The re-audit also VERIFIED (independently, at production scale on EC2): v1.29.4's field-report torn-write fix holds clean for **95 min / 85 mid-flush writer restarts** (v1.29.3 corrupts after 2), and the whole flat/IVF feature matrix + the sparsevec OOM guard are solid. Known issues still tracked for a dedicated release (graph-kind only): graph concurrent-insert lost-update, graph wrong-results at dim>=512, graph under-return at 128d, graph build un-cancellable, `turbovec_check` graph-adjacency-blind, and the scan sentinel-ctid projection. The graph kind (`WITH (graph = true)`) should be treated as experimental. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '1.29.5';` — no REINDEX, no downtime beyond the `.so` swap + reconnect. ## [1.29.4] — 2026-08-14 **Data-corruption fix: VACUUM-INDEPENDENT torn-write of the `.tvim` relfile on writer interrupt/restart.** Patch bump — **no wire change (stays v7), no SQL surface change, no REINDEX to upgrade** (existing indexes are read + written correctly in place by the new binary). A running index that already corrupted needs a one-time `REINDEX` to clear the bad state, but the *upgrade itself* is `ALTER EXTENSION` + restart. v1.29.2/1.29.3 fixed the concurrent-VACUUM lost-update races. A deployment (pg.ddx.io/agora) that runs **no VACUUM at all** (`autovacuum_count = 0`, no DELETE) still corrupted continuously — signature always `id 0` in more than one slot, ~every 30 min, on a 1.84M-row 768d `bit_width=4` partial flat index driven by a continuous `INSERT … ON CONFLICT DO UPDATE` upsert writer that the orchestrator **restarts every ~15 min**. Root cause (reproduced + dumped on EC2, unpatched corrupts in ~40 s): the deferred `aminsert` PreCommit flush rewrites the whole relfile via `write_full_inner`, which polls `check_for_interrupts!()` inside `write_chain_at` and commits each `GenericXLog` batch as a **physical, WAL-logged page change that is NOT rolled back on transaction abort**. The writer-restart is a `pg_terminate_backend` (SIGTERM → `ProcDiePending`); when it lands mid-flush it longjmps (`FATAL`) out of the chain write. With the pre-fix ordering — **meta page written FIRST, then the codes/scales/ids chains** — an interrupt after the meta commit but before the ids chain finished left `meta.n_vectors = N` pointing at an ids chain whose newly-appended (or in-place) region was still the zero-initialised page bytes. On reload that reads back as a contiguous run of `id 0` slots — `duplicate ids … id 0 appears in more than one slot` (XX001) + SIGABRT — with `count_matches` still TRUE (the chain is physically N slots long). On disk we dumped the exact fingerprint: a contiguous zeroed run equal to one batch's append count (16 slots at n=1,840,016; 730 slots == the tail of a single ids page at n=1,840,149), never scattered garbage. VACUUM is not involved. Fixed by making `write_full_inner` write the row **CHAINS FIRST and the META PAGE LAST** (the same crash-safety invariant the build path's `write_blocked_phase_and_meta` already documents). An interrupt during the chain writes now leaves block 0 (the meta) UNCHANGED, so readers and the next reconcile-flush observe the previous, fully-valid `n_vectors` + chains; any pages written past the old count are simply unreferenced until the next full rewrite / `RelationTruncate` (which stays after the meta write, so a shrink never leaves the meta referencing a page past EOF). The meta write is a single `GenericXLog` page op with no interrupt poll, so it cannot tear. Also added a **belt-and-suspenders guard on the deferred-flush /reconcile path** (`reconcile_and_write_flush`): before splicing this transaction's upserts onto the current on-disk ids and re-persisting, it runs `first_duplicate_id` over the ids it just re-read and, if the on-disk state is ALREADY corrupt (e.g. a hole left by an older binary), ABORTS the transaction with a `REINDEX INDEX` hint instead of entrenching it — a retryable abort beats propagating on-disk corruption. A NON-restarting, long-lived single writer (no `statement_timeout`, no cancels, never `pg_terminate_backend`'d) never hits the interrupt window and does not trigger this bug — a valid immediate operational workaround. ### Validation - Fail-before/pass-after unit test `torn_flush_on_interrupt_meta_last_stays_clean`: injects the exact interrupt window and shows meta-FIRST + tear → id-0/duplicate on reload, meta-LAST (fixed) + tear → clean bijection. - Sustained no-VACUUM A/B on the reproduction (1.84M-row 768d bit_width=4 partial flat index, upsert writer, writer restarted via `pg_terminate_backend`): unpatched corrupts in ~40 s at the first restart; patched runs 2 h+ across dozens of restart cycles with `turbovec_check` `is_corrupt = false` throughout AND real forward progress (n_vectors climbs cleanly). ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '1.29.4';` — no REINDEX, no downtime beyond the `.so` swap + reconnect. An index already corrupted by a pre-1.29.4 binary needs a one-time `REINDEX INDEX ;`. ## [1.29.3] — 2026-08-14 **Runtime-hardening patch from a full code audit.** Patch bump — **no wire change (stays v7), no SQL surface change, no REINDEX** (`ALTER EXTENSION pg_turbovec UPDATE` is sufficient; existing indexes are read and written unchanged). All fixes are defensive guards that turn a crash/OOM into a clean `ERROR`; none change results on valid input. - **Adversarial-input OOM guard on the `sparsevec` densify paths.** `sparsevec` permits a dimension up to 1e9, but the `::vector` cast (`sparsevec_to_vector`) and `sum(sparsevec)` (`SparsevecAccum:: ensure_dim`) densified into `vec![0.0; dim]` (up to 4–8 GB) BEFORE the 16000-element vector cap was checked — so one unprivileged `SELECT '{}/1000000000'::sparsevec::vector` (or `sum()` over such a row) could OOM a backend. Both paths now reject `dim > 16000` before allocating. Test: `sparsevec_oversized_dim_densify_errors_not_ooms`. - **Clean `ERROR` (not a Rust panic across the FFI boundary) on a torn/corrupt scan slot lookup.** `ReadOnlyIndex::search` / `search_masked` / `id_at_slot` and the out-of-core IVF id remap indexed `slot_to_id[slot]` directly; a corrupt/torn read yielding a slot past the id table panicked (a backend-abort risk under load). Now `id_at_slot_checked` / `id_at_global_checked` raise a PG `ERROR` with a `REINDEX INDEX` hint. - **VACUUM stays cancellable.** The flat swap-remove loop (`vacuum.rs`) holds the exclusive relfile-rewrite lock across every dead slot; added `check_for_interrupts!()` at the loop top so a VACUUM deleting millions of rows can be cancelled (and doesn't hold the lock uncancellably against readers) between slots. The audit also confirmed the v1.29.2 flat/IVF corruption fixes, determinism gates, dup-id guards, GUC safety, and 1-bit build fencing are solid. It surfaced one CRITICAL follow-up NOT fixed here: the **graph-kind incremental `INSERT`** still has the pre-v1.29.2 lost-update window (a non-atomic read-then-rewrite) that reconcile-on- flush closed for the flat/IVF kinds — tracked for a dedicated fix release with a concurrent-inserter reproduction test + sustained-load validation per the corruption HARD MANDATE. Graph inserts should stay serial/one-writer until then (the documented build-then-serve model). ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '1.29.3';` — no REINDEX, no downtime beyond the `.so` swap + reconnect. ## [1.29.2] — 2026-08-13 **Data-corruption fix: concurrent VACUUM + deferred-flush lost-update in the flat/single relfile.** Patch bump — **no wire change (stays v7), no SQL surface change, no REINDEX to upgrade** (existing indexes are read + written correctly in place by the new binary). A running index that already corrupted needs a one-time `REINDEX` to clear the bad state, but the *upgrade itself* is `ALTER EXTENSION` + restart. Continuous concurrent inserts + autovacuum into a large flat `turbovec.vector` index corrupted the `.tvim` id table ("duplicate ids … id 0 appears in more than one slot", XX001) and SIGABRT-crashed backends. Root cause was a set of concurrency races between the whole-relfile operations, NOT on-disk-format corruption: 1. **Deferred-flush lost-update (the load-bearing bug).** Each writer loads the whole index into a per-backend in-memory snapshot on its first mutation and, at `PreCommit`, rewrote the ENTIRE relfile from that snapshot. A concurrent VACUUM that shrank the relfile (deleted dead rows) in between was **clobbered** — the committing writer resurrected the deleted rows from its stale pre-VACUUM snapshot, reintroducing dead/duplicate ids. Fixed with **reconcile-on-flush**: under the exclusive rewrite lock at `PreCommit`, re-read the CURRENT on-disk `(codes, scales, ids)` and splice ONLY this transaction's upserted ids onto it (`relfile::reconcile_flush_image` / `reconcile_and_write_flush`), never a blind whole-snapshot overwrite. VACUUM's deletes and other backends' committed inserts survive. 2. **VACUUM stale-snapshot write.** `ambulkdelete` read the ids chain, computed dead slots, and only THEN took the write lock for the swap-remove — using the pre-lock `meta`/`dead_slots`. A flush completing in that window left VACUUM swap-removing against a just-rewritten chain. Fixed by taking the exclusive rewrite lock for the WHOLE bulkdelete (ids read → dead-slot compute → swap → shrink → truncate) on a snapshot re-read under the lock. 3. **`read_rotation` unlocked window.** `install_whole_index` / `install_graph_index` read the chains under the shared lock, then read the rotation chain UNLOCKED; a concurrent rewrite could move or truncate it. Fixed by bracketing the whole install (chains + rotation + adjacency + tombstones) under one outer shared lock. 4. **Stale-meta torn reads in `read_ids_only`** (used by `turbovec.turbovec_check` and the scan visibility path): it read the ids chain with the caller's pre-lock `meta`. Fixed by re-reading meta under the shared lock, atomically with the ids. 5. **Hardening (last line of defense):** the persist path now refuses to write an id-0 or a duplicate id for the bijective flat/single kind, aborting the transaction rather than persisting the corruption signature. Proven on an i4i.8xlarge (PG18) against a 1.76M-row 768d flat index: a comprehensive sustained-load run (VACUUM every 3s + 4 writers doing batch-16/128 `ON CONFLICT DO UPDATE` with fresh-backend-first-insert churn + 3 scanners) that **corrupts the pre-fix binary in 60 s** (id-0 duplicate, 35 SIGABRT/recovery events) ran **70.5 minutes on the fixed binary with ZERO corruption**: 46/46 in-flight `turbovec_check` reads clean, 0 dup-id, 0 SIGABRT, 0 recovery, ~19,470 write transactions concurrent with VACUUM. Insert throughput regresses ~13% (15→13 rows/s on a 1M-row flat index, single writer) — the reconcile adds a full ids+scales re-read at flush; the dominant per-flush whole-relfile rewrite is unchanged. Warm kNN and build time unaffected (1M build 72 s, warm kNN ~112 ms/query server-side). Repro harness in `benches/corruption-repro/`. ### Migration `ALTER EXTENSION pg_turbovec UPDATE` + restart the backend to pick up the new `.so`. No wire change, no REINDEX for the upgrade. An index that has *already* corrupted under the old binary should be `REINDEX`ed once to clear the bad on-disk state. ## [1.29.1] — 2026-08-11 **Packaging fix: ship in-place `ALTER EXTENSION UPDATE` scripts.** Patch bump — no wire change (stays v7), no code change, no REINDEX. v1.28.4 added `turbovec.turbovec_check()` to the full-install schema but **not to any upgrade path**, so operators who ran `ALTER EXTENSION pg_turbovec UPDATE TO '1.28.4'` in place never got the function the changelog advertised (agora report 2026-08-11). Root cause: the repo shipped only full-install `pg_turbovec--.sql` and no `pg_turbovec----.sql` upgrade scripts, so PostgreSQL's `ALTER EXTENSION UPDATE` had no SQL delta to apply — only the `.so` changed. - Adds the runnable upgrade scripts under `sql/`: the `1.28.3->1.28.4` edge `CREATE`s `turbovec_check`; forward edges (`1.28.4->1.29.0`, `1.29.0->1.29.1`) are documented no-ops so the update chain resolves. - New `scripts/drift-check.sh` gate (#12) fails the build if a release lacks its `sql/pg_turbovec----.sql` upgrade edge, so this can never silently regress again. - Also codifies the **2026-08-11 hard mandate** in `AGENTS.md`: no corruption ever; non-major upgrades must be zero-format-change or online-upgradable in place (no REINDEX-from-corpus for a minor); even major format breaks must offer an offline in-place converter tool. ## [1.29.0] — 2026-08-07 **Partitioned-scale support (toward 1T+ vectors) + the 1-bit quantization foundation.** Minor bump — purely ADDITIVE SQL surface; no wire-format change (`MetaPageData::version` stays 7), **no REINDEX**. `ALTER EXTENSION pg_turbovec UPDATE TO '1.29.0';` is sufficient. ### Partitioned scale (Phase S-0) PostgreSQL caps a single heap at 32 TB (~11B rows @768d), so 1T+ vectors must live in a partitioned table. pg_turbovec now documents this directly, and the surprising-but-verified core finding is that **PostgreSQL's native `Merge Append` over a hash-partitioned parent already does a correct, LAZY global top-k across per-partition turbovec indexes with zero new code** — the scatter→gather→merge is byte-identical to a single-table exact top-k (PoC: 10/10 overlap), and the merge pulls ~`k + probed` rows, not `k·N`. - **`docs/PARTITIONED_SCALE.md`**: a cookbook for scaling to 1–10B vectors *today* with no AM changes — partition sizing (10M–50M/partition), embarrassingly-parallel per-partition build orchestration (the key build-time lever: ~2.2 days at 800-way vs ~47 days naive for 1T), native insert routing, per-partition `VACUUM`/`REINDEX CONCURRENTLY`, and the query patterns (native parent `ORDER BY emb <=> q LIMIT k`, plus the explicit UNION-ALL fallback). PoC at `benches/poc/scatter_gather_partitioned_topk.sql`. - **Partition pruning for very large N** (Phase S-1, the partition-level coarse quantizer) is **designed** (doc §6) but **deferred** to a later release — its first cut didn't pass its own correctness `#[pg_test]` when run, and this project does not ship unproven partition-selection code (a silent recall bug is worse than a deferral). It lands alongside the 1-bit completion, real-PG validated. The free scatter-gather above covers 1–10B today. ### 1-bit (sign binary quantization) foundation `WITH (bit_width = 1)` is now **accepted** by the reloption (range `1..=4`; 1 = sign-BQ, 2/3/4 = TurboQuant; the GUC default stays `2..=4`, so 1-bit is opt-in and never a default). The turbovec kernel hard-rejects `bit_width < 2`, so 1-bit is a distinct sign-BQ scheme (per-coord sign bit + Hamming coarse + exact heap rerank), matching pgvector/DiskANN/Qdrant. The exact-heap-rerank floor auto-widens for a 1-bit index (reusing `hi_dim_rerank`/`search_k`/`oversample` — no new mechanism). The sign-BQ core is implemented + unit-tested: **mean-centering** (the mandatory footgun fix — raw sign-at-zero collapses to R@10≈0 on non-zero-centered data), degenerate all-same-sign detection, and half-of-2-bit storage (`dim/8` codes/vec, no scale). **Not yet functional end-to-end:** a `CREATE INDEX ... WITH (bit_width = 1)` currently **ERRORs clearly** ("not yet implemented; use bit_width = 2, 3, or 4") — it never half-builds or ships a silent all-ones landmine. The end-to-end 1-bit encode/scan path requires a wire bump to v8 and lands in a later release **after full real-PG validation of the new scan kernel + wire format** (deliberately not rushed into this release). The reloption is accepted now for forward compatibility. ## [1.28.4] — 2026-08-07 **Corruption fix: eliminate the dual row-counter drift that could persist a `.tvim` id table claiming more rows than it holds, plus a new `turbovec.turbovec_check(regclass)` integrity function.** Patch bump — no wire-format change (stays v7), **no REINDEX required by the upgrade itself** (a currently-corrupt index still needs `REINDEX` or `DROP` + `CREATE`; see Migration). Reported 2026-08-06 (agora / pg.ddx.io) on a 1.74M-row PG18.4 index: the `.tvim` id table developed "duplicate ids (id 0 appears in more than one slot)," write-blocking the whole indexed table, and `REINDEX` didn't durably repair it (only `DROP` + `CREATE` held). This release fixes the root cause on the write path and adds an operator-facing way to detect the corruption without attempting a write. ### B (root cause) — single source of truth for the persisted row count The deferred `aminsert` flush (`xact::flush_to_relfile`) persisted `PersistState.n_vectors` — a SEPARATELY-incremented `i64` counter — as the on-disk row count, passed to `relfile::write_full_with_prepared` as an INDEPENDENT argument alongside `idx.slot_to_id()`. Those two counters can drift (an `IdAlreadyPresent` remove+re-add, a `mark_dirty` closure that didn't run, a future insert-path refactor). If `n_vectors` ever exceeded `slot_to_id.len()`, the meta page claimed more rows than the ids chain held, and reload (`read_full`) over-read the ids chain into zeroed trailing slots — surfacing as "id 0 appears in more than one slot" (a real CTID never encodes to 0, so an id-0 slot is always zeroed bytes). - The flush now DERIVES the persisted count from `idx.slot_to_id().len()` (the authoritative array whose bytes actually land in the ids chain) via the new pure, unit-tested `xact::reconciled_row_count`, making the drift class structurally impossible on the write path. - A hard runtime guard aborts the transaction (the matching `Abort` callback then evicts the dirty cache entry, so the next access reloads clean committed state) if the mirror ever drifts — belt-and-suspenders with the pre-existing `assert_eq!` in `relfile::write_full_inner` (now documented as a hard release-build persist-site guard that must NOT be downgraded to `debug_assert!`). - `PersistState.n_vectors` is retained only as a mirror (and the scan-visibility snapshot); the on-disk truth is `slot_to_id.len()`. ### A (recovery) — REINDEX now durably repairs With B fixed, a `REINDEX` (which allocates a fresh relfilenode; the per-backend cache drops the stale entry on the relfilenode mismatch in `am_lookup_for_mutation`) produces a clean id bijection that the write path can no longer re-corrupt. The pre-1.28.4 "REINDEX reports success but the corruption returns minutes later" behavior was the ongoing insert workload re-corrupting the freshly rebuilt index via the drift above (or a crash — see C); B removes the write-path source. ### C (crash-safety) — KNOWN REMAINING GAP, scoped honestly This release does NOT make the `.tvim` id table fully WAL-crash-safe. An unclean shutdown / `pg_resetwal` can still discard WAL that extended the id chain, leaving two slots claiming id 0. Full WAL-crash-safety of the id table is a larger durability change, deliberately not attempted inside a corruption patch (shipping a half-done risky durability change would be worse than the documented gap). Mitigations in place: the insert AND read paths already ERROR loudly on such a relfile with `HINT: REINDEX INDEX ;` (v1.28.2), and `turbovec_check()` (below) now makes it detectable without a write. Recovery: `REINDEX INDEX ;` (now durable per A), or `DROP` + `CREATE` for an index corrupted before 1.28.4. ### D (integrity check) — `turbovec.turbovec_check(regclass)` New read-only, ownership-checked function that reads the meta + ids chain and reports enough for an operator to detect the corruption WITHOUT attempting a write (the only signal pre-1.28.4 detection gave, which blocks the whole table): ```sql SELECT * FROM turbovec.turbovec_check('my_idx'::regclass); -- wire_version | kind | n_vectors | slot_count | count_matches -- duplicate_id | is_corrupt | tombstone_density ``` `is_corrupt` (true when a flat index has a duplicate id OR the counts disagree) is the column monitoring should alert on. Reuses `scan::first_duplicate_id`. Takes only `AccessShareLock`, so it never blocks writers; non-owners get a permission-denied ERROR (portable across PG13-19 via `pg_class_ownercheck` / `object_ownercheck`). ### Tests - `xact::reconcile_tests` — load-independent pure-Rust drift-guard unit tests (run under `cargo test --lib`, no cluster needed): the regression gate for B. - `persist_large_batch_no_duplicate_id_after_reload` (`#[pg_test]`) — builds a populated flat index, adds a 128-row batch of new ids via the same path aminsert uses, flushes, and re-reads the persisted relfile asserting no duplicate id and `meta.n_vectors == slot_to_id.len()`. - `persist_row_count_drift_aborts` (`#[pg_test]`) — a deliberately drifted `PersistState` must abort the flush, never persist a corrupt relfile. - `turbovec_check_reports_healthy_flat_index` (`#[pg_test]`) — D. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '1.28.4';` is sufficient and cannot fail on existing indexes. No wire change (stays v7), no REINDEX required by the upgrade. **A currently-corrupt index still needs recovery**: `REINDEX INDEX ;` (now durable) or `DROP` + `CREATE`. The fix STOPS new write-path corruption; it does not repair an index already corrupted by a prior crash. ## [1.28.3] — 2026-07-31 **Managed-PostgreSQL readiness (audit follow-up): interrupt handling + portability gates + deployment docs.** Patch bump — no wire-format change (stays v7), no SQL-surface change, **no REINDEX**. Addresses a managed/hosted-PostgreSQL readiness audit of v1.28.2. - **P0 — interrupt handling.** The extension had zero `CHECK_FOR_INTERRUPTS` call sites, so long scans/builds/flushes were uninterruptible and `statement_timeout` / `pg_cancel_backend()` (and the platform admin signals delivered by the same mechanism) were ignored for seconds. Added polling at every point pg_turbovec controls where **no buffer content lock is held**: - each `amgettuple` entry (makes the iterative-scan refill loop promptly cancellable), - each `CREATE INDEX` build-stage boundary (k-means / assign-sweep / quantize-encode / prepare-and-persist), - the top of each `GenericXLog` batch in the relfile write path (large `aminsert` PreCommit flushes). **Known limitation:** a single large *flat-kind* `search()` call is still internally uninterruptible (one call into the Postgres-free kernel crate). Finer-grained polling there needs a kernel block-scoring API and is tracked follow-up; prefer the IVF/graph kinds for large corpora meanwhile. - **P2 — portability drift-check gates.** `scripts/drift-check.sh` now **fails the build** if (11a) any raw-WAL / raw-storage primitive appears at a call site (`log_newpage`, `XLogInsert`, `smgrwrite`, `smgrextend`, `PageSetLSN`, `FlushRelationBuffers`, ...) — durability must stay 100% `GenericXLog`; or (11b) any GUC is registered with a non-`Userset` context. Both invariants hold today; the gate stops a future commit from silently reintroducing a managed-adoption blocker. - **P2 — docs.** New `docs/DEPLOYING_ON_MANAGED_POSTGRES.md` consolidating durability/replication, restricted-superuser compatibility, read-replica behavior, cancellation, the memory / O(n)-per-transaction insert model, build cost, and the upgrade / wire-format policy. **Deferred (tracked, larger scope):** making the in-place index rewrite atomic (write the meta page last, as the single commit point; today it self-heals via the `xs_recheckorderby` heap backstop); a full out-of-process TAP suite (crash recovery / replication / failover / cancellation-latency / codec fuzz); a kernel block-scoring API for fine-grained flat-scan cancellation. See the audit for detail. ## [1.28.2] — 2026-07-31 **Bug fix: detect duplicate-id corrupt `.tvim` relfiles and fail loudly with a `REINDEX` hint, instead of silently mis-serving reads while failing every write.** Patch bump — no wire-format change (stays v7), no SQL-surface change, **no REINDEX required by the upgrade itself** (a corrupt index does need one; see below). Reported 2026-07-30 (agora / pg.ddx.io): a 1.6M-row index on PG18.4 rejected **100% of `INSERT`s for ~1.5 days** (~47k logged failures) with an opaque `aminsert: corrupt relfile pages: duplicate ids in .tvim file`, while `pg_index.indisvalid` stayed **true** and scans kept answering — so the corruption was invisible to health checks and a backfill retry-loop turned it into an availability incident. The likely origin was several crash-recovery / `pg_resetwal` cycles that left the on-disk id table with duplicate ids. Root cause was an asymmetry: the **write** path validated the id table is a bijection (turbovec's `from_id_map_parts` rejects duplicates) and hard-errored, but the **read** path (`ReadOnlyIndex::from_parts`, which only builds `slot_to_id`) never checked — so a corrupt index failed writes yet still served (duplicated / mis-ranked) scan results and looked valid. Fix: - The read/open path (`scan::assert_ids_unique_or_reindex`, called in the whole-index, graph, and out-of-core cache-install paths) now detects duplicate ids once per backend at cache-install and `ERROR`s (`ERRCODE_DATA_CORRUPTED`) with `HINT: REINDEX INDEX ;` — the same actionable shape as the pre-v7 legacy gate. A corrupt index now fails loudly on **reads** too, rather than silently mis-serving. - The insert-path error carries the same `REINDEX` hint, so a backfill loop gets a clear "rebuild me" signal instead of retrying an opaque error forever. - **Recovery:** `REINDEX INDEX ;` (or `... CONCURRENTLY`) rebuilds a clean id bijection from the heap. There is no lighter in-place dedup — a rebuild is the supported repair. Gate: `scan::duplicate_id_tests::detects_duplicate_ids` (pure-Rust) covers the bijection predicate incl. the exact reported id. **Still open (tracked, larger scope):** the `.tvim` id table is not fully crash-safe against `pg_resetwal` / unclean shutdown — WAL-logging it well enough to recover cleanly (or refusing to come up `valid` when it can't) is the deeper durability fix and is deferred to a future release. v1.28.2 makes the corruption **detectable and actionable**; it does not yet prevent it from being introduced. ## [1.28.1] — 2026-07-28 **Nix flake packaging.** Patch bump — packaging only; no code change, no wire change (stays v7), no SQL-surface change, **no REINDEX**. A customer installing on PostgreSQL 18 via Nix found the v1.28.0 tag had no `flake.nix` ("no Nix output"). This release adds one: - `nix build github:gburd/pg_turbovec#pg_turbovec_NN` for NN in 13–19 (`default` = PG18), built with nixpkgs' `buildPgrxExtension` + a flake-local `cargo-pgrx` 0.19.1 (nixpkgs' pinned set stops at 0.18.x). The turbovec git dependency's vendor hash is pinned via `cargoLock.outputHashes`. - Output layout: `lib/pg_turbovec.so` + `share/postgresql/extension/pg_turbovec--.sql` + `.control`. On PG18+ a stock server can use it in place via `extension_control_path` / `dynamic_library_path` (validated: CREATE EXTENSION + distance fns + a turbovec index scan against nixpkgs' postgresql_18, served straight from the store path). - `nix develop`: dev shell with rust, cargo-pgrx 0.19.1, clang/bindgen, openblas, and the PG-from-source deps (bison/flex/readline/zlib/icu). - The `#[pg_test]` suite stays CI's job (needs a live cluster); the flake package build itself runs no tests (`doCheck = false`). ## [1.28.0] — 2026-07-28 **PostgreSQL 19 (beta1) support via the pgrx 0.17 → 0.19.1 upgrade.** Minor bump — framework/platform change; no wire-format change (stays v7), no SQL-surface change, **no REINDEX**. Existing indexes and queries are unaffected; `ALTER EXTENSION pg_turbovec UPDATE` suffices. - **New:** `pg19` Cargo feature + CI matrix leg. PG19 is upstream **beta** (19beta1); support is experimental until PG19 RC/GA, at which point it will be re-validated. - **pgrx 0.19.1** (from 0.17.0, a two-major jump): - The `pgrx_embed` second-compilation-pass model is gone in pgrx 0.18+ — deleted `src/bin/pgrx_embed.rs` and the `[[bin]]` stanza; SQL entity metadata now lives in the `.so`'s own linker section. - Rust **edition 2024**, MSRV **1.96** (was 2021 / 1.85). The edition-2024 `unsafe_op_in_unsafe_fn` lint is allowed crate-wide: the index-AM callbacks are `unsafe extern` FFI boundaries whose entire bodies are unsafe by construction; the per-fn `# Safety` contracts remain the audit surface. - cargo-pgrx 0.19.1 required (`cargo install --locked cargo-pgrx --version 0.19.1`). - **PG19 C-API deltas handled** (all version-gated; pg13–18 builds byte-identically unaffected): - `LockBuffer`'s `mode` parameter changed from `int` (i32) to `BufferLockMode::Type` (u32) — new `relfile::lock_buffer_mode` shim used at all 5 call sites. - `relfilenode_from_relation` gained the pg19 arm (same `rd_locator.relNumber` layout as pg16+; the cfg just didn't cover pg19, leaving the fn body empty on pg19). - Verified: all 7 versions (pg13–pg19) compile with 0 errors and 0 warnings; pure-Rust determinism gates (reservoir, IVF k-means, build-pool invariance) green on pg13 and pg19; the full `#[pg_test]` suite gate runs in CI across the 7-leg matrix. ## [1.27.3] — 2026-07-12 **Phase Q-4c: clear the IVF build cliff — batch the k-means reservoir rotation into one parallel GEMM.** Patch bump — build-SPEED change only. The persisted IVF centroids + codes are **byte-identical to v1.27.2** for a fixed (corpus, seed, lists, dim); no wire-format change (stays v7), no SQL-surface change, **no REINDEX**. Profiling a real 1M×1024/lists=4096 build on a 32-vCPU AVX-512 host showed the v1.27.2 lazy-per-row-rotation was **still ~85% of a ~14.5-min build**: the scalar O(dim²) `rotate_unit` ran ~600k times (reservoir fill + replacements) single-threaded, and at dim=1024 each rotation is ~1M FLOPs. Reducing the call count (v1.27.2) wasn't enough — the per-row *scalar* rotation was the wrong primitive. The fix: the reservoir now stores **normalised-but-unrotated** rows and rotates the whole training sample **once at drain** via `ivf::rotate_corpus_into` — a parallel BLAS GEMM that already exists and is already gated byte-identical (`rotate_corpus_bit_identical_across_pool_sizes`). The per-row hot path is now just the O(dim) normalise; the O(dim²) work is one GEMM over the ≤`256·lists` kept rows across all cores. Byte-identical output: same rows selected (RNG/seen/cap unchanged, value-independent), same rotation math (GEMM `corpus @ R^T` == scalar `unit @ R^T`). **Measured (real corpus, 32-vCPU AVX-512, PG17.10):** - 1M×1024, lists=1024: ~870s → **138s (~6.3×)**, identical 540 MB index. - **10M×1024, lists=4096: 1562s (~26 min), 5343 MB — previously did NOT complete** (DNF, cancelled at 33.7 min / 16%). The build cliff is cleared; the build now scales roughly linearly (per-stage trace shows no single dominant stage). - Query correctness verified at both scales (self-top-1 returns the query id). Also in this release (verified on the same host): **the v1.27.2 G3 concurrency fix holds** — the 1M in-RAM kNN throughput sweep rises monotonically and **plateaus at ~40 TPS through 64 connections** (no collapse), where the pre-fix per-query-pool-churn build collapsed from 46 TPS @16 conns to 1.7 TPS @32 conns. Latency saturates gracefully (797 ms @32 → 1590 ms @64) instead of thrashing. Gate: `reservoir_tests::deferred_batch_rotation_is_byte_identical_to_eager` (pure-Rust) proves the store-unrotated-then-batch-rotate sample matches the original rotate-every-row path; the IVF determinism + pool- invariance gates are unchanged and green. ## [1.27.2] — 2026-07-11 **Phase Q-4b: kill the IVF build cliff's real bottleneck — the per-row rotation in k-means reservoir sampling.** Patch bump — build-SPEED change only. The persisted IVF centroids + codes are **byte-identical to v1.27.1** for a fixed (corpus, seed, lists, dim); no wire-format change (stays v7), no SQL-surface change, **no REINDEX**. Profiling the real 10M×1024/lists=4096 build (v1.27.1 still did not complete in budget despite Q-4a's parallel k-means) found the dominant cost was **not** k-means (~4.4%, flat in `n`) but the single-threaded scalar `rotate_unit` inside `BuildState::ivf_reservoir_push` — **~92% of the projected build.** It ran an O(dim²) rotation on **all N accepted rows** during the heap scan, even though the reservoir only ever keeps `cap = 256·lists` samples. At 10M/lists=4096 that is ~9M O(dim²) rotations computed and immediately discarded. The fix defers the rotation: the reservoir *selection* uses only the RNG stream + seen-count + cap (none depend on the rotated values), so `rotate_unit` now runs **only for the ≤ cap rows that actually land in the sample**. Rotation cost drops from O(N·dim²) to O(cap·dim²), removing the `n`-scaling that was the cliff — while producing the **identical** k-means training sample (hence identical centroids, identical codes). The RNG draw stays unconditional in the replacement branch, exactly as before, so the same source rows land in the same reservoir slots and receive the same rotation. Verified: `index::build::reservoir_tests:: lazy_rotation_is_byte_identical_to_eager` reproduces both the old (eager-rotate-always) and new (lazy-rotate-on-keep) reservoir algorithms with a shared ChaCha8 seed across 4 corpus sizes × 3 seeds (exercising both the fill and the replacement paths) and asserts the sample buffers are bit-identical (`f32::to_bits`). Pure-Rust, load-independent, green. The full `cargo pgrx test pg16` gate (including `ivf_streaming_build_determinism_byte_identical`) must be re-run on a quiet host before tagging — the local box was under load ~73 at authoring time. ## [1.27.1] — 2026-07-11 **Phase Q-4a: parallelize the IVF k-means build.** Patch bump — build-SPEED change only. The persisted IVF centroids + assignment are **byte-identical to v1.27.0** for a fixed (corpus, seed, lists, dim); no wire-format change (stays v7), no SQL-surface change, **no REINDEX**. This only makes NEW `WITH (lists = N)` builds faster. Q-1 exposed the IVF build cliff as the blocker for scale: a real 10M×1024 IVF build did not complete in a 90-min budget. Two remaining serial hot loops in the k-means build are now parallelized **bit-identically** (the rest was already parallel from v1.20.0/v1.22.1): - **`gemm_lloyd_assign`'s per-row argmin + top-2 tie-break** — run every Lloyd iteration over the whole sample — is now a data-parallel `par_iter_mut().enumerate()` writing each `assign[i]` to a fixed index, so the result is byte-identical regardless of thread count. - **`rotate_corpus_into`'s GEMM** — `Parallelism::None` → `Rayon(0)`, on the same "gemm tiling never reduces across threads" guarantee v1.22.1 established for the Lloyd cross-term GEMM. Left serial deliberately: the k-means++ next-seed CDF pick (sequential by construction) and the empty-cell reseed scan (data-dependent mutation, not a parallel-safe map). **Measured ~1.91× build speedup** (lists=4096, n=131072, dim=256, 8-core, release: 418.2s → 219.4s) with bit-identical output. Sub-linear because Lloyd iterations are sequentially dependent (only within-iteration work parallelizes) — the honest ceiling for k-means, unlike the embarrassingly-parallel graph build (v1.26.0). Verified: `ivf_coarse_model_bit_identical_across_pool_sizes` + `rotate_corpus_bit_identical_across_pool_sizes` assert byte-identical centroids + assignment across pool sizes {1, 2, auto}; `kmeans_deterministic_across_pool_sizes` stays green; full suite 338/338. This is the first step toward the scale-and-heavy-load build goals; the empty-cell/reseed serial remainder and a larger-scale (10M→100M) build validation are the follow-ups. **Migration:** `ALTER EXTENSION pg_turbovec UPDATE TO '1.27.1';` — a no-op upgrade for existing indexes; new IVF builds are just faster. ## [1.27.0] — 2026-07-10 **Phase Q-0: de-duplicate the on-disk quantized-codes storage, roughly halving the per-vector index footprint.** Minor bump (wire-format change — **REINDEX required**; additive capability, no SQL-surface removal). This clears the storage blocker for large single-node indexes. **The problem.** Every prior version persisted each vector's quantized codes TWICE: the row-major bit-plane `packed_codes` chain AND the SIMD-`blocked` chain (the output of `pack::repack(packed_codes, …)`). That doubled the dominant O(n) storage term. At 100M vectors that is the difference between ~78 GB and ~39.6 GB for 768d/4-bit (double- storage vs single). **The fix (Option A — persist only the packed codes).** The blocked layout is a PURE FUNCTION of the packed codes, so v7 drops the blocked chain entirely from disk and recomputes it once per backend at index-open via `pack::repack`. This is the same O(n) one-time compute a pre-Phase-P index already paid lazily on first scan; it's paid once per `(backend, am_version)` at cache-install and cached in the per-backend `ReadOnlyIndex`, so **warm per-query latency is unchanged** and scan results (recall, ordering) are **bit-identical** to before (the recomputed blocked layout equals the layout that used to be persisted). Option A was chosen over Option B (persist blocked, recompute packed via the new `pack::unblock`) because `repack` is the forward, already-used direction and the OOC path never touches the blocked chain — A keeps every hot path simplest. **Measured storage reduction** (`phase_q0_storage_is_deduplicated` #[pg_test], 512×128d 4-bit): the dropped blocked chain is ≥ the packed codes chain, i.e. persisting it doubled the code-storage term. Per-vector on-disk code bytes: `dim/8 * bit_width` stored ONCE (was twice). 100M projections: 768d/2-bit **19.8 GB** (was 39.6), 768d/4-bit **39.6 GB** (was 78), 1536d/2-bit **39.6 GB** (was 78). **Wire format (v6 → v7), NOT additive.** Unlike the additive v4→v5→v6 per-kind bumps, dropping a persisted chain is a real break for EVERY kind: single-vector, ColBERT, IVF, and graph indexes all now emit wire version 7 (the `kind` byte still discriminates). A pre-v7 index (v1..v6) is detected by the new `MetaPageData::is_legacy_v6()` predicate (`version < 7`); `ambeginscan` raises a clear `ERROR` that NAMES the index with a `HINT: REINDEX INDEX ;` at the first scan — never silent corruption. Verified by `ambeginscan_errors_on_legacy_v6_meta` (and the retained v1/v2 forgeries, which now also hit the unified v7 gate). No SQL surface change (no operators/types/functions/GUCs/opclasses added or removed). All index kinds (flat, IVF, graph, ColBERT) continue to work. **Migration:** 1. `ALTER EXTENSION pg_turbovec UPDATE TO '1.27.0';` 2. `REINDEX INDEX ;` — once per turbovec index, ANY kind. Until an index is REINDEXed under v7, scans against it ERROR with the REINDEX hint (they do NOT return wrong results). See `docs/UPGRADING.md`. ## [1.26.0] — 2026-07-10 **Phase G-2d(a): a partitioned/merge parallel build for the graph index kind, so it scales past the single-pass build's serial ceiling.** Minor bump (new GUC `turbovec.graph_build_partitions`, new build path, additive). **NO wire-format change** (stays v6, byte-identical on-disk CSR), no new operators/types/functions, **no REINDEX**. The single-pass Vamana build is serial by necessity (each insertion navigates the graph every prior insertion left; G-2c showed thread-parallelizing it doesn't amortize) and did not complete at 5M rows (>2h26m). This adds a structurally-parallel build: **partition** the corpus into P shards (contiguous ranges of the deterministic shuffled insertion order — each shard a uniform random sample), **build each shard's sub-graph in parallel** across the bounded rayon pool, then **stitch** via a parallel cross-shard refinement pass (greedy-search the merged graph from a global medoid entry + RobustPrune per node) and a deterministic parallel reverse-edge pass. No explicit persisted bridge edges are needed — the refinement pass creates cross-shard navigability, so the frozen v6 CSR is sufficient. **Measured (verified, not just claimed):** - **Recall parity — partitioned MATCHES or BEATS single-pass.** In a findable high-recall regime (queries drawn from the corpus, 20k×64): single-pass R@10=0.958, partitioned (P=8) **R@10=0.996** (+0.038). In a cross-cluster regime: 0.630 → 0.754 (+0.124). The refinement pass's greedy search over the merged graph surfaces better cross-shard neighbours than incremental insertion, so the partitioned graph is HIGHER quality, not a speed-for-recall trade. - **~8× parallel build speedup** (relative serial-vs-partitioned wall-clock, 8-core AVX2 box, 200k×64): P=16 ≈ 7.99×, P=8 ≈ 6.45×. Both the per-shard builds and the refinement/reverse passes parallelize — unlike the single-pass build. - **Deterministic**: bit-identical for a fixed (corpus, seed, P) AND across rayon pool sizes {1, 2, auto} (partition is a pure function of the seed; every parallel stage uses an index-ordered collect / staged fixed-id writes). **`turbovec.graph_build_partitions`** (int, default `auto`): `auto` derives P from corpus size + build-pool budget (single-pass below a size threshold); `0`/`1` forces the single-pass reference build; `N` forces N shards. The single-pass `build_vamana` is retained intact (all its tests pass) and is what `P<=1` runs — verified identical. **Migration:** `ALTER EXTENSION pg_turbovec UPDATE TO '1.26.0';`. No REINDEX — existing v6 graph indexes are unaffected; the parallel build only changes how NEW `WITH (graph = true)` indexes are constructed, to the identical on-disk shape. The 5M/10M a cloud VM gate re-run (now that the build completes at scale) is the follow-up this unblocks. ## [1.25.1] — 2026-07-09 **Release-tooling + docs/benchmark patch. No shippable code change — the compiled binary is byte-identical to v1.25.0** (no `src/` runtime change; wire format stays v6; no SQL-surface change; **no REINDEX**). This is the release that first exercises the new automated publish pipeline. - **Automated release publishing** (`.forgejo/workflows/release.yml`, `scripts/make-dist.sh`, `META.json.in`, `ci/announce.sh`): a `vX.Y.Z` tag on Codeberg now gates (compile + drift-check), builds a PGXN **source** distribution (renders `META.json` from Cargo.toml's version, runs `cargo pgrx schema` for the install SQL, zips a PGXN-layout archive), attaches it to a Codeberg release, uploads to PGXN, and submits a postgresql.org news announcement (feeds pgsql-announce). Runs on a self-hosted Forgejo runner (Codeberg's hosted 10-min cap can't fit a pgrx build); each publish step no-ops cleanly if its secrets are unset. The PGXN dist is a **source archive built with `cargo pgrx install`**, not a `pgxn install`-able package — published for discoverability + version-pinning. See `RELEASING.md` for the one-time runner + secrets setup. - **Qdrant + ANN-Benchmarks-protocol competitive benchmark** (`benches/results/qdrant_annbench_20260709/`): validated v1.25.0's `turbovec.hi_dim_rerank` at real scale — on GIST-960-1M, `auto` lifts R@10 from **0.876** (the `off` ceiling) to **0.953**, crossing the ≥0.90 and ≥0.95 bands the pre-fix engine never reached. vs Qdrant (in-RAM, a benchmark host, 1M): Qdrant wins raw latency 3–18×, pg_turbovec wins storage (SIFT 141 MB = 7–8× smaller; GIST 983 MB = 5.4× smaller than Qdrant's 5.3 GB). At 10M×960 only Qdrant built in a 90-min budget (turbovec's single-threaded k-means build cliff). - **Docs correction**: the 2026-07-08 competitive doc's pg_turbovec SIFT R@10=0.99 latency (16.84 ms) was the out-of-core path; the in-RAM number is **2.6 ms** (verified R@10=0.9915 @ 2.61 ms). Footnoted in place. **Migration:** `ALTER EXTENSION pg_turbovec UPDATE TO '1.25.1';` — a no-op upgrade (nothing in the SQL surface or on-disk format changed). ## [1.25.0] — 2026-07-09 **Gap-B fix: `turbovec.hi_dim_rerank`, a dimension-aware exact-L2 rerank-window widening that recovers high-dimensional recall.** Minor bump (one new GUC, additive). **No wire-format change** (stays v6, byte-identical to v1.24.0), no new operators/types/functions, **no REINDEX**. An offline investigation (using FAISS's trusted quantizers as measurement vehicles) established that the high-dim recall gap — e.g. GIST-1M/960d capping ~0.86 where pgvector HNSW and VectorChord reach 0.95-0.98 — is **NOT retrieval- bound**. The true nearest neighbours DO land in the probed IVF cells (measured cell recall 0.978 at probes=64, 0.996 at probes=128). The real bottleneck is **in-cell quantized ranking**: at high dim the lossy 4-bit score is noisy enough that a true neighbour often sits at rank ~200-800 *within* the probed cells, below a small `search_k`, so it never enters the exact-L2 reorder recheck. This corrects the prior "retrieval-recall ceiling" framing (corrections noted in-place). The cure is scan-side only: fetch a **wider candidate set** so the always-on exact-L2 reorder queue (`xs_recheckorderby`) re-ranks enough survivors to recover the true top-k. Measured: an SQ4 analog of TurboQuant's per-coordinate scalar quant lifts R@10 from 0.666 to **0.978** at 960-dim by reranking ~800 candidates instead of ~64. This must not blanket-widen: at low dim recall already plateaus by `search_k≈25`, so widening there is pure latency tax. Hence a **dimension-aware floor**. `turbovec.hi_dim_rerank`: - `auto` (default): apply a candidate floor of `clamp(dim, 256..=1024)` **only** for indexes with `dim >= 256`, and only ever RAISE the count — an explicit `search_k`/`oversample` override past the floor always wins. SIFT-128 is untouched (zero latency cost); GIST-960, OpenAI-1536, and the 512-768d embedding families get the wider window out of the box. - `on`: apply the floor regardless of dim. - `off`: honour `search_k`/`oversample` exactly (pre-1.25.0 behaviour). The result set is identical to setting `search_k`/`oversample` by hand to the same candidate count — this is a smarter default, not a new mechanism. Verified end-to-end by the `hi_dim_rerank_raises_high_dim_recall` `#[pg_test]` (384-dim corpus: `auto` measurably beats `off`, never regresses any query) plus the `hi_dim_rerank_tests` unit tests on the dim-scaling decision function. **Migration:** `ALTER EXTENSION pg_turbovec UPDATE TO '1.25.0';`. No REINDEX. The new `auto` default improves high-dim recall out of the box at a small high-dim-only latency cost; `SET turbovec.hi_dim_rerank = off` restores exact pre-1.25.0 scan-candidate behaviour. ## [1.24.0] — 2026-07-08 **Phase G-2b: VACUUM + incremental INSERT for the graph index kind.** Both operations previously raised a clear `ERROR` against a `WITH (graph = true)` index (v1.23.0 was build+scan only, correctness- first); this release turns them into working functionality. No wire- format change — wire format stays **v6**, existing v4/v5/v6 indexes decode byte-identical, **no REINDEX**. Minor bump because a previously-`ERROR`ing operation becomes real capability (same reasoning v1.23.0 used to justify G-2a as a minor). **VACUUM (`ambulkdelete`)** now uses the same per-slot tombstone bitmap mechanism IVF already uses (the generic `relfile::read_tombstones` / `write_tombstones_and_meta` path, confirmed to have zero IVF-specific assumptions). Tombstoned nodes never enter scan results and their out-edges are never followed; a fully-tombstoned corpus returns empty cleanly. `graph_search` gained a `tombstones` parameter threaded through the scan path. **INSERT (`aminsert`)** does a whole-relfile rewrite per insert (read every chain back, quantize+append the new vector, run a Vamana insertion via `insert_one_node_via_oracle`, persist). This is a **deliberate O(n)-per-insert cost** — explicitly NOT the deferred/batched path other kinds get — appropriate for the build-then-serve model the graph kind targets; heavy incremental churn should still REINDEX. The insertion distance oracle uses the exact quantized-code scan kernel for `dist(new, existing)` and a triangle-inequality lower bound for the `dist(existing, existing)` RobustPrune diversity check (documented safe-direction approximation: can only make pruning less aggressive, never drops a genuinely diverse edge, still hard-capped at degree R). **Two real bugs found and fixed during G-2b's own test-writing:** 1. **Relfile corruption on insert-after-VACUUM.** `write_tombstones_and_meta`'s block-offset formula for placing a new tombstone chain omitted `+ graph_count`, so a graph index's tombstone chain (once VACUUM wrote one) computed an offset that collided with the already-persisted graph adjacency chain, corrupting it on the next incremental insert (`corrupt graph adjacency chain: graph offsets[n]=0 != neighbors.len()=...`). Root-caused via instrumentation showing `graph_first == tombstone_first`. `graph_count` is `0` for every non-graph kind, so the fix is a no-op for flat/IVF/ColBERT. A second, related fix re-persists a pre-existing tombstone bitmap after the main graph write (which plans a fresh meta from scratch and would otherwise silently drop it). 2. **VACUUM entry-point dead-end.** The entry-point fallback fired only when the entry point itself was tombstoned, never when the entry point SURVIVED but every one of its out-neighbors was tombstoned — a low-degree entry point whose only edge points at a now-dead slot is a genuine dead-end (first beam-search hop expands to zero live candidates). Both the VACUUM-side fallback selection and the scan-side entry pick now treat "no live out-neighbor" as equally disqualifying as "dead" and prefer a fallback that itself has a live neighbor. **Also corrected a test-harness bug** (not a shipped-code bug): the graph #[pg_test] corpora were generated with an uncorrelated `random()` subquery that PostgreSQL hoisted and evaluated once, making every test row identical (`n_distinct = 1`) and producing spurious "recall collapse" numbers that had nothing to do with the feature under test. Correlating the inner `generate_series` to the outer row restored genuine per-row randomness; the insert/vacuum/quantization paths were correct all along. **Migration:** `ALTER EXTENSION pg_turbovec UPDATE TO '1.24.0';` only. No REINDEX, no wire change, no SQL-surface change. A v1.23.0 graph index that was built and never mutated is unaffected; graph indexes can now be VACUUMed and incrementally inserted into in place. Still deferred (unchanged from v1.23.0): G-2c (SIMD traversal + build parallelism), G-2d (the 5M-scale AVX2 HNSW-latency gate). ## [1.23.0] — 2026-07-06 **Phase G-2a: `WITH (graph = true)`, a new Vamana-style navigable- graph index kind** — the first step toward matching HNSW's query latency while keeping TurboQuant's storage compression, per the long-standing roadmap. Builds a real single-pass Vamana graph (DiskANN's algorithm: greedy search + RobustPrune per node, in a deterministic randomized insertion order) over the full corpus. A graph node's vector storage is identical to a flat index's (same TurboQuant encode path); the new adjacency chain (CSR: offsets + flat neighbor ids) and an entry-point slot id are the only new on-disk structures. Scan is a greedy beam search from the entry point, feeding results through the existing `xs_recheckorderby` machinery every other kind already uses — no new correctness surface there. **Wire format v6, additive** (same pattern as v5's `KIND_COLBERT`): existing v4 (single-vector) and v5 (ColBERT) indexes decode byte-identical under the v6 binary — verified by dedicated tests. No REINDEX for any existing index. A graph index is a brand-new shape only a v6 binary produces. **Determinism: relaxed for this kind only**, per the plan doc's explicit "three fundamental tensions" framing — deterministic for a fixed seed on one machine/thread-count (required for the test suite and for `REINDEX` reproducibility on a given host), but NOT byte-identical across machines/ISAs the way flat/IVF/ColBERT indexes are. WAL/streaming replication are unaffected either way (replicas replay the primary's actual page bytes, never rebuild independently). **Scope of this release (G-2a, correctness-first — see for the full sub-phase breakdown)**: build + scan work end to end with real, verified recall against an exact linear scan on test corpora. Explicitly NOT yet done, each tracked as a numbered follow-up: - G-2b: VACUUM/tombstone integration. `ambulkdelete` and `aminsert` against a graph index currently raise a clear `ERROR` rather than silently corrupting the graph — rebuild the index after bulk loading/deleting, the same operational model IVF's original build-then-query-only phase used before tombstones landed. - G-2c: SIMD-optimized traversal + build parallelism. This release's distance computation is plain scalar Rust; the build is correctness-first, not speed-optimized. - G-2d: the real 5M-scale, AVX2-hardware HNSW-latency gate measurement this whole feature is ultimately judged against (the gate: p50 ≤ 1.3× HNSW AND storage ≥ 6× smaller AND recall ≥ IVF at matched budget, at 5M rows/ R@0.96 on AVX2). **Not run in this release** — no latency or recall-vs-HNSW claim is made here; that measurement is the honest next step before this feature can be called production-ready, and the plan doc's own framing ("the bet failed, keep IVF" is a real possible outcome) still stands. New reloption `WITH (graph = true)`, mutually exclusive with `WITH (lists = N)` on the same index (enforced in `amoptions`). 24 new tests (287 total, up from 263), drift-check/compile-matrix/fmt clean, CI green. **Migration**: `ALTER EXTENSION pg_turbovec UPDATE TO '1.23.0';` is sufficient. No REINDEX for existing indexes. ## [1.22.2] — 2026-07-06 **Raises `turbovec.probes`'s default from 8 to 16 — the out-of-the- box recall floor was unreasonably low.** Scan-side default change only, no wire change, no SQL surface change, no REINDEX. The v1.22.1 a cloud VM competitive re-benchmark measured pg_turbovec's shipped defaults (`probes=8, search_k=32`) capping at **R@10=0.796 on SIFT-1M and R@10=0.407 on GIST-1M** — both far below any reasonable recall SLO, and a real footgun for anyone who runs `CREATE INDEX ... USING turbovec` without reading the tuning docs. `probes=16` (same `search_k=32`) measured: | Corpus | `probes=8` (old default) | `probes=16` (new default) | |---|---|---| | SIFT-1M | R@10=0.796, p50=3.0ms | R@10=0.918, p50=4.8ms | | GIST-1M | R@10=0.407, p50=18.5ms | R@10=0.557, p50=20.4ms | Roughly 1.5-1.6× the latency for +12-15 recall points on both corpora — the better point on the curve than `probes=32` (which roughly triples latency for a similar recall gain). Existing sessions/deployments that explicitly `SET turbovec.probes` are unaffected; this only changes the compiled-in default. New regression test `index_am_probes_defaults_to_16` guards the default against silent drift (matches the `index_am_iterative_scan_ defaults_to_off` precedent from v1.20.1). **Migration**: `ALTER EXTENSION pg_turbovec UPDATE TO '1.22.2';` is sufficient. No REINDEX. ## [1.22.1] — 2026-07-05 **Closes a real fraction of the IVF build-cliff gap — scan/build-path only, no wire change, no SQL surface change, no REINDEX.** v1.20.0/v1.21.0's parallel-build work row-blocked the normalize/rotate/assign-sweep stages of `ambuild`, measuring only a modest ~1.27× speedup on a 64-core box. A FLOPs analysis (triggered by the v1.22.0 GUC audit) found the real dominant cost was never row-blocked: `gemm_lloyd_assign`'s cross-term GEMM runs over the *whole* k-means training sample (`n_sample = lists × 256`, which equals the full corpus size at high `lists`) once per Lloyd iteration, up to 25 times, single-threaded (`Parallelism::None`, kept that way for on-disk determinism). At GIST-1M/960d/`lists=4096` scale this GEMM is ~26-112× more FLOPs than the already-parallelized stages — explaining why the earlier fix barely moved the needle. **The fix**: `gemm` 0.18's own internal `Parallelism::Rayon(n)` tiling produces bit-identical output to `Parallelism::None` for every shape/seed/thread-count tested — a GEMM's output tiles are independent dot-product reductions over the shared contraction dimension, so (unlike a cross-thread SUM, which does need the fixed-partition-order bookkeeping k-means' centroid-update step already has) thread count can never perturb a GEMM's per-element result. One line changed: `Parallelism::None` → `Parallelism:: Rayon(0)`. Via `rayon::current_num_threads()`, this automatically and correctly respects `turbovec.build_parallelism`'s bounded pool with zero extra plumbing (`train_kmeans` already runs inside `build_pool::install(pool, ..)`). **Measured on real hardware, real scale** (16-core AVX-512 a cloud VM instance, GIST-1M corpus shape: `n_sample=1,048,576, dim=960, lists=4096`, full 25-iteration k-means training): | Variant | Wall clock | vs today | |---|---:|---:| | `Parallelism::None` (v1.22.0, shipped) | 2686.6s (~44.8 min) | baseline | | `Parallelism::Rayon(0)` (this release) | 768.4s (~12.8 min) | **3.50×** | Centroids confirmed bit-identical between the two runs. **The investigation took two wrong turns before this number, both worth recording rather than hiding**: (1) an early microbenchmark of `gemm` at specific shapes segfaulted with a real gdb backtrace into `gemm-common` internals, looking exactly like a crate memory-safety bug — root cause was the *test harness's own* transposed-read stride bug (a genuine ~16M-element out-of-bounds read), not a `gemm` bug; replicating the real call site's exact strides showed no issue. (2) A later a cloud VM timing harness's naive "serial baseline" — an explicit 1-thread `rayon::ThreadPoolBuilder` wrapped around the whole training call, meant to isolate the GEMM's own parallelism — also accidentally forced the *unrelated, already-parallel* k-means++ seeding phase down to 1 thread, making the "today" baseline look far slower than v1.22.0 actually behaves in production (where seeding always runs on the real `build_parallelism` pool regardless of the GEMM's parallelism setting). Both were caught and retracted before being reported as findings; the final harness varies only pool size as an independent axis and compares GEMM modes strictly within each pool size, matching what the real code actually does. New regression test `kmeans_deterministic_across_pool_sizes` (sized to exceed `gemm`'s `DEFAULT_THREADING_THRESHOLD` so it genuinely exercises multi-threading, not a no-op) asserts byte-identical `CoarseModel.centroids` across pool sizes `{1,2,3,4,8}`. 263/263 tests (1 ignored), drift-check clean, compile-matrix clean all 6 PG versions. **Migration**: `ALTER EXTENSION pg_turbovec UPDATE TO '1.22.1';` is sufficient. No REINDEX — this changes build wall clock only, not the on-disk bytes (centroids/assignment/everything downstream of `train_kmeans` is byte-identical to v1.22.0 for the same input). ## [1.22.0] — 2026-07-04 **Repo cleanup, no functional change.** Prompted by an audit for "silent GUC" traps, unfinished/debugging artifacts, and general repo hygiene before a release. No wire-format change (`MetaPageData::version` stays 5), no REINDEX. ### Removed - **`turbovec.mmap_static_blocked`** — a deprecated no-op GUC since v1.19.0 (it toggled a relfile-mmap fast path that v1.19.0 deleted). Removed after a three-minor deprecation window (v1.19.0 warn → v1.20.0/v1.21.0 still-warning → v1.22.0 remove), per AGENTS.md's SQL-surface-removal policy (two-release minimum). `SET turbovec.mmap_static_blocked = ...` now errors like any other unknown GUC instead of silently no-op'ing. - `.woodpecker/ci.yaml` — an orphaned CI config from before the project moved to Forgejo Actions (last touched at v0.3.0, never referenced by any current doc or actually run). It was also the only place `cargo fmt --all -- --check` was ever wired up, which is how 244 formatting violations accumulated across the tree without CI ever catching them (see below). - `relfile_mmap_static_round_trip_matches_buffer_manager` (test) — compared the mmap read path against the buffer-manager fallback; meaningless now that there is only one read path. Also incidentally wrote a stray debug file (`/tmp/pg_turbovec_phase_r3_smoke.txt`) on every run. ### Fixed - **`cargo fmt`'d the whole tree** (244 pre-existing violations, purely mechanical/cosmetic — no behavior change). `fmt-check` is now wired into `.github/workflows/test.yml` and `.githooks/pre-push` so this can't silently reaccumulate. - **Literal `\uXXXX` escape-sequence artifacts** (e.g. `\u2014` instead of an actual em dash) in, , `src/extras.rs`, `src/index/cost.rs`, and the now-removed `.woodpecker/ci.yaml` — cosmetic (doc comments, not code behavior), but a real artifact of a write-tool double-escaping bug worth stamping out repo-wide rather than file-by-file. - A stale dead-code compiler warning: `highdim_oversample_recovers_ recall`'s unused `lists: i64 = 141` local (the `CREATE INDEX` right below it hardcoded the literal `141` instead of interpolating the variable). - `src/guc.rs`'s own module-doc GUC table was missing `turbovec.search_k` and `turbovec.probes` — two of the most-used GUCs in the extension — and had the wrong range for `turbovec.cache_size_mb` (documented as `1..=65536`; the actual registered range is `0..=65536`, and 0 has real, documented meaning: it disables caching). Fixed in `src/guc.rs` and propagated to `docs/ARCHITECTURE.md` §9's GUC table, which had drifted to only 6 of the 17 real GUCs. - `README.md`'s "Operations note: shared_buffers" section still described the v1.5.0–v1.18.x mmap-era guidance ("1.5× the index size is no longer required", "shared_buffers size no longer bounds warm-scan latency") as current fact. It's the opposite of current reality since v1.19.0 removed mmap: `shared_buffers` sizing matters again, and pg_turbovec's 7–15× compression is what makes fitting the hot index in `shared_buffers` achievable. Rewritten to describe the actual current (buffer-cache-only) read path. - `docs/BUFFER_CACHE_ONLY_DESIGN.md` was never actually committed to git despite describing a change that shipped in v1.19.0 (it sat untracked in the working tree for multiple sessions). Committed now with its status header corrected from "DESIGN" (proposal) to "IMPLEMENTED" (it already is, and has been since v1.19.0). ### Documentation-only, for context - Added an explicit warning to `turbovec.bit_width_default`'s GUC description: the name is `turbovec.bit_width_default`, not `turbovec.bit_width` — PostgreSQL silently accepts `SET turbovec.` as a no-op placeholder custom GUC when the name doesn't match one this extension actually registered (a generic PostgreSQL behavior, not a pg_turbovec bug), so a typo'd `SET turbovec.bit_width = N` neither errors nor does anything. A benchmark driver script hit exactly this during the v1.21.0 Phase G-1 validation (see that release's CHANGELOG entry) — every "bw=2" row in the original Phase G-0 results was silently built at bw=4. Use the `bit_width` **index reloption** (`WITH (bit_width = N)`) to set it at `CREATE INDEX` time. ### Migration **No REINDEX.** Wire stays v5; no SQL surface change besides the removed deprecated GUC. `ALTER EXTENSION pg_turbovec UPDATE TO '1.22.0';` is sufficient. ## [1.21.0] — 2026-07-03 **Phase G-1: centroid graph for sublinear IVF coarse-cell selection.** (Gated in by the finding that the IVF-vs-HNSW latency gap at SIFT-1M didn't clear the bar for a full corpus graph, so this release attacks the coarse-probe cost instead.) In-memory / scan-path only; **no wire change** (`MetaPageData::version = 5`), one new GUC, **no REINDEX**. ### Added - **Centroid graph coarse-cell selection.** For an out-of-core IVF index (`turbovec.out_of_core` cell-scoped path) with `lists >= 4096`, `coarse_probe` can now navigate a small fixed-out-degree (16) undirected graph over the coarse centroids — a Vamana/ HNSW-lite greedy beam search — instead of scoring every centroid. The graph is built **once per backend, in-memory**, from the already-persisted coarse centroids (`ivf::build_centroid_graph`); nothing new is persisted, so **existing IVF indexes get it for free on the next scan, no REINDEX**. - **Undirected by construction.** A pure directed k-NN graph (each centroid's own nearest-16 others) can strand a cell that's someone else's close neighbour but has no close neighbours of its own pointing back — a real navigability gap for greedy search that a randomized-corpus test caught during development. `build_centroid_graph` symmetrizes every edge (adds the reverse of each directed edge) before search ever runs, which is what makes the recall-preservation guarantee below actually hold. - **Byte-deterministic.** The directed pass is per-row independent (parallel-safe); the symmetrization pass is a fixed sort+dedup over the full edge list. Same centroids ⇒ byte-identical graph. Verified by `centroid_graph_build_deterministic` (unit) and `ivf_coarse_graph_build_is_deterministic_across_cache_rebuilds` (`#[pg_test]`, end-to-end through the relfile + cache). - **Recall-preserving.** `graph_probe`'s beam width (`ef = max(nprobe*4, 32)`, the classic HNSW-style slack budget) is sized so the graph-navigated result SET matches the exact linear scan's `nprobe`-nearest cells exactly at the tested scales (verified by `graph_probe_matches_linear_scan_exactly`, 150 random queries across 5 `nprobe` values, and the end-to-end `ivf_coarse_graph_matches_linear_scan` `#[pg_test]`). Existing recall-floor tests (`index_am_recall_floor_{2,3,4}bit`) still pass unmodified. - **`turbovec.coarse_graph`** (GUC, enum, default `auto`): `auto` builds/uses the graph only when `lists >= 4096` (below that the plain linear scan is already cheap and a graph's build + per-query overhead isn't worth paying — see `ivf::GRAPH_MIN_LISTS`'s doc); `on` forces it regardless of `lists`; `off` always uses the exact linear scan. `ivf_coarse_graph_auto_falls_back_below_threshold` proves the small-`lists` fallback is correctness-neutral (matches `off` exactly) and that forcing `on` below the threshold still matches too. ### Honest notes - **G-1 is scoped to the out-of-core (`OocIvfIndex`) scan path only.** The whole-load path (`ivf_setup_and_search` in `src/index/scan.rs`) re-reads centroids fresh from the relfile on every scan-open (there is no per-backend cache of that struct today), so building an `O(lists²)` graph there per scan would be pure overhead, not the "build once per backend" amortised cost the plan requires. That path is also gated to comfortably-RAM-resident indexes, i.e. the small-`lists` regime where the linear scan is already cheap — the OOC path is also where `lists >= 4096` (the scale G-1 targets) actually shows up in practice. - **Correcting a docs-drift bug found while implementing this release**: the v1.20.0 CHANGELOG entry below claims a "sublinear two-level coarse quantizer" (`O(lists)→O(√lists)`) shipped in that release. **That was never implemented.** v1.20.0's actual diff (verified against `git show`) only parallelized k-means++ seeding and the build-time assign-sweep, and added `turbovec.scan_parallelism` for the fine-scan. `coarse_probe` remained the plain `O(lists·dim)` linear scan through v1.20.1. There is no `TwoLevelCoarse` type or equivalent anywhere in the v1.20.0–v1.20.1 source tree. v1.21.0 (this release) is the first to actually ship sublinear coarse-cell selection. The v1.20.0 CHANGELOG entry and `docs/UPGRADING.md`'s corresponding row are left as historical record (not rewritten) but this release's `docs/UPGRADING.md` row calls out the correction explicitly. - Measured effect is a rough local sanity check (small-corpus pgrx test host, not a benchmark AVX-512/AVX2 latency benchmark): correctness and recall-preservation are verified by the tests above; a proper before/after p50 comparison at `lists` in the low-to-high thousands (where `auto` actually engages) on an AVX2+ host is follow-up bench work, not part of this patch. ### Migration **No REINDEX.** Wire stays v5; existing v4/v5 IVF indexes benefit from the centroid graph (when `lists >= 4096`) with no rebuild. `ALTER EXTENSION pg_turbovec UPDATE TO '1.21.0';` is sufficient. ## [1.20.1] — 2026-07-03 **CRITICAL PERF FIX: `turbovec.iterative_scan` default flipped `relaxed_order` → `off`.** Wire format unchanged (`MetaPageData::version = 5`); no SQL surface change; **no REINDEX needed** — `ALTER EXTENSION pg_turbovec UPDATE` is sufficient and the new default takes effect on the next backend. ### The bug PostgreSQL's reorder queue (`IndexNextWithReorder` in `nodeIndexscan.c`) can only return a candidate tuple early when the index AM's advertised `ORDER BY` value for that tuple is *exact*. pg_turbovec always advertises `f64::NEG_INFINITY` for every tuple (deliberately opclass-agnostic — correct across L2/cosine/inner- product without per-opclass bounds logic), so that exactness condition can never be satisfied. Under the old default (`relaxed_order`), this forced the executor to drive the AM's own iterative-refill schedule (probe-widening, `search_k` doubling, up to `turbovec.max_scan_tuples` = 20,000 and `turbovec.max_probes` = 64) all the way to completion on **every** `ORDER BY dist LIMIT n` query — no matter how small `n` was — before the executor's reorder queue could signal it was safe to return even the first row. Measured on an AVX-512 a cloud VM host (SIFT-1M/128d IVF, `probes=8`, otherwise-default GUCs): **~2 ms with `turbovec.iterative_scan = off` vs ~900 ms with the old default `relaxed_order`** — a **450x** latency tax paid by every default-configuration KNN query since `relaxed_order` first shipped as the default in v1.8.0. Every benchmark and load-test in this repository explicitly set `turbovec.iterative_scan = off`, which is why this went undetected for eleven releases: the bug only manifests when a caller does NOT override the GUC, and no internal benchmark left it at its default. It was caught while measuring the Phase G-0 IVF-vs-HNSW frontier, which (deliberately) exercised the untouched defaults for the first time. ### Changed - `turbovec.iterative_scan` now defaults to `off` (was `relaxed_order`). `off` matches pgvector's own `hnsw.iterative_scan` default and only under-returns on a *selective* `WHERE` filter combined with `ORDER BY ... LIMIT` — a much rarer shape than the plain unfiltered KNN query this bug taxed. Opt back into `relaxed_order` (`SET turbovec.iterative_scan = relaxed_order;`) if your workload relies on the under-return-avoidance guarantee for selective filters; see `docs/FILTERING.md` and `docs/PRODUCTION.md`. - Added `index_am_iterative_scan_defaults_to_off` regression test (`src/lib.rs`) that asserts the compiled-in default — without an explicit `SET` — caps at `search_k` rather than draining to `max_scan_tuples`, so this can't silently regress back. - Fixed one pre-existing test (`ivf_lists_scan_matches_flat`) that was implicitly relying on the old `relaxed_order` default to guarantee finding an exact self-match under quantization noise; it now opts into `relaxed_order` explicitly, since that's a correctness anchor for k-widening behaviour, not a test of the compiled-in default. - Corrected the documented default in `src/guc.rs`'s module-level GUC table, `docs/PRODUCTION.md`, `docs/FILTERING.md`, `docs/MIGRATING_FROM_PGVECTOR.md`, and `docs/PARITY_GAPS.md`. ### Migration `ALTER EXTENSION pg_turbovec UPDATE TO '1.20.1';` (empty migration file, `migrations/032_pg_turbovec_v1.20.1.sql`). No REINDEX. The new GUC default applies to new backends/sessions; a long-lived backend that already read the old compiled-in default at connection start keeps using it until it reconnects (ordinary PostgreSQL GUC semantics — not specific to this fix). ## [1.20.0] — 2026-07-02 **IVF scaling — parallel build + parallel scan + sublinear coarse quantizer.** Surfaced by the benchmark A/B/C benchmark (the benchmark host, AVX-512). Scan-path / build-path / in-memory only; **no wire change** (`MetaPageData::version = 5`), one new GUC, **no REINDEX** — existing indexes benefit with no rebuild. ### Added - **Sublinear two-level coarse quantizer** (the key scaling enabler). For an IVF index with `lists > 4096`, cell selection is now `O(√lists)` instead of `O(lists)`. The two-level structure is computed **in-memory at index-open** from the already-persisted coarse centroids (a deterministic function of them) — nothing new is persisted, so **existing IVF indexes get it for free on the next scan, no REINDEX**. Measured **39–364× fewer centroid distance computations per query at recall 1.0**, which breaks the coarse-probe wall: `lists` can grow large (tiny cells) without the coarse step becoming the bottleneck — the enabler for IVF at 10M+. - **`turbovec.scan_parallelism`** (GUC, int, default `0` = auto = `min(cores, 4)`; `1` = serial). Parallelizes the per-query IVF fine-scan across probed cells (out-of-core path), cutting single-query latency at high dimension. Conservative default to protect aggregate QPS under concurrency. **Results identical to the serial scan** (same top-k; verified). ### Changed - **Parallel IVF build.** k-means training + the assign-sweep now use the bounded build pool (`turbovec.build_parallelism`) across cores, **memory-bounded and byte-identical** (the reduction order is a fixed function of the input, independent of thread count — so the relfile is reproducible across machines with different core counts). ### Honest notes - The parallel **build** measured only **~1.27×** on 64-core GIST-960d: the single-threaded GEMM (`Parallelism::None`, for bit-exactness) remains the dominant term, and row-blocking around it can't parallelize the GEMM itself. This release makes the parallel build **safe and correct** (no OOM, deterministic) but it is **not** the full build-cliff fix — a deterministic parallel GEMM (or an IVF bit-exactness policy change enabling BLAS threads) is scoped follow-up work. - The high-dim recall ceiling was investigated and found **retrieval-bound** (addressed by the sublinear coarse + more probes), **not** quantization-bound (widening the reorder-rescore pool recovered ~0 recall). - Query-latency-at-scale and 10M-build validation on a cloud VM are pending (the 10M run OOM'd before the memory fix in this release; re-run needed to confirm). ### Migration **No REINDEX.** Wire stays v5; existing v4/v5 indexes benefit from the sublinear coarse + parallel scan with no rebuild. `ALTER EXTENSION pg_turbovec UPDATE TO '1.20.0';` is sufficient. Tests: 241 → 249. ## [1.19.0] — 2026-06-18 **All index reads through the PostgreSQL buffer manager.** Read-path architecture change; **no wire change** (`MetaPageData::version` unchanged), no SQL-surface change, **no REINDEX**. Required for managed/sandboxed Postgres and any environment that restricts direct file access. ### Changed - **Removed the direct relfile `mmap`.** Every byte of index data is now read through PostgreSQL's shared-buffer cache (`ReadBufferExtended`) — there is no `mmap`/`pread` of the relfile. The buffer manager is the single source of truth for page access (consistent pinning/locking; clean crash + streaming-replication semantics). `src/index/mmap_static.rs` is deleted and the `memmap2` dependency dropped (net −700 lines). The buffer-manager readers this routes through already existed (they were the mmap fallback), so this is mostly a deletion. See `docs/BUFFER_CACHE_ONLY_DESIGN.md`. - **Out-of-core (>RAM) IVF serving is preserved without mmap.** The cell-scoped gather (`OocIvfIndex::search_ooc` → `relfile::gather_codes_ranges`) reads **only the probed cells' pages** through the buffer manager, so the per-backend resident set stays `O(probes * cell_size)`, not `O(n)`. Out-of-core serving needs cell-contiguous layout + range-scoped reads — **not** mmap. ### Deprecated - **`turbovec.mmap_static_blocked`** is now a no-op (ignored). It is retained for one minor release so an existing `SET` does not error, and will be removed in a future minor. ### Performance notes - **Warm queries: unchanged** (the prepared index is cached per-backend; warm scans never touch the buffer manager). - **Cold cache-fill on a >`shared_buffers` index: slower** than the old mmap path (per-page lookup/pin/lock + a buffer-manager copy). Mitigation: size `shared_buffers` to hold the hot index — pg_turbovec's 7–15× compression is what makes "the index fits `shared_buffers`" practical where fp32 HNSW could not. For a >RAM index, use IVF + `turbovec.out_of_core` so only probed cells' pages are read. ### Migration **No REINDEX.** Read-path only; wire format unchanged. `ALTER EXTENSION pg_turbovec UPDATE TO '1.19.0';` is sufficient. Tests: 241 (unchanged) on pg16; all 6 PG versions compile; drift-check clean. ## [1.18.0] — 2026-06-18 **Tier-1 IVF latency optimizations (scan-path).** No SQL-surface change, **no wire change** (`MetaPageData::version = 5`; single-vector stays v4), **no REINDEX**. Closes the Tier-1 backlog — narrowing the IVF-vs-HNSW latency gap by attacking the actual per-query floor, with evidence rather than speculation. ### Changed - **Default `turbovec.search_k` lowered 100 → 32 (#1a).** The dominant per-query cost is the executor's reorder-recheck of *every* returned candidate (a heap-tuple fetch + an exact full-precision distance recompute each) — not the vector scan. The new `searchk_recall_frontier` test shows recall@10 **plateaus by `search_k`≈25** (25/50/100/200 identical), so the old default of 100 over-provisioned the recheck ~3× for zero recall gain. The real-corpus `recall_floor_{2,3,4}bit` tests pass at the new default, confirming recall safety. Raise it for `LIMIT > ~20` or a hard corpus; lower it (toward 16) for the lowest latency. ### Added - **`assign_dups_probes_pareto` test + guidance (#2).** Demonstrates that raising `WITH (assign_dups = M)` (soft multi-assignment) lets a query reach a matched recall while probing **fewer cells** (best recall@10 climbs 0.173 → 0.207 → 0.240 as `assign_dups` 1 → 2 → 4 on the test corpus; min-probes-to-matched-recall is non-increasing in `assign_dups`). Opt-in (a build-time layout choice; `assign_dups > 1` needs a REINDEX). The default (1) is unchanged. - Frontier artifacts: `benches/results/searchk_recall_frontier_2026-06-18.json`, `benches/results/assign_dups_probes_pareto_2026-06-18.json`. ### Investigated and rejected / deferred (documented, not built) - **#1b (advertise a tighter ORDER BY distance) — rejected as a no-op.** PostgreSQL's `IndexNextWithReorder` rechecks (heap fetch + exact recompute) every candidate unconditionally under `xs_recheckorderby`, before reading the advertised value; a tighter bound reduces zero work (identical PG 13–18). Documented in `src/index/scan.rs`. - **#3 (SIMD `coarse_probe`) — assessed, deferred:** it is a fixed-floor term, not the dominant cost, and a SIMD horizontal-sum risks cross-ISA reduction-order divergence → recall drift. - **#4–6 — not warranted by the data** (zero effect on the OOC-gathered benchmark; build-when-profiled). ### Honest caveat The **latency confirmation** of #1a/#2 (does p50 actually drop ~3×?) is **deferred to a quiet AVX2 host** — both `floki` and `arnold` were saturated by unrelated work during this release. The **recall safety** of every change is host-independent and verified here; only the latency *number* awaits a quiet window. The projected effect : ~10–13 ms at recall@10≈0.96 at 500k–1M, matching HNSW ef40–ef100. ### Migration **No REINDEX.** Scan-path / default-tuning only; wire stays v5. `ALTER EXTENSION pg_turbovec UPDATE TO '1.18.0';` is sufficient (the new `search_k` default applies to new sessions). Tests: 239 → 241. ## [1.17.1] — 2026-06-18 **ColBERT recall win confirmed cross-domain.** Docs + bench-results release; **no source, SQL-surface, or wire change** (`MetaPageData::version = 5`, single-vector still v4); **no REINDEX**. ### Confirmed The Phase F-2 index-native ColBERT recall gain (shipped v1.17.0) was **replicated on a second, out-of-domain corpus** (BEIR/NFCorpus, 3,633 docs, medical/nutrition, entity-heavier), exercising the persistent `vec_colbert_ops` index on `floki` (AVX2): - **+0.044 nDCG@10 / +0.037 Recall@10** vs the Phase-D pooled+rerank baseline at the value operating point (`candidate_n=256`), rising to **+0.065 nDCG at low candidate budget** (`candidate_n=128`, where the pooled baseline collapses to 0.220 while `colbert_search` holds at 0.285) — same sign, same mechanism, same low-budget shape as the SciFact gate at *every* config. - Quantization signal intact (2-bit ≈ 4-bit, ≤0.0001 nDCG; 2-bit index = 43 MB). - The **persistent index built cleanly** (561k token slots, 42 s / 43 MB at 2-bit, no OOM) and **served from disk — the F-1 ~28 MB/call backend-RSS leak is gone** (RSS plateaus flat at ~360 MB; ~1.4 KB/call warm). **The qualified GO is upgraded to an established cross-domain recall win.** Data: `benches/results/colbert_f2_confirm_floki_nfcorpus_20260618.json`; harness: `benches/scripts/colbert/`. ### Docs and updated to record index-native late interaction as **DONE** (was a future phase): pg_turbovec is one of two PostgreSQL extensions (with VectorChord) with index-native multivector/MaxSim, and the only one also 7–15× smaller than HNSW. ### Migration **No REINDEX.** Docs + bench only; wire stays v5 (single-vector v4). `ALTER EXTENSION pg_turbovec UPDATE TO '1.17.1';` is sufficient. Tests unchanged (239). ## [1.17.0] — 2026-06-18 **Phase F-2 — persistent index-native ColBERT late interaction.** New index kind; **additive wire bump 4 → 5** (single-vector indexes stay **byte-identical to v4**); **no REINDEX** for any existing index. This makes pg_turbovec one of only two PostgreSQL extensions (with VectorChord) to offer index-native multivector/MaxSim — and the only one that is also 7–15× smaller than HNSW. ### Added - **Persistent ColBERT token index.** `CREATE INDEX ON docs USING turbovec (tokens vec_colbert_ops)` over a `turbovec.vector[]` column (per-doc token arrays) builds a v5 on-disk **token** index: `ambuild` unnests each doc's `vector[]` into per-token slots (the doc's heap TID repeated per token — the IVF soft-assign synthetic-slot-id machinery), laid out IVF cell-contiguous and spilled via the Phase B-4 BufFile (n_tokens ≫ n_docs, so the spill is load-bearing). Determinism: tokens are unnested in array order. - **`turbovec.colbert_search` now reads the persistent index.** It locates the `vec_colbert_ops` index on the token column and runs stage-1 candidate generation against the on-disk relfile (warm cache or cold read) instead of rebuilding a backend cache every call. Stage-2 still exact-MaxSim-reranks heap tokens by ctid. **The F-1 ~28 MB/call backend-RSS leak is eliminated on the persistent path** (and bounded on the no-index fallback). - **`vec_colbert_ops`** operator class over `turbovec.vector[]` — support function `max_sim`, **no order-by operator**, so the planner can never select a ColBERT index for `ORDER BY` (the forbidden `amrescan` scan-key path is untouched). A ColBERT index ERRORs on an `ORDER BY` scan with a HINT to use `turbovec.colbert_search`. - **VACUUM** reuses the IVF tombstone path unchanged: a deleted doc's TID marks all its token slots dead (the many-slots-one-TID shape is identical to IVF soft-assign dups); cells stay contiguous (tombstone, never swap-remove); `colbert_search` masks tombstoned slots. ### Wire format (additive v5, per index kind) A new `kind` byte at page offset 30 (formerly a reserved zero) discriminates `KIND_SINGLE` (0, single-vector, wire v4) from `KIND_COLBERT` (1, multivector, wire v5). A single-vector build never sets it, so it emits **wire version 4 + kind 0 — byte-identical to v1.16.0** (guarded by `single_vector_still_emits_v4_bytes` and `v4_single_vector_index_byte_identical`). A v4 meta decodes as `kind = KIND_SINGLE`, so `is_legacy_v4()` never trips. `EXPECTED_WIRE_FORMAT_VERSION` is now 5. ### Migration **No REINDEX.** Existing single-vector indexes are byte-identical and read unchanged under the v5 binary. A ColBERT index is a brand-new shape, built fresh (no in-place conversion from a single-vector index). `ALTER EXTENSION pg_turbovec UPDATE TO '1.17.0';` registers the new opclass and is sufficient. See `docs/UPGRADING.md`. ### Tests 230 → 239 (+`colbert_persistent_build_and_search`, `colbert_persistent_recovers_single_token_match`, `colbert_persistent_survives_vacuum`, `colbert_persistent_deterministic`, `colbert_index_rejects_orderby_scan`, `v4_single_vector_index_byte_identical`, + 3 page.rs unit tests). All six PG versions (13–18) compile; drift-check clean (after the VERSION-5 / minor-bump pairing this release provides). ## [1.16.0] — 2026-06-17 **Phase F-1 — index-native late interaction (ColBERT stage-1).** Additive SQL function in the `turbovec` schema; **no wire change** (`MetaPageData::version = 4`), **no index-AM change**, **no REINDEX**. Closes the last acknowledged feature gap vs Qdrant/VectorChord at the level the analysis showed actually matters (stage-1 recall) —. ### Added - **`turbovec.colbert_search(rel, id_col, token_col, query vector[], k, per_token_k = 64, candidate_n = 256, bit_width = 4)`** (`src/colbert.rs`) — the index-accelerated stage-1 of ColBERT late interaction. Stage 1 builds a **backend-cached flat token index** (one slot per token across all docs, doc-id repeated; synthetic unique slot-ids fed to `IdMapIndex`, real doc-ids kept separately — the IVF soft-assign trick), batch-searches all `|Q|` query tokens, and unions the hit doc-ids into a candidate set. Stage 2 reads each candidate's full token array from the heap and scores it with the exact `max_sim` kernel (Phase D). Returns the top-`k` documents. - **The value over the Phase D pooled-vector + `max_sim` re-rank pattern is stage-1 recall:** a document is retrieved by its **best single token**, not its pooled mean — so a doc whose pooled vector is far but which has one token near a query token (the entity/rare-term/long-doc case ColBERT is built for) is still found. Proven by the `colbert_search_recovers_single_token_match` test. ### What it is / isn't The token index lives **only in the backend cache** (the `turbovec.knn` model) — there is **no relfile, no `CREATE INDEX`, and no wire-format change**. It is the index-native *stage 1* over `max_sim`'s exact *stage 2*; it is **not** the full persistent multivector index AM (per-token relfile + MaxSim-aware scan + PLAID pruning). That persistent AM (Phase F-2) is **gated** on a measured recall/latency win over this F-1 path on a real ColBERT corpus — the plan explicitly refuses to build a 32–512×-larger persistent index on faith. Tuning: `per_token_k` / `candidate_n` trade recall for work; raise `per_token_k` under heavy (2–3 bit) token quantization. ### Migration Additive function; no wire change, no REINDEX. `ALTER EXTENSION pg_turbovec UPDATE TO '1.16.0';` is sufficient. Tests: 224 → 230 (+`colbert_search_basic`, `colbert_search_recovers_single_token_match`, `colbert_search_matches_bruteforce_maxsim`, `colbert_search_empty_query`, `colbert_search_deterministic`, `colbert_search_rejects_bad_k`). ## [1.15.1] — 2026-06-17 **Cross-version build fix (pg13 / pg14 / pg15 / pg18).** Build-only patch; **no wire change** (`MetaPageData::version = 4`), **no behaviour change** on the versions that already compiled (pg16 / pg17), **no REINDEX**. ### Fixed - The Phase B-4 out-of-core IVF build (v1.12.0) called `pg_sys::BufFileReadExact` (PG16+ only) and passed `BufFileWrite` a `*const` pointer (PG13–15 declare it `*mut`). The extension compiled on pg16/pg17 — the local dev target — but **failed to compile on pg13, pg14, pg15, and pg18**, silently breaking those CI matrix legs from **v1.12.0 through v1.15.0**. Now both calls go through version-gated shims (`buffile_write` / `buffile_read_exact` in `src/index/build.rs`), backing the read with the universally-present `BufFileRead` + an explicit short-read check. All six PG features (13–18) compile again. ### Added (CI hardening) - **`scripts/compile-matrix.sh`** — `cargo check`s every `pgNN` feature in `Cargo.toml` (compile-only, ~20s each, no test cluster), so version-specific C-API breaks are caught locally before tagging. Wired into `.githooks/pre-push` alongside `drift-check.sh`. Skips via `COMPILE_MATRIX_SKIP=1` on hosts without every pgrx toolchain. This is the gate that would have caught the v1.12.0 regression; `cargo pgrx test pg16` alone never could. ### Migration **No REINDEX.** Build-only; wire stays v4; pg16/pg17 runtime unchanged. `ALTER EXTENSION pg_turbovec UPDATE TO '1.15.1';` is sufficient. Tests: 224 (unchanged) on pg16; the fix is verified by all six PG features compiling. ## [1.15.0] — 2026-06-17 **Phase C follow-up — operator-path allowlist on flat + IVF.** Additive GUC + function in the `turbovec` schema; **no wire change** (`MetaPageData::version = 4`), **no index-AM scan-key rewrite**, **no REINDEX**. Brings the in-kernel allowlist pushdown (previously `turbovec.knn()`-only, flat-only) to the `ORDER BY emb <=> q LIMIT k` operator path. ### Added - **`turbovec.allowlist`** (session string GUC, default `""`) — a CSV of heap TIDs (encoded as bigint). When set, the index-AM scan ANDs the allowed slots into the slot mask it hands the SIMD kernel, so the kernel short-circuits 32-vector blocks with no allowed slot before any LUT work — the same in-kernel block-skip `knn(..., allowed)` gets, now on the operator path. On an **IVF** index the allowlist is ANDed with the probed-cell mask, scoping the skip to *probed cells ∧ allowed slots*; the out-of-core cell-scoped path gets it too. Empty/unset = exact prior behaviour with **zero added hot-path cost** (no slot-bool is ever built). Parsed once per scan (refills reuse it); a non-integer token ERRORs the scan. - **`turbovec.tid_to_bigint(tid) -> bigint`** — the ergonomic encoder for building the allowlist from `ctid` (returns the `(block << 32) | offset` value the AM stores per slot), so users never hand-write the bit-twiddling. Verified bit-identical to the raw encoding (`tid_to_bigint_matches_raw_encoding`). ### Notes / honest limitation The allowlist is a set of **heap TIDs, not an `id` column** — the index AM keys vectors by heap TID, never a heap `id` column; `turbovec.knn(..., allowed)` remains the id-column path. This is a **pre-materialized id-set** channel, **not** arbitrary-`WHERE` pushdown (which would require scan-key reinterpretation — the forbidden `amrescan` rewrite — or payload columns in the index). See `docs/FILTERING.md` §§ 3.5, 6, 7. Composes with tombstones (a vacuum-deleted row is excluded even if allowlisted) and with `probes >= lists` (exact over the allowed set); returns the same rows as `knn()` for the same id-set. ### Migration Additive GUC + function; no wire change, no REINDEX. `ALTER EXTENSION pg_turbovec UPDATE TO '1.15.0';` is sufficient. Tests: 215 → 224 (+`allowlist_guc_restricts_ordered_scan_flat`/`_ivf`, `allowlist_guc_matches_knn`, `allowlist_guc_empty_is_unfiltered`, `allowlist_guc_composes_with_tombstones`, `allowlist_guc_probes_all_exact`, `allowlist_guc_rejects_bad_token`, `allowlist_guc_out_of_core`, `tid_to_bigint_matches_raw_encoding`). ## [1.14.0] — 2026-06-17 **Phase D — breadth parity (multivector + hybrid fusion).** Additive SQL surface in the `turbovec` schema; **no wire-format change** (`MetaPageData::version = 4`), **no index-AM change**, **no REINDEX**. Closes the multivector / hybrid-fusion breadth gap vs VectorChord / Qdrant at the SQL layer. ### Added - **`turbovec.max_sim(vector[], vector[])` / `max_sim_cosine(...)`** (`src/hybrid.rs`) — ColBERT-style late-interaction MaxSim: `sum_{q in Q} max_{d in D} sim(q, d)` over per-token `vector[]` arrays. `max_sim` uses dot-product similarity (correct for L2-normalised tokens); `max_sim_cosine` uses cosine similarity (`1 - cosine_distance`). All token vectors across both arrays must share one dimension (ERROR on mismatch); an empty query or empty doc scores `0.0` (ColBERT convention). This is a **re-rank** primitive (ANN-retrieve candidates on a pooled vector, MaxSim-rerank the top-N) — the token arrays are not indexed, and index-native late interaction remains a documented future phase. - **`turbovec.rrf_score(rank integer, k integer DEFAULT 60)`** (`src/hybrid.rs`) — reciprocal rank fusion term `1.0 / (k + rank)` for fusing a dense ANN ranking with a sparse / keyword ranking. Pairs with the documented CTE recipe; non-positive denominator raises ERROR. - **`docs/HYBRID_SEARCH.md`** — the canonical breadth guide: multivector MaxSim re-rank (signature, conventions, the two-stage retrieve-then-rerank pattern, the honest index-native limitation), the dense+sparse RRF recipe (full `ROW_NUMBER()` + `rrf_score` CTE for both full-text and `sparsevec`), and the named-vector multi-column schema pattern. - Cross-links from `README.md`, `docs/PRODUCTION.md`, `docs/PARITY_GAPS.md`, and `docs/MIGRATING_FROM_PGVECTOR.md`; the multivector / hybrid rows now read "SQL surface SHIPPED; index-native late interaction is a future phase." ### Notes - Out-of-core BUILD (roadmap Phase D-3) already shipped in v1.12.0 (streaming IVF build); no re-implementation. - Named vectors (multiple vector columns per row) are a documented schema pattern, not new code. ### Migration Additive SQL functions only; no wire change, no REINDEX. `ALTER EXTENSION pg_turbovec UPDATE TO '1.14.0';` is sufficient (the new `turbovec.max_sim` / `max_sim_cosine` / `rrf_score` functions are created by the update script). Tests: 203 → 215 (+`max_sim_basic`, `max_sim_dim_mismatch_errors`, `max_sim_empty`, `max_sim_cosine_normalised`, `max_sim_rerank`, `rrf_score_values`, `hybrid_rrf_recipe`, plus 5 in-module unit tests). ## [1.13.1] — 2026-06-17 **Phase C — metadata-filtering docs + measured allowlist crossover.** Docs + benchmark release; **no source-logic, SQL-surface, or wire change** (`MetaPageData::version = 4`); **no REINDEX**. Bench-results and documentation only. ### Added - **`docs/FILTERING.md`** — the canonical guide to pg_turbovec's three working metadata-filter mechanisms, with a cardinality×selectivity×corpus decision matrix: 1. **Partial index** (`CREATE INDEX ... WHERE tenant_id = X`) — native PG predicate pushdown; the default for known, low-cardinality filters. 2. **In-kernel allowlist** `turbovec.knn(rel, id_col, vec_col, query, k, bit_width, allowed bigint[])` — true in-kernel pushdown (the SIMD kernel skips 32-vector blocks with no allowed slots before any LUT work), flat-only, for selective per-query id sets. 3. **Iterative scan + `WHERE`** (v1.8.0) — the `ORDER BY emb <=> q LIMIT k` AM path; the executor rechecks the predicate, the AM widens `k`/`probes` (capped by `max_scan_tuples`). Includes the honest limitation: no true in-traversal pushdown on the `ORDER BY` AM path (the index stores only vector codes + TID, no payload columns), and a C-4 design sketch for a future phase. - **Measured allowlist selectivity crossover** (floki, AVX2, 300k× 256-d, 4-bit, k=10): allowlist latency decreases monotonically as the filter tightens (17.9 ms → 0.48 ms, ~37×) while the naive post-filter is flat (~7 ms); crossover at ~7–10% selectivity, up to **14.7× faster at 0.1%**. `benches/allowlist_crossover.rs` + `benches/results/allowlist_crossover_floki_v1_13_0_20260617.json`. ### Fixed (docs drift) -: refreshed v1.10.1/v1.11.0 → v1.13.0; the **>500k IVF build ceiling** and the **>RAM** gaps are now marked **CLOSED** (out-of-core build v1.12.0 + out-of-core query v1.13.0); the metadata-filtering row reflects the three real patterns instead of "post-filter only". - `docs/PARITY_GAPS.md`: added the metadata-filtering row; corrected the stale "Parallel index build | GAP — single-threaded" row (parallel build shipped v1.8.0, `turbovec.build_parallelism`). - `docs/MIGRATING_FROM_PGVECTOR.md`: filtered-ANN section lists all three patterns and links `FILTERING.md`; `knn()` signature matches `src/knn.rs`. - `README.md` + `docs/PRODUCTION.md`: cross-link `FILTERING.md`. - Fixed three pre-existing broken benchmark sources (`concurrent_knn`, `recall_vs_pgvector`, `recall`) that referenced the pre-`d3d468e` `IdMapIndex::new` signature (now returns `Result`); `cargo check --benches` is green again. (`cargo pgrx test` never compiled benches, so they did not gate tests.) ### Migration **No REINDEX.** Docs + bench only; wire stays v4. `ALTER EXTENSION pg_turbovec UPDATE TO '1.13.1';` is sufficient. Tests unchanged (203). ## [1.13.0] — 2026-06-17 **Out-of-core IVF query (>RAM serving)** — an IVF index larger than RAM can now be queried, not just built (v1.12.0). Wire format unchanged (`MetaPageData::version = 4`); **no REINDEX**. Completes the out-of-core arc for the >5M production deployment . ### Added — cell-scoped IVF serving (Phase B-1/B-2) The scan previously loaded the **whole index** into a per-backend cache (`read_full` + a copy of the blocked-codes chain off the mmap), so the resident set was `O(n)` and an index that exceeded RAM could not be served. - **Cell-scoped scan.** The backend now caches only bounded metadata (coarse centroids, cell directory, rotation, codebook, per-slot scales/ids) plus a `MAP_PRIVATE` mmap of the relfile, and per query copies **only the probed cells'** contiguous code ranges off the mmap into a compact throwaway sub-index (cells are contiguous from the build-time permutation). Resident set drops to `O(probes * cell_size + faulted pages)`; hot cells stay in the OS page cache, cold cells fault from disk on demand. - **`turbovec.out_of_core`** (enum `off | auto | on`, **default `auto`**). `auto` goes cell-scoped only when the index codes exceed `0.5 * turbovec.cache_size_mb` — an in-RAM index loads whole (no per-query gather/reblock cost); only a genuinely large index pays the bounded-memory-for-CPU tradeoff. `on` forces cell-scoped; `off` forces the pre-v1.13.0 whole-load. - No wire change, no turbovec fork change (reuses `from_parts_with_prepared_borrowed`). Added `mmap_static::gather_slot_ranges` + a buffer-manager twin for the fresh-index fallback. ### Measured 200k×256-d×4-bit IVF (52 MB on disk): per-backend `VmHWM` whole-load **140.8 MB → cell-scoped 44.1 MB** (~3.2× lower). Under a tight cgroup `MemoryMax`, the whole-load backend was OOM-killed (postmaster recovered cleanly, no corruption) where cell-scoped stayed within bound. Warm p50 **82 ms (whole-load) → 199 ms (cell-scoped)** — the expected per-query reblock cost, paid by `auto` only when the index is too large to keep whole. ### Compatibility Scan-path only. Results identical to the whole-load path (`probes >= lists` still reduces to the exact flat scan; tombstones masked; soft-assign deduped). MVCC backstops (reorder queue + heap visibility) preserved. Flat (`lists = 0`) / vacuum-degraded indexes keep the whole-index load (no cells to scope; still `O(n)`-resident — use IVF for >RAM). ### Migration **No REINDEX.** Scan-path change; wire stays v4. `ALTER EXTENSION pg_turbovec UPDATE TO '1.13.0';` is sufficient. ### Tests 197 → 203 (+`ivf_ooc_results_match_whole_load`, `ivf_ooc_probes_all_equals_flat`, `ivf_ooc_tombstones_masked`, `ivf_ooc_soft_assign_dedup`, `ivf_ooc_installs_cell_scoped_handle`, `ivf_ooc_auto_is_size_aware`). drift-check clean. ## [1.12.0] — 2026-06-17 **Out-of-core IVF build** — IVF indexes can now be built at 1M–5M+ rows on a RAM-constrained host. Wire format unchanged (`MetaPageData::version = 4`, byte-identical relfile); **no REINDEX**. Driven by the >5M production deployment . ### Fixed — the 1M+ IVF build OOM (Phase B-4) The `WITH (lists = N)` build held the **full f32 corpus twice** in RAM (`ivf_flat` ~4 GiB + `perm_flat` ~4 GiB at 1M×1024-d) plus the growing index and GEMM scratch — a ~14 GiB peak that `maintenance_work_mem` did not bound, OOM-killing 1M+ builds on a 31 GiB host. IVF was effectively unbuildable at the production scale. - **Disk spill.** The corpus now spills to a PostgreSQL `BufFile` temp file (in `pgsql_tmp`, respecting `temp_tablespaces` / `temp_file_limit`) during the heap scan, wrapped in a `CorpusSpill` RAII type. Cleanup is double-covered: the resource owner unlinks on (sub)transaction abort (even when `ereport(ERROR)` longjmps past Rust destructors) and `Drop` unlinks on success. - **Three streamed passes**, each bounded by `maintenance_work_mem`: (1) spill + bounded reservoir sample for k-means; (2) GEMM-assign cells over disk-backed row-blocks, keeping only the per-row cell-id array (not a corpus copy); (3) feed the quantizer in cell order by re-reading the spill at permuted offsets in bounded chunks. The full f32 corpus is **never resident**; the only RAM term that scales with row count is the **quantized** `packed_codes` (7–15× smaller than the f32 corpus). - **Measured** (1M×1024-d, lists=1024, 30 GiB host): peak RSS **~14 GiB (OOM) → ~7.1 GiB (completes)**; index 1030 MB, spill ~3.9 GB on disk. **5M projected ~8–10 GiB** — buildable on the 31 GiB production host. ### Determinism / compatibility Byte-identical relfile to a v1.11.x in-memory build for the same input, and **`maintenance_work_mem`-invariant** (TQ+ calibration is fit on a fixed cell-ordered prefix, independent of chunk size). The flat (`lists = 0`) build path is unchanged (already Phase-W streamed). No GUC added — `maintenance_work_mem` is the knob. ### Migration **No REINDEX.** Build-internal; wire stays v4. `ALTER EXTENSION pg_turbovec UPDATE TO '1.12.0';` is sufficient. Existing indexes are unaffected; the benefit applies to the next `CREATE INDEX` / `REINDEX`. ### Tests 193 → 197 (+`ivf_streaming_build_determinism_byte_identical`, `ivf_streaming_build_chunk_size_invariant`, `ivf_streaming_build_bounded_memory_completes`, `ivf_streaming_build_temp_file_cleanup`). drift-check clean. ## [1.11.1] — 2026-06-16 **Bench-results-only release. Wire format unchanged from v1.11.0 (`MetaPageData::version = 4`); no REINDEX.** Zero source change. ### Benchmark — IVF latency frontier vs HNSW + ivfflat (Phase A-2) The honest at-scale measurement (isolated AVX2 on `arnold`, `taskset`-pinned, contention-gated, warm, 300 queries/config, Cohere-wiki 500k×1024-d) answering "does IVF beat/equal HNSW at scale." At recall@10 ≈ 0.96: | engine | config | recall@10 | warm p50 | |---|---|---:|---:| | **pgvector HNSW** | ef=200 | 0.966 | **7.9 ms** | | **pg_turbovec IVF** | lists=707, probes=64 | 0.960 | 18.5 ms | | pgvector ivfflat | probes=100 | 0.978 | 117.4 ms | | pg_turbovec flat (exact) | all cells | 1.000 | 41.4 ms | **Honest verdict:** HNSW wins latency at 0.96 (7.9 vs 18.5 ms, ~2.3×). But IVF is now in HNSW's **order of magnitude** (not the 490× flat-scan gap), **beats pgvector's own ivfflat 3–6×** at every matched recall, beats its own exact flat scan, **wins the ≥0.99 recall tail** (0.99 @ 25 ms via probes=256; this HNSW config never reaches 0.99), and is **7.5× smaller** (518 MB vs HNSW 3902 MB). The earlier ~40 ms projection was pessimistic; real p50 at 0.95 is 18.5 ms. ### Critical finding — 1M IVF build OOMs (motivates Phase B-4) The **1M IVF build OOM-killed the postmaster** (~14 GiB peak on a 31 GiB host): the `lists > 0` build holds the full flat corpus + a permuted copy + k-means scratch — a structural peak `maintenance_work_mem` does not bound. Largest IVF index that built on arnold: **500k**. 1M/5M IVF are **blocked on Phase B-4** (streaming / out-of-core build). The IVF *query* path is unaffected. Files: `benches/results/ivf_frontier_arnold_cohere-wiki_2026-06-16.json`, `docs/BENCHMARKS.md`, . ## [1.11.0] — 2026-06-16 Production hardening for IVF: it now **survives VACUUM** instead of silently degrading, and builds **~7.8× faster**. Wire format stays `MetaPageData::version = 4` (additive); **no REINDEX**. Driven by the >5M production deployment + (Phases A-1, E-2). ### Fixed — IVF survives VACUUM (Phase E-2, the production landmine) An IVF index used to **silently degrade to a flat `O(n)` scan after VACUUM** — swap-remove moved the last vector into the deleted slot, breaking cell contiguity, so `has_ivf()` flipped false and queries fell back to the ~seconds full scan with no operator signal. On a churning multi-million-row index that's a latency cliff. - **Tombstones.** The IVF `ambulkdelete` path now leaves dead slots in place and ORs them into a **persisted per-slot tombstone bitmap** (a new v4-additive relfile chain). No rows move, `n_vectors` and the cell directory are untouched, cells stay contiguous, `has_ivf()` stays true, and the scan keeps cell-restricting. The flat (`lists = 0`) path keeps the unchanged swap-remove. Tombstoned slots are masked out of the initial scan **and every probe-widening refill**, so deleted rows are never returned. - **Observability** for any residual fallback: a throttled scan-time `WARNING` (once per backend per index, with a `HINT: REINDEX`) and a new SQL function **`turbovec.index_is_degraded(regclass) -> bool`**. `write_meta_shrink_in_place` now preserves `lists` and flips an `ivf_degraded` meta flag rather than blanking the IVF identity, so the cliff is detectable. - `docs/PRODUCTION.md` gains an IVF + VACUUM operational section. ### Performance — 7.8× faster IVF k-means (Phase A-1) A 200k×256-d / lists=448 build's k-means training was ~295 s (scalar Lloyd, fixed 25 iters) — prohibitive at 5M+. Now GEMM-batched Lloyd assignment (each iteration's nearest-centroid step is one `V@Cᵀ` cross-term GEMM, single-threaded `Parallelism::None` + exact top-2 scalar tie-break) + convergence early-exit (`KMEANS_TOL = 1e-6`). Training **295 s → 38 s = 7.8×**; per-iteration centroids byte-identical to the scalar path, determinism preserved. Training cost is bounded by the `256×lists` reservoir sample regardless of corpus size, so 5M builds train in low-minutes. Build-internal; no surface or wire change. ### Migration **No REINDEX.** Wire stays v4; the tombstone bitmap and `ivf_degraded` flag are additive — pre-1.11.0 v4 indexes read as not-degraded / no-tombstones. The new `index_is_degraded()` function is registered by `ALTER EXTENSION pg_turbovec UPDATE TO '1.11.0';`. ### Tests 187 → 193 on pg16 (+`ivf_survives_vacuum`, `ivf_tombstoned_rows_not_returned`, `ivf_degradation_is_observable`, fast-k-means + page.rs wire-format coverage). drift-check clean. ## [1.10.1] — 2026-06-16 **Bench-results-only release. Wire format unchanged from v1.10.0 (`MetaPageData::version = 4`); no REINDEX.** Zero source-code change. ### Benchmark — IVF warm-p50 on AVX2 Records the AVX2 IVF warm-p50 measurement confirming the IVF cell-skipping **latency** win that `meh` (pre-AVX2 scalar fallback) could not produce. Host `floki` (Intel Core Ultra 7 258V, AVX2), v1.10.0 release build, 200k × 256-d, `lists = 448`, 4-bit, warm cache, 50 timed queries per `probes`: | probes | warm p50 | vs full scan | |-------:|---------:|-------------:| | 4 | 0.74 ms | 5.4× faster | | 16 | 0.78 ms | 5.1× faster | | 448 (= lists, full exact scan) | 3.97 ms | baseline | **At `probes = 16`, ~5× faster than the full exact scan**, on AVX2. The IVF latency win is real on AVX2 hardware. Honest caveat: recall@10 = 1.000 at all probes in this run is an artifact of the synthetic corpus's strong cluster structure, not a general guarantee — the host-independent recall-vs-probes frontier (v1.10.0) is the honest recall/probes trade-off. A full isolated 1M+ × 1024-d sweep on a quiet AVX2 host remains future work. Files: `benches/results/ivf_warmp50_floki_avx2_2026-06-16.json`, `docs/BENCHMARKS.md` ("IVF warm-p50 (AVX2)" section). ## [1.10.0] — 2026-06-16 **Adds the IVF coarse-quantizer layer — a real sublinear ANN structure over the quantized codes.** First wire-format change since v1.4.0 (`MetaPageData::version` 3 → 4), but **existing v3 indexes do NOT need a `REINDEX`**: a v1.10.0 binary reads a v3 index as a flat (`lists = 0`) index. Only users who opt into IVF rebuild. . ### Why The v1.9.1 AVX2 benchmark established that pg_turbovec's flat `O(n·dim)` quantized scan is ~490× slower than pgvector HNSW at 1M×1024-d. IVF partitions the corpus into `lists` Voronoi cells (coarse k-means centroids) and scans only the `probes` nearest cells per query, dropping query work to roughly `(probes/lists)` of the corpus — the architectural path to a competitive latency story while keeping the 10–15× storage win. ### Added — IVF (opt-in) - `WITH (lists = N)` reloption (default 0 = flat / today's exact scan; `N` = number of coarse cells, recommended `≈ sqrt(n)`). - `WITH (assign_dups = M)` reloption (default 1 = single assignment; `M > 1` = soft assignment: boundary vectors stored in their top-M nearest cells to raise recall@10 at a fixed `probes`). - `turbovec.probes` GUC (default 8) — cells scanned per query; the recall/latency dial (the `ivfflat.probes` / `hnsw.ef_search` analogue). `probes >= lists` reduces exactly to the flat exact scan. - `turbovec.max_probes` GUC (default 64) — under `iterative_scan = relaxed_order`, a selective `WHERE` filter that under-returns triggers probe-WIDENING (scan more cells) up to this cap; the `ivfflat.max_probes` analogue. ### How it composes - **Iterative scan** (v1.8.0): refill widens `probes` for IVF indexes (vs growing `k` for flat). - **Oversampling** (v1.9.0): widens the candidate set within the probed cells. - **Reorder queue**: exact-distance recheck, unchanged. - turbovec's SIMD mask SKIPS scan work (block-level early-exit), and cells are stored contiguous, so fewer probed cells = real latency reduction. ### Performance / determinism - The IVF build (k-means + assignment) is GEMM-batched (corpus rotation as `block @ R^T`, cell assignment via a `V @ C^T` cross-term GEMM + scalar top-2 tie-break). Without this the per-vector scalar loops were ~10^12 FLOPs and a 1M build ran 60+ min. Single-threaded `gemm` keeps it bit-deterministic. - Builds are deterministic: same table + same `lists`/`assign_dups` ⇒ byte-identical relfile (seeded k-means++, `IVF_SEED`). - Recall-vs-probes frontier (host-independent; recall is CPU-independent): on a hard random 16k×64-d corpus, `probes=16` → R@10 0.53 skipping 85% of blocks, `probes=lists` → 1.000. Real clustered embeddings reach high recall at far lower probes. Absolute AVX2 warm-p50 latency is deferred to a quiet `arnold` window (`meh` is pre-AVX2). See `docs/BENCHMARKS.md`. ### Migration **No REINDEX for existing (v3, flat) indexes** — they read as `lists = 0` under the v1.10.0 binary. Opt into IVF by rebuilding with `WITH (lists = N)`. `ALTER EXTENSION pg_turbovec UPDATE TO '1.10.0';` registers the new reloptions + GUCs. `MetaPageData::version` is 4; `EXPECTED_WIRE_FORMAT_VERSION = 4`. ### Tests 150 → 185 on pg16 (IVF build/scan/soft-assign/determinism/ probes-frontier coverage; distinct-id assertions throughout). drift-check clean. ## [1.9.1] — 2026-06-15 **Bench-results-only release. Wire format unchanged from v1.9.0 (`MetaPageData::version = 3`); no REINDEX needed.** Zero source-code changes — this release bundles the AVX2 latency-frontier benchmark and the honest positioning correction it produced. ### Benchmark — AVX2 latency frontier on `arnold` The latency numbers `meh` (a pre-AVX2 Xeon) physically could not produce. Run on `arnold` (i9-12900H, AVX2), isolated via `taskset -c 2-5` CPU-pinning to dedicated P-cores with per-batch contention measurement (observed 1-min load ≤ 1.05 throughout, zero CPU steal, no contended batches, no re-runs). Cohere wikipedia 1M × 1024-d, 1000 held-out queries, brute-force exact GT. - **Correctness on the AVX2 SIMD path: recall@10 = 1.000** (the AVX2 kernel, not just meh's scalar fallback, is correct on v0.9.0). - **The hard truth: pg_turbovec loses to HNSW on latency by ~490× at 1M rows** — warm p50 ~2552 ms (flat `O(n·dim)` quantized scan, recall 1.000) vs pgvector HNSW ~5 ms (sublinear graph, recall 0.96). AVX2 makes the correct scan ~15–25× faster than meh's scalar fallback (2.5 s vs 41.6 s), but a 1M-row full scan is seconds, not ms, by design. - **Retracted:** the earlier "26.8 ms on `meh` / we win 2.3× warm p50" claim. That came from the **pre-AVX2 scalar-fallback bug** (fast-but-WRONG, fixed in v1.7.3) and never represented correct behaviour. ### Positioning correction `docs/PARITY_GAPS.md` and updated to the honest scoreboard. pg_turbovec's durable wins are **storage** (10–15× smaller), **exact recall** (1.000 vs HNSW's ~0.96), and **build memory** — NOT query latency at scale. Honest positioning: *"best storage efficiency + exact recall for PG vector search where an O(n) scan fits the latency budget,"* NOT "beat HNSW on every axis." The architectural path to a latency story at scale is an IVF / coarse-quantizer layer (turning the O(n) scan into O(n/nlist + probes)) — a planned future major arc. ### Files - `benches/results/latency_frontier_arnold_cohere_1m_v1_9_0_2026_06_15.json` - `benches/scripts/vectordbbench/sweep_latency_isolated.py` - `docs/BENCHMARKS.md` (arnold AVX2 section) - `docs/PARITY_GAPS.md` (corrections) ## [1.9.0] — 2026-06-15 Oversampling (tunable recall), test-coverage hardening, and the first published head-to-head benchmark. **Wire format unchanged** (`MetaPageData::version = 3`); **no `REINDEX`** — `ALTER EXTENSION pg_turbovec UPDATE TO '1.9.0';` suffices. The one new GUC defaults to a no-op. ### Added — `turbovec.oversample` (differentiator #5) Turns quantization from a fixed accuracy point into a tunable recall lever, matching Qdrant's oversampling / VectorChord's rerank. - `turbovec.oversample` (float, default 1.0, range 1.0..=100.0): the scan fetches `ceil(search_k * oversample)` quantized candidates and the executor's reorder queue (`xs_recheckorderby = true`) trims to the exact top-k. Widening the candidate set recovers true neighbours the lossy quantized ranking placed just outside `search_k`. - No separate rescore path: oversampling + the always-on reorder queue together ARE the rescore mechanism (the reorder queue already re-ranks by exact full-precision distance). Measured: recall@10 climbs 0.81 (oversample 1.0) → 1.0 (oversample 4.0) on a 4-bit / 3000×64 corpus; latency rises ~linearly. - Composes with iterative scan: oversample sets the initial `k`; refill doubles from there, capped by `max_scan_tuples`. - Default 1.0 is identical to v1.8.0 behaviour. ### Testing — scale + distinct-id + recall-floor regression guards The pre-AVX2 wrong-results bug (fixed in v1.7.3) shipped because no test exercised more than ~2000 rows or asserted distinct result ids. Closed those gaps: - Medium-scale (20k×128-d) recall-floor `#[pg_test]` per bit_width {2,3,4}, with a brute-force ground-truth comparison. - `assert_distinct_ids` on EVERY ANN-scan test — the cheapest guard against the whole wrong-ranking bug class (a duplicate-id assert would have caught the pre-AVX2 bug instantly). - `docs/TESTING.md` documenting coverage + honest gaps: CI is AVX2-only (the scalar fallback runs only in turbovec's upstream tests + pre-AVX2-host validation on turbovec bumps); unit tests cap at 20k rows (the benchmark is the large-scale evidence); the 15 "ignored" items are benign ` ```ignore ` doctests. ### Benchmark — first published head-to-head (`docs/BENCHMARKS.md`) Cohere wikipedia 1M × 1024-d (real embeddings, 1000 held-out queries, brute-force GT) vs pgvector HNSW, with a full reproducible harness under `benches/scripts/vectordbbench/`. - **recall@10 = 1.000** on the fixed v1.8.0+ build at every config — the same pre-AVX2 host scored 0.0 on the old v1.7.1 build, so this is the definitive confirmation that the pre-AVX2 fix works on real embeddings at scale. - Storage: pg_turbovec 4-bit **7.6× smaller** (1026 MB vs HNSW 7806 MB), 2-bit **15.2× smaller** (512 MB). Build 1.9–2.1× faster. - pgvector HNSW frontier (its own SIMD): R@10 0.849/9.4 ms (ef40) → 0.979/20.1 ms (ef400). - **pg_turbovec latency frontier is DEFERRED to an AVX2 host.** The bench host `meh` is a pre-AVX2 Ivy Bridge Xeon; turbovec takes its scalar fallback (~1000× slower than its AVX2/AVX-512 kernels: ~42–69 s/query full-corpus scan). EXPLAIN confirmed Index Scan (not a seq-scan artifact). Latency/QPS benchmarks require an AVX2+ host (`arnold`); see `BENCHMARKS.md` for the explicit TODO and full caveats. (Updated `AGENTS.md` bench-host guidance accordingly — SIMD class matters more than RAM for turbovec latency.) ### Migration **No migration; no REINDEX.** On-disk format byte-identical to v1.7.x / v1.8.x. The new `turbovec.oversample` GUC defaults to 1.0 (no-op). `ALTER EXTENSION pg_turbovec UPDATE TO '1.9.0';` resolves against the empty `migrations/014_pg_turbovec_v1.9.0.sql`. ### Tests 142 → 150 on pg16 (+5 oversampling, +3 recall-floor; distinct-id assertions added to existing tests). drift-check clean. ## [1.8.0] — 2026-06-15 Four competitive-parity features in one minor release. **Wire format unchanged** (`MetaPageData::version = 3`); **no `REINDEX` needed** — `ALTER EXTENSION pg_turbovec UPDATE TO '1.8.0';` is sufficient. All four additions are scan-side, build-side, or additive SQL surface; none touch the on-disk relfile layout. . ### Added — iterative index scan (parity gap #1, the correctness fix) The one true correctness gap vs pgvector. `amgettuple` used to run a single `search_k`-sized batch and return false when drained, so a selective `WHERE filter ORDER BY emb <=> q LIMIT k` silently under-returned (e.g. 3 rows when 10 were asked for) — exactly what pgvector shipped `hnsw.iterative_scan` (0.8.0) to fix. - When the executor exhausts the candidate batch and the filter hasn't been satisfied, the scan re-runs the turbovec search with a **doubled `k`** and feeds the new candidates, capped by a new `turbovec.max_scan_tuples` GUC (default 20000, matching pgvector's `hnsw.max_scan_tuples`). - Controlled by `turbovec.iterative_scan` — an enum GUC `off | relaxed_order` (default `relaxed_order`). `strict_order` is deferred; our existing reorder-queue model (`xs_recheckorderby = true`) already restores exact per-tuple ordering on top of `relaxed_order`. - Dedup across refills via a returned-TID `HashSet` (turbovec's `search` isn't a stable prefix across `k` due to an unstable sort on score ties; the set is robust and bounded by `max_scan_tuples`). - Regression test demonstrates `off` under-returns and `relaxed_order` returns the full `LIMIT`. ### Added — parallel index build (parity gap #2) pgvector parallelises HNSW/IVFFlat builds across `max_parallel_maintenance_workers`; pg_turbovec's `ambuild` was single-threaded. - Option B (rayon): the CPU-heavy `encode` + SIMD-`repack` phases (which dominate build CPU, not the heap scan) are parallelised over heap-scan chunks via a rayon pool. Chunks are processed in heap-scan order then concatenated deterministically. - New `turbovec.build_parallelism` GUC (default 0 = derive from `max_parallel_maintenance_workers + 1`; positive pins the pool). - **Byte-for-byte identical relfiles** regardless of thread count — asserted by a unit test — so the wire format and any reproducibility guarantees hold. Memory stays bounded by the Phase W `maintenance_work_mem` cap. ### Performance — cold-scan latency (parity gap #3) Cold-scan p50 was ~1256 ms (1 M × 1536-d) vs HNSW's ~100 ms. - **Lazy `id_to_slot` on the read path.** Profiling the per-backend cache-fill showed the dominant residual term — once Phase P pre-baked the SIMD-blocked layout and Phase R-2 persisted the rotation — was the `id_to_slot: HashMap` that `IdMapIndex::from_id_map_parts*` builds eagerly (~50 ms at 200 k rows, linear in `n`). The index-AM scan path never reads `id_to_slot` (`search` returns slots, mapped via the `slot_to_id` `Vec`). The scan path now installs a lightweight `cache::ReadOnlyIndex` (no HashMap); the build is deferred to the first `aminsert`/`remove`. A read-only / pooled-connection backend that only scans never pays it. Read-only constructor: ~50 ms → ~0 ms. - Key correctness test `mutation_after_readonly_scan_is_correct` verifies the deferred HashMap builds correctly on first insert. - **Deferred follow-ups** (see `docs/PARITY_GAPS.md` § cold scan): read-path mmap of the codes/scales/ids chains; a header-gap-free on-disk layout for true zero-copy mmap (VERSION 3 → 4, a future minor); a cross-backend DSA/DSM shared cache. ### Added — `||` concat + halfvec arithmetic (parity gap #4) pgvector has `||` concat for vector+halfvec and `+`/`-`/`*` for both; pg_turbovec had `+`/`-`/`*` for `vector` only. - `turbovec.vector || turbovec.vector -> vector` (concat) - `turbovec.halfvec || turbovec.halfvec -> halfvec` (concat) - `turbovec.halfvec` `+`/`-`/`*` element-wise (Hadamard for `*`) - Matches pgvector overflow semantics (error on non-finite result) and dim-mismatch errors. ### Migration **No migration needed; no REINDEX.** The on-disk relfile format is byte-identical to v1.7.x. Drop in the new shared library, restart, scan; existing indexes work unchanged. The new GUCs default to the pgvector-equivalent behaviour (`iterative_scan = relaxed_order`). `ALTER EXTENSION pg_turbovec UPDATE TO '1.8.0';` resolves against the empty `migrations/013_pg_turbovec_v1.8.0.sql`. ### Tests 123 → 142 on pg16 (+19: iterative-scan, parallel-build, cold-scan, and arithmetic-parity coverage). drift-check clean. ## [1.7.3] — 2026-06-15 ### Fixed — pre-AVX2 x86_64 wrong-results bug (turbovec fork → v0.9.0) Wire format unchanged from v1.6.0 / v1.7.x (`MetaPageData::version = 3`); **no `REINDEX` needed** to upgrade. - **Root cause.** The Phase A1 "regression" (index `ORDER BY emb <=> probe LIMIT N` returning the same `id` N times at 10 M scale on the `meh` bench host) was traced to an **upstream turbovec kernel bug, not pg_turbovec**. The pinned turbovec v0.7.0-era fork (`6e80a59`) had a scalar fallback that, on x86_64 CPUs **without AVX2**, read the perm0-interleaved (FAISS-style) SIMD code layout as if it were sequential — producing silently-wrong / repeated top-k. `meh` is an Intel Xeon E5-2697 v2 (Ivy Bridge, 2013): `avx` but no `avx2`, so it hit the buggy path. AVX2 (Haswell 2013+), AVX-512, and ARM NEON hosts always took a correct SIMD path — which is why the bug never reproduced on AVX2 dev boxes or `arnold`, only on `meh`. - **Fix.** Upstream turbovec fixed this in PR #108 (issue #106, "V5"), released in v0.8.0, adding a correct `score_query_into_heap` x86_64 scalar fallback plus a `FORCE_SCALAR_FALLBACK` regression test. v1.7.3 upgrades the `gburd/turbovec` fork from the v0.7.0-era `6e80a59` to a fork rebased onto upstream **v0.9.0** (`d3d468e` on branch `pg_turbovec-integration-v0.9.0`). - **Also brought in, inert here:** - TQ+ per-coordinate calibration fields, constructed as identity (empty) on the relfile path — **no recall change, no wire change** in v1.7.3. Persisting them for a recall gain is a future minor release (VERSION 3 → 4 + REINDEX). - Security hardening: `MAX_DIM = 65536`, NaN/Inf/huge-magnitude input rejection, checked-mul `.tv`/`.tvim` loaders. - **Zero pg_turbovec source churn** — the upgrade is a `Cargo.toml` rev bump only; the fork kept the `prepare_eager` alias and passes TQ+ through internally so every `from_id_map_parts*` call site is unchanged. - **Toolchain note.** turbovec v0.9.0 uses `avx512` `target_feature`s requiring **Rust ≥ 1.89**. Builds with the default `stable` toolchain (1.95). See `AGENTS.md` for the refreshed openblas store path and the `-fuse-ld=bfd` linker note. - Tests: 123/123 on pg16. drift-check clean. ### Migration **No migration needed; no REINDEX.** The on-disk relfile format is byte-identical to v1.6.x / v1.7.x. Drop in the new shared library, restart, scan. **Pre-AVX2 x86_64 users specifically** should upgrade to clear the wrong-results bug and can drop any `SET enable_indexscan = off;` workaround. `ALTER EXTENSION pg_turbovec UPDATE TO '1.7.3';` resolves against the empty `migrations/012_pg_turbovec_v1.7.3.sql`. ## [1.7.2] — 2026-05-27 ### Added — Phase Y: automated upgrade-matrix validation Wire format unchanged from v1.6.0 / v1.7.0 / v1.7.1 (`MetaPageData::version = 3`); **no `REINDEX` needed** to upgrade. v1.7.2 is a test-only patch release. Production-confidence foundation: previously the upgrade matrix in `docs/UPGRADING.md` and the `is_legacy_v{1,2}()` detection predicates in `src/index/page.rs` were promises with no automated end-to-end test. Phase Y closes that gap. - **`alter_extension_path_140_to_171_runs_clean`** (new `#[pg_test]`) replays every `migrations/0NN_pg_turbovec_*.sql` from v1.3.0 onward against the live test cluster. Catches a release engineer who lands a syntactically-broken DDL change in one of the post-v1.3 migration files (which are otherwise intentionally empty). - **`ambeginscan_errors_on_legacy_v1_meta`** and **`ambeginscan_errors_on_legacy_v2_meta`** (new `#[pg_test]`s) build a real v1.7.2 index, forge the meta-page version byte to 1 or 2 via the new cfg-gated `relfile::force_meta_version()` helper, and assert that `ambeginscan` ERRORs at first scan with the documented primary message + `REINDEX INDEX` HINT. Exercises the Phase Q (v1.3.0) + Phase R-2 (v1.4.0) hard migration boundaries without having to keep pre-v1.4 binaries around. - **`alter_extension_update_chain_resolves`** (new `#[pg_test]`) asserts the installed extension version matches `Cargo.toml`, catching version-number drift between `Cargo.toml`, `pg_turbovec.control`, and the migration file naming convention. - **`migration_files_cover_documented_versions`** (new `#[pg_test]`) asserts the set of `migrations/*.sql` sigils matches the documented release history. If you tag a new release without adding the migration file, this test fails before the bad tag escapes CI. - **`scripts/drift-check.sh` § 9** (new check) cross-checks `migrations/*.sql` against the `From` column of the migration matrix in `docs/UPGRADING.md`. Catches release-time drift between adding a tag and forgetting to add the migration file. - **`relfile::force_meta_version()`** (new test-only helper) is gated on `cfg(any(test, feature = "pg_test"))` and patches the version byte of the meta page in place via a `GenericXLog` record. Only the pgrx test suite (and a future feature-gated build) can reach it; production builds never compile it. ### Migration **No migration needed; rebuild not required.** The on-disk format is byte-identical across v1.6.0 / v1.7.0 / v1.7.1 / v1.7.2. Drop in the new shared library, restart, scan; existing indexes built under any of these versions continue to work unchanged. ## [1.7.1] — 2026-05-27 ### Reverted — Phase W-2 split-write design (regression) Wire format unchanged from v1.6.0 / v1.7.0 (`MetaPageData::version = 3`); **no `REINDEX` needed** to upgrade or downgrade between any of these. v1.7.1 is a behaviour-only revert. - **Phase W-2 (v1.7.0) reverted.** Validation on `meh` (24-core, 125 GiB RAM NixOS host, head commit `a289870`) at 10 M × 1536-d × 4-bit showed the split-write `ambuild` path introduced in v1.7.0 made the build **53% slower** (5052 → 7748 s), used **2.7 GiB of swap** (vs 0 in v1.6.0), and slightly **raised** peak RSS (22.5 → 23.04 GiB). The predicted ~15 GiB peak never materialised. Full data: `benches/results/phase_w_2_validate_meh_10m_2026_05_27.json`. | metric | v1.6.0 | v1.7.0 (W-2) | v1.7.1 (revert) | |-------------------|-------:|-------------:|----------------:| | Peak RSS (GiB) | 22.5 | 23.04 | 22.5 (= v1.6.0) | | Swap used (GiB) | 0 | 2.67 | 0 (= v1.6.0) | | Build time (s) | 5,052 | 7,748 | 5,052 (= v1.6.0)| - **Why Phase W-2 didn't work.** The hypothesis was that dropping the ~7.7 GiB row-major `packed_codes` Vec mid-finalise (via `IdMapIndex::take_packed_codes()`) would shave the peak RSS by ~7.7 GiB. It didn't, because the intervening `write_packed_phase` pins those bytes in `shared_buffers` before `take_packed_codes()` runs, and `ps -o rss` counts mapped shared memory as part of the backend's resident set. The 7.7 GiB of "freed" heap simply migrated to pinned shared memory; same RSS budget, plus the cost of an extra `GenericXLog` flush phase. See .7.1" for the full analysis. - **What was reverted.** - `src/index/relfile.rs::write_full_inner` — restored to the v1.6.0 single-pass batched-`GenericXLog` flow: meta page, then codes / scales / ids chains, then blocked / rotation chains, then `RelationTruncate` for shrinking REINDEX. - `src/index/build.rs::ambuild` — restored to the v1.6.0 sequence: `prepare_eager()` first, then a single `write_full_with_prepared` call. The `take_packed_codes()` call is dropped from this code path. - `src/lib.rs` — `ambuild_drops_packed_codes_before_blocked_write` renamed to `ambuild_round_trip_after_phase_w_2_revert` and kept as a generic ambuild round-trip smoke (still passes via the v1.6.0 code path). - **What was kept.** - `relfile::write_packed_phase`, `relfile::write_blocked_phase_and_meta`, and `relfile::PackedPhaseLayout` remain in the source as parked dead code, marked `#[allow(dead_code)]`. They have no callers after the revert but may be useful for a future Phase W-3 attempt that takes a different angle (e.g. streaming `pack::repack`). - The turbovec fork pin at rev `6e80a59f473292cc9e04d575ba1596f3e23321c5` (turbovec 0.7.0) stays. `IdMapIndex::take_packed_codes()` on the fork is harmless additive API; we just don't call it. ### Migration **No migration needed; rebuild not required.** The on-disk format is byte-identical across v1.6.0 / v1.7.0 / v1.7.1. Drop in the new shared library, restart, scan; existing indexes built under any of these versions continue to work unchanged. ## [1.7.0] — 2026-05-27 ### Added — mid-finalise drop of `packed_codes` in `ambuild` (Phase W-2) Wire format unchanged from 1.6.x (`MetaPageData::version = 3`); **no `REINDEX` needed** to upgrade. v1.7.0 is a build-side change only: the on-disk index format is byte-identical to v1.6.x. - **Reorder finalisation writes so packed_codes and blocked are never co-resident.** Phase W (v1.6.0) capped the heap-scan staging buffer, dropping peak `ambuild` RSS from 121 GiB to 22.5 GiB at 10 M × 1536-d on `meh`. The remaining 22.5 GiB peak was `IdMapIndex`'s row-major `packed_codes` (~7.7 GiB) plus the SIMD-blocked derived layout (~7.5 GiB) plus allocator slack + GenericXLog page-assembly buffers, all alive during the single-call `relfile::write_full_with_prepared` flush. Phase W-2 splits that call into two phases: 1. `relfile::write_packed_phase` streams `packed_codes`, `scales`, and `slot_to_id` to relfile pages while `packed_codes` is the only large in-memory Vec. 2. `IdMapIndex::prepare_eager()` materialises the SIMD-blocked layout, codebook, and rotation matrix (transient peak: packed + blocked). 3. `IdMapIndex::take_packed_codes()` (new turbovec 0.7.0 API) swaps the row-major Vec out and `shrink_to_fit`s it; the `OnceLock`-backed blocked cache is unaffected. 4. `relfile::write_blocked_phase_and_meta` streams the blocked + rotation chains and stamps the meta page LAST. Expected peak at 10 M × 1536-d: **~15 GiB** (down from 22.5 GiB). Combined with Phase W's 121 → 22.5 GiB cut, that's an **8× total reduction vs pre-Phase-W**. Validation on `meh` at 10 M scale is a follow-up bench phase; the v1.7.0 code change ships with local unit-test coverage of the split write (`ambuild_drops_packed_codes_before_blocked_write`). - **Meta page is now written LAST.** `write_full_inner` used to write the meta page first and the chains second, which left a crash window where block 0 referenced not-yet-written chain pages. v1.7.0 routes both the legacy `write_full` / `write_full_with_prepared` and the new split path through `write_blocked_phase_and_meta`, which writes the meta page AFTER all chain pages — matching the standard PG hash/gist AM "meta page is the atomic-complete signal" pattern. A crash before the meta-page WAL record commits leaves block 0 in its previous state (zero-filled for fresh build, previous meta for REINDEX), and `ambeginscan` rejects the index as empty/legacy. No on-disk format change — readers never observed the intermediate state in any released version. - **Turbovec fork bump 0.6.0 → 0.7.0** (rev `6e80a59f473292cc9e04d575ba1596f3e23321c5`, branch `pg_turbovec-integration` on `gburd/turbovec`). Adds `TurboQuantIndex::take_packed_codes(&mut self) -> Vec` and the matching `IdMapIndex::take_packed_codes`. Additive minor; no breaking changes for embedders that don't call the new API. - **Phase W-3 deferred.** The remaining ~15 GiB peak is dominated by the SIMD-blocked Vec materialised by `prepare_eager()` plus GenericXLog page-assembly slack. Dropping the blocked peak further (to ~10 GiB) would require streaming `pack::repack` so the blocked layout never has to be fully resident; that's substantial turbovec internals work and is out of scope for v1.7.0. ### Migration **No migration needed; rebuild not required.** The on-disk format is byte-identical to v1.6.x. Drop in the new shared library, restart, scan; existing indexes continue to work unchanged. ## [1.6.1] — 2026-05-27 ### Bench-results-only release. Wire format unchanged from 1.6.0; no REINDEX needed. - **Phase W validation on `meh` (commit `8efb89c`).** Re-ran the Phase V 10M × 1536-d build against v1.6.0 to confirm the streaming `ambuild` change actually drops peak RSS as designed. Result: **121 GiB → 22.5 GiB peak** (5.4× reduction), **60 GiB → 0 GiB swap usage**. Build time within 0.1 % of Phase V (5048 → 5052 s); index size unchanged (15 GiB); warm-scan p50 identical (21.2 ms vs Phase V's 21–49 ms band). The remaining 22.5 GiB peak is `IdMapIndex`'s row-major `packed_codes` (~7.7 GiB) + the SIMD-blocked prepared layout (~7.5 GiB) + allocator slack + PG backend baseline, all held simultaneously during end-of-build finalisation. Tracked as Phase W-2. Files: `benches/results/phase_w_validate_meh_10m_2026_05_27.json`, `benches/results/phase_w_warm_sanity_meh_10m_2026_05_27.json`, `benches/results/build_tv_meh_10m_v1_6_0_2026_05_27.{log,psql.log,rss.tsv.gz}`, `docs/RECALL.md` § 2.7 follow-up. - **Phase X: RISC-V architecture comparison (commit `a8fbd87`).** First non-x86 host bring-up. 100 k × 384-d synthetic on `rv` (RISC-V 64, 8 cores, 7.7 GiB RAM, Ubuntu 24.04 LTS): index 39 MB (5× compression), build 13.97 s, warm p50 **242.64 ms** (50-query stdev 0.73 ms — extremely tight). Verdict: **arch_works**. The latency multiplier vs x86 (~10–25× depending on corpus comparison) reflects turbovec's AVX2/SSE inner loop falling back to scalar on RISC-V; RVV intrinsics are upstream-future work. Operational note for non-NixOS hosts: the postmaster needs `LD_PRELOAD=libopenblas.so.0` because `cblas_sgemm` is a deferred symbol not in the .so's NEEDED entries. Files: `benches/results/recall_warm_rv_100k_v1_6_0_2026_05_27.json`, `docs/RECALL.md` § 2.8 (new section). ### Migration **No migration needed; rebuild not required.** The on-disk format is byte-identical to v1.6.0. Drop in the new shared library, restart, scan; existing indexes continue to work unchanged. ## [1.6.0] — 2026-05-26 ### Added — streaming heap scan in `ambuild` (Phase W) Wire format unchanged from 1.5.x (`MetaPageData::version = 3`); **no `REINDEX` needed** to upgrade. v1.6.0 is a build-side change only: the on-disk index format is byte-identical to v1.5.x. - **Build-time memory cap.** Phase V measured `CREATE INDEX` peak RSS at **121 GiB** on a 10 M × 1536-d × 4-bit corpus on `meh` (24 cores, 125 GiB RAM), with 60 GiB of swap usage. The dominant offender was `BuildState::flat: Vec` in `src/index/build.rs::ambuild` accumulating the entire heap-scan output before passing it to `IdMapIndex::add_with_ids`. At 10 M × 1536-d that buffer alone is 61 GiB. - **Phase W: stream the heap scan.** `BuildState` now carries two bounded staging buffers (`pending_flat`, `pending_ids`) sized off `maintenance_work_mem`. Every `chunk_rows` rows the callback flushes into `IdMapIndex::add_with_ids` and `shrink_to_fit`s the buffers back to zero capacity, returning the bytes to the allocator. A trailing flush after the heap-scan loop drains the partial chunk. - **Chunk sizing formula** (in `BuildState::compute_chunk_rows`): `chunk_bytes = min(maintenance_work_mem_kb * 1024 * 3 / 4, 1 GiB)`; `chunk_rows = max(chunk_bytes / (dim * 4), 1)`. The GUC is read in **kilobytes** (PG convention; the global is `pg_sys::maintenance_work_mem: c_int` whose unit is KB despite the name). 75% allocation leaves headroom for the IdMapIndex's own growth; the 1 GiB ceiling caps the staging buffer even with a `SET maintenance_work_mem = '8GB'`. - **Expected peak at 10 M × 1536-d:** ~16 GiB (down from 121 GiB). Validation on `meh` at 10 M scale is a follow-up phase — the v1.6.0 code change ships with local unit-test coverage of the streaming path; the multi-hour memory-cap validation runs separately. - **Phase W-2 deferred.** The IdMapIndex still holds `packed_codes` (~7.7 GiB at 10 M × 1536-d × 4-bit) in memory alongside `blocked_codes` after `prepare_eager()`. Dropping it would save another ~7.7 GiB at peak but requires a turbovec fork API change (`IdMapIndex::drop_row_major_codes(&mut self)` on branch `pg_turbovec-integration`). Tracked as a follow-up; out of scope for v1.6.0. - **One new `#[pg_test]`:** `ambuild_streams_heap_scan_under_maintenance_work_mem` exercises the streaming path with `maintenance_work_mem = '4MB'` and a 1000-row table. Test count 116 → 117. - **Docs.** `docs/UPGRADING.md` migration matrix gets a `1.5.x → 1.6.0` no-op row; records the diagnosis, the formula, and the Phase W-2 follow-up parking lot. ### Migration **No migration needed; rebuild not required.** The on-disk format is byte-identical to v1.5.x. Drop in the new shared library, restart, scan; existing indexes continue to work unchanged. `ALTER EXTENSION pg_turbovec UPDATE TO '1.6.0';` resolves against the empty `migrations/007_pg_turbovec_v1.6.0.sql`. This is a **minor** bump rather than a patch because the build-time memory profile is observably different: a host that used to OOM on 10 M × 1536-d will now succeed. That's a behaviour change worth a minor even though no on-disk format changed. ## [1.5.1] — 2026-05-26 ### Bench-results-only release. Wire format unchanged from 1.5.0; no REINDEX needed. - **Phase U-1: cache works correctly.** A debug-only tracepoint in `cache::lookup` confirmed 50/50 hits across a 50-query warm sweep (zero misses of any class). The Phase S agent's hypothesis that the per-backend cache misses on every warm scan was wrong; what they saw in `perf` was the one-shot `finalise_from_inner` build during the cold-cache install, amortised over the sampling window. Tracepoint reverted before the build that produced the Phase U-2 measurements. - **Phase U-2: Phase S delivers no win on RAM-rich hosts.** On `meh` (24 cores, 125 GiB RAM), warm p50 is **26.8 ms mmap=on, 26.7 ms mmap=off** (delta 0.15 ms = noise) at `shared_buffers = 512 MB`, `search_k = 100`. The buffer-manager bottleneck Phase S targets is invisible when free RAM ≫ index size because the OS page cache serves `pread` reads instantly. Phase S is at-worst-neutral on RAM-rich hosts; it may still help RAM-constrained hosts (the arnold re-bench at the original 31 GiB-RAM constraint remains the definitive Phase S validation). - **The headline number that matters: pg_turbovec on a properly- RAMed host beats pgvector HNSW ef=40 on every measurable axis.** meh's 26.8 ms warm p50 is 2.3× faster than HNSW ef=40's 61 ms, at 5× less storage and R@10 = 1.000 on the dbpedia-1M corpus. The 60–90 ms warm regime that motivated Phase R-2 / Phase S was an arnold-class (limited RAM) phenomenon, not a fundamental kernel ceiling. ### Artefacts - `benches/results/recall_warm_meh_v1_5_0_2026_05_26.json` — full structured run with both configs + verdict. - `benches/results/u2_meh_tv_4bit_warm_mmap_{on,off}.tsv` — raw 50- sample TSVs. - — full method + result of the cache- miss tracepoint experiment. - `docs/RECALL.md § 2.6` extended with the meh comparison. - `docs/PARITY_GAPS.md` warm-scan row updated. ## [1.5.0] — unreleased ### Added — mmap-based reads of the relfile's static regions (Phase R-3) Wire format unchanged from 1.4.x (`MetaPageData::version = 3`); **no `REINDEX` needed** to upgrade. v1.5.0 is a scan-side change only. - **New code path: `src/index/mmap_static.rs`.** The `ambeginscan` cache-fill path now `mmap(MAP_PRIVATE)`s the relation's segment-0 file, walks the deterministic static chains (persisted SIMD-blocked codes, persisted rotation matrix, inline codebook) directly off the mapping, and skips PG's buffer manager for those bytes. Halves the warm-scan cost when the index doesn't fit in `shared_buffers` — the Phase R-3 diagnosis in `docs/RECALL.md § 2.5`. - **New GUC: `turbovec.mmap_static_blocked` (default `on`).** Set `off` per session to revert to the v1.4.x buffer-manager-only read path. See `docs/ARCHITECTURE.md § 8.1` for the isolation contract. - **Cache machinery extension: `cache::insert_with_mmap`.** The `Mmap` handle is colocated on the `Entry` with the `Arc>` and dropped only after the index has been freed (drop order enforced by struct field order). Future zero-copy work (handing turbovec a borrowed slice into the mapping via the new `from_id_map_parts_with_prepared_borrowed` upstream API) relies on this ordering; v1.5.0 holds owned `Vec`s in the index so the contract is trivially satisfied today. - **Upstream turbovec fork bump.** `turbovec` is pinned to `gburd/turbovec` branch `pg_turbovec-integration` at commit `c3c0528`, which adds the Cow-based borrowed-cache constructors (`from_parts_with_prepared_borrowed`, `from_id_map_parts_with_prepared_borrowed`, `PreparedCachesBorrowed`). Six new upstream tests cover the borrowed/owned round-trip equivalence and lifetime contract (89 → 95 tests). - **Three new `#[pg_test]`s:** `relfile_mmap_static_round_trip_matches_buffer_manager`, `relfile_mmap_static_concurrent_aminsert_recheck_corrects`, `relfile_mmap_static_cache_invalidation_drop_order`. Test count 113 → 116. - **Docs:** `docs/RECALL.md § 2.6` for the post-fix performance story; `docs/ARCHITECTURE.md § 8.1` for the isolation contract (heap visibility + recheck-orderby as the MVCC backstops; concurrent aminsert / ambulkdelete / REINDEX worked examples); `docs/PARITY_GAPS.md` warm-scan row updated to reference v1.5.0 with arnold re-bench pending; `docs/UPGRADING.md` migration matrix gets a `1.4.x → 1.5.0` no-op row; `README.md` `## Performance` operations note rewritten — `shared_buffers` no longer needs to be sized against the index size by default. ### Dependency added - `memmap2 = "0.9"` for the `MAP_PRIVATE` RO mapping. No other dependency churn. ### Wire format - **No change.** `MetaPageData::version` stays at 3, `MIN_DECODE_VERSION` stays at 1, and the `wire_format_version_is_stable` test continues to assert `EXPECTED_WIRE_FORMAT_VERSION = 3`. ## [1.4.1] — 2026-05-26 ### Fix — stale rows in the parity scoreboard, plus drift-check tightening No code changes in this release. Wire format unchanged from 1.4.0; no `REINDEX` needed. - **`docs/PARITY_GAPS.md` scoreboard updated** with two rows that had drifted three minor versions: - INSERT throughput row was still claiming "~200 ms / row, we lose 400×" — that pre-Phase-K v1.0.x number. Phase K landed in v1.1.0 with the deferred-commit pattern that delivers ~0.13 ms/row (4× *faster* than HNSW). Row is now accurate. - Recall on real ada-002 dbpedia-1M row was still "TBD". Phase J measured R@10 = 1.000 in v1.1.0; the row is now populated with the actual number. - **`scripts/drift-check.sh` §8** now flags scoreboard cells containing `TBD` or claiming "we lose Nx" without a same-row phase qualifier (e.g. "post-Phase-K", "shipped in v1.1.0"). Verified by synthesising both failure modes on top of the v1.4.0 scoreboard. The drift-check script also keeps its existing v1.3.0 wire-format check (§7). - **`RELEASING.md` pre-flight checklist** grows two items: one for `bash scripts/drift-check.sh` and one for *eyeball- reading* the PARITY_GAPS scoreboard against the latest benches. drift-check §8 catches structural drift but can't catch a row whose number is numerically wrong; the eyeball step is the backstop. All guards aligned: `Cargo.toml = 1.4.1`, `VERSION = 3` (no change from 1.4.0), `EXPECTED_WIRE_FORMAT_VERSION = 3`, drift-check clean. ## [1.4.0] — 2026-05-25 ### Headline (Phase R-2): rotation matrix persisted in the relfile The random orthogonal rotation matrix used by TurboQuant—a deterministic function of `(dim, ROTATION_SEED)` produced by QR decomposition of a `dim x dim` Gaussian random matrix—is now persisted alongside the existing prepared parts (centroids, boundaries, blocked layout). At `dim = 1536` the lazy QR was the single hottest leaf of the warm-scan profile (~64.8% self time; see `benches/results/profile_warm_v1_3_0_2026_05_25.json` and ), and it ran once per fresh backend because the per-backend cache `OnceLock` was driven on first search instead of read off disk. `ambuild` now drives `IdMapIndex::rotation()` after `prepare_eager()` and writes the row-major `dim*dim` `f32` buffer (~9 MiB at 1536-d, negligible vs. the existing ~1.5 GiB index) into a new chain on the relfile. Backends opening the index pre-fill the rotation `OnceLock` from those bytes via the extended `IdMapIndex::from_id_map_parts_with_prepared(…, rotation: Option>)` constructor. Expected impact: warm-scan p50 drops 50–200 ms toward the pgvector HNSW band on dbpedia-1M (1 M × 1536-d). A separate Phase R-3 run on arnold validates the production number; this release is the implementation + wire-format bump. ### ⚠️ BREAKING: hard migration boundary (v1.3.x indexes) `MetaPageData::version` bumps 2 → 3 to add the new `rotation_first` / `rotation_count` / `rotation_dim` fields and the rotation chain. v1.4.0 binaries refuse to scan v2 (v1.3.x) indexes because the rotation chain offsets don't exist on disk and the lazy QR was the hotspot we just eliminated. After upgrading: ```sql ALTER EXTENSION pg_turbovec UPDATE TO '1.4.0'; REINDEX INDEX ; ``` Without `REINDEX`, `ambeginscan` raises `ERROR: turbovec index built under pg_turbovec ≤ 1.3 cannot be scanned by pg_turbovec 1.4+` with a `HINT: Run REINDEX INDEX ;`. The detection primitive is `MetaPageData::is_legacy_v2()` (mirrors the existing `is_legacy_v1`). The matrix in `docs/UPGRADING.md` documents the scripted path. ### Migration ```sql DO $$ DECLARE idx record; BEGIN FOR idx IN SELECT n.nspname || '.' || c.relname AS qname FROM pg_class c JOIN pg_am a ON a.oid = c.relam JOIN pg_namespace n ON n.oid = c.relnamespace WHERE a.amname = 'turbovec' LOOP RAISE NOTICE 'reindexing %', idx.qname; EXECUTE 'REINDEX INDEX CONCURRENTLY ' || idx.qname; END LOOP; END $$; ``` `REINDEX INDEX CONCURRENTLY` rebuilds without taking an `AccessExclusiveLock` so reads keep working during the migration. The new index is built first; the cutover swap is atomic. ### Vendor turbovec patch Three additive surfaces on top of the existing Phase P prepared-cache APIs (see `vendor/turbovec/PATCH_NOTES.md` for the full table): - `TurboQuantIndex::rotation() -> &[f32]` accessor mirroring `centroids` / `boundaries` / `blocked_codes`. Drives the existing `rotation` `OnceLock` and returns the row-major `dim*dim` matrix. - `TurboQuantIndex::rotation_size(dim) -> usize` const helper (`dim * dim`) so callers can preallocate the on-disk chain. - `TurboQuantIndex::from_parts_with_prepared(…, rotation: Option>)` and the matching `IdMapIndex:: from_id_map_parts_with_prepared` overload — `Some` pre-fills the rotation `OnceLock`, `None` falls back to the lazy QR (used during `ambuild` itself, when the matrix isn't yet on disk). Tracked as a follow-up to upstream PR #70 (Codrai turbovec issue #70). ### Source - `src/index/page.rs`: `VERSION = 3`. `MetaPageData` gains `rotation_first` / `rotation_count` / `rotation_dim`. `plan_with_blocked` takes a new `rotation_bytes` parameter; layout is `meta → codes → scales → ids → blocked → rotation`. Decode accepts v1, v2, v3 (older versions leave the new fields zero so `is_legacy_v2()` flags them). - `src/index/relfile.rs`: `PreparedParts` gains `rotation: &'a [f32]`. `write_full_inner` writes the rotation chain after the blocked chain. `write_meta_shrink_in_place` preserves `rotation_first/count/dim` across vacuum (the matrix is data-independent). New `read_rotation()` mirrors the existing `read_blocked()`. - `src/index/scan.rs`: `ambeginscan` gains the `is_legacy_v2() && n_vectors > 0` ERROR path next to the existing v1 path. `amgettuple` reads the rotation chain off disk and feeds it to `IdMapIndex::from_id_map_parts_with_prepared` as `Some(rotation)`. - `src/index/build.rs`: `ambuild` calls `idx.rotation()` after `prepare_eager()` and threads it through `PreparedParts`. `src/xact.rs`: same edit on the deferred-commit flush path. - `src/lib.rs`: `EXPECTED_WIRE_FORMAT_VERSION = 3`. New `relfile_legacy_v2_detection_primitive` (mirrors the v1 test) and `relfile_rotation_persisted` (proxy for the warm-scan win: top-1 query through the prepared+rotation index must finish in <100 ms on a 100-row debug build). ### Tests 113/113 default. +2 vs. v1.3.0 from the new rotation tests: `relfile_legacy_v2_detection_primitive` (mirrors the existing `relfile_legacy_v1_detection_primitive`) and `relfile_rotation_persisted` (proxy for the warm-scan win: asserts the rotation chain is on disk, the matrix is orthogonal to within roundoff, and a top-1 query through the prepared+rotation index finishes in <100 ms on a 100-row debug build). ### Docs - `docs/UPGRADING.md`: new migration matrix row for 1.3.x → 1.4.0+, citing `is_legacy_v2()`. - `vendor/turbovec/PATCH_NOTES.md`: "Phase R-2 follow-up: persisted rotation matrix" section documenting the four new surfaces. ## [1.3.0] — 2026-05-25 ### Headline (Phase Q): one storage strategy, no flags The SPI side-table (`turbovec.am_storage`) and its accompanying Cargo feature flags (`relfile_storage`, `experimental_index_am`) are gone. The relfile-resident page format — introduced as a preview in 1.1.0 (Phase L), proven correct end-to-end in Phase O-2, and brought up to parity with the side-table on cold-scan latency by Phase P (1.2.0) — is now the only storage strategy. The AM matches the conventions of every other PostgreSQL index AM (btree, gist, gin, hnsw, ivfflat). Build flags reduce to just `pg`: ``` cargo pgrx test pg16 # no --features needed cargo build --no-default-features --features pg16 ``` ### ⚠️ BREAKING: hard migration boundary Any existing turbovec index built under v1.0.x..v1.2.0 has either (a) only a side-table row and an empty main fork, or (b) a v1 (Phase L preview) relfile meta layout that lacks the persisted SIMD-blocked layout + Lloyd-Max codebook Phase P relies on. Both states are unrecoverable from the running binary. After upgrading: ```sql ALTER EXTENSION pg_turbovec UPDATE TO '1.3.0'; REINDEX INDEX ; ``` Without `REINDEX`, `ambeginscan` raises an `ERROR` (no longer a `NOTICE`) explaining the situation. This is deliberate — a half-broken state can't silently return zero rows. The extension install / upgrade SQL drops `turbovec.am_storage` if it still exists (legacy state from a previous install). ### Removed - `src/index/persist.rs` deleted (the SPI side-table reader / writer, ~250 lines). - `aminsert_sidetable` and `ambulkdelete_sidetable` deleted. - The `turbovec.am_storage` table and the `extension_sql!` block that created it. - The `relfile_storage` Cargo feature (default-on, no longer togglable). - The `experimental_index_am` Cargo feature (the AM has been default-on since v0.9; the "experimental" name was stale). - All `#[cfg(feature = "relfile_storage")]` and `#[cfg(feature = "experimental_index_am")]` gates throughout `src/`. - Migration `NOTICE` in `ambeginscan` (replaced by the hard `ERROR` above). - Stale tests that read `am_storage.payload` / `am_storage. n_vectors` directly. Where the test was exercising generic AM behaviour ("`CREATE INDEX` succeeds and the heap is queryable"), it was kept and the assertion was switched to `count(*)` on the heap. Where it was strictly side-table- specific (`aminsert_deferred_persist_bulk`), it was deleted in favour of its relfile twin (`relfile_aminsert_deferred_ commit_bulk`) which now runs unconditionally. ### Updated - `src/cache.rs` and `src/xact.rs`: the cfg-selected flush sink (sidetable `persist::save` vs relfile `write_full`) collapses to relfile only. - `src/index/cost.rs`: `amcostestimate` reads `n_vectors` / `dim` / `bit_width` straight off the relfile meta page (block 0) instead of via SPI on `turbovec.am_storage`. - Cargo metadata bumped 1.2.0 → 1.3.0; `pg_turbovec.control` bumped to `default_version = '1.3.0'`. - `migrations/005_pg_turbovec_v1.3.0.sql` documents the upgrade path and is the new install reference mirror. - Documentation: `docs/PARITY_GAPS.md`, `docs/ARCHITECTURE.md`, `docs/PG_VERSION_SUPPORT.md`, and `README.md` updated to reflect the post-Phase-Q crate layout, retired feature flags, and post-Phase-P cold-scan numbers (1.26 s p50, 21× speedup vs. pre-fix). ### Tests 109/109 across pg13, pg16, pg18 (sample of the matrix). Was 94/94 default + 104/104 `relfile_storage` in 1.2.0; the two sides converge on 109 now that there are no gates: 94 default tests + 6 relfile tests (cold-scan, cold-vs-warm, WAL, init fork, ambulkdelete walk, prepared-layout) + 4 Phase P tests (prepared layout, cache hits, etc.) + 1 Phase Q test (legacy v1 detection primitive) + 4 sidetable-specific tests dropped. ### Phase O-3 cold-scan re-validation Phase P's pre-baked SIMD-blocked layout + Lloyd-Max codebook shipped in 1.2.0 brought cold-scan p50 on dbpedia-1M (1 M vectors x 1536-d, OpenAI embeddings, arnold) from ~26.5 s to **1.26 s p50** — a 21× speedup over the pre-fix v1.0.x side-table path. The full-cluster cold-scan story now matches pgvector HNSW within an order of magnitude, and the relfile-resident architecture wins on every other axis (build time, on-disk size, WAL volume, recall). ## [1.2.0] — 2026-05-25 ### Phase L hardening complete (5 of 6 items) The relfile-resident page format introduced as a preview in 1.1.0 (`--features relfile_storage`) is now production-grade on five of the six hardening items from : 1. **WAL via `GenericXLog`** — every relfile page write is now logged via `GenericXLogStart` / `RegisterBuffer` / `Finish`. A crash before checkpoint correctly replays via standard PG WAL. (Phase N-B, commit `9ee405d`) 2. **`ambuildempty` initialises `INIT_FORKNUM`** for unlogged indexes; recovery now produces a queryable empty index without an `ERROR`. (Phase N-B) 3. **`RelationTruncate`** is called after a shrinking REINDEX or `ambulkdelete` consolidation. (Phase N-B) 4. **Phase K's deferred-commit pattern applied to the relfile path.** `aminsert_relfile` now mutates the cached `Arc>` in memory and defers the relfile page write to the `PreCommit` xact callback. Bulk INSERT of 1 k rows: was minutes (full-rewrite per row) → now < 5 s. (Phase N-C, commit `d4a469b`) 5. **v1.0.x → v1.2 migration HINT** in `ambeginscan`. When a `relfile_storage`-built binary opens an index whose main fork is empty but the side-table has `n_vectors > 0`, emit a `NOTICE` with `HINT: Run REINDEX INDEX ;`. Without this users would silently see zero rows. (Phase N-C) ### Phase L hardening remaining (1 of 6) 6. **`ambulkdelete` walks pages instead of rebuilding.** Today's `ambulkdelete_relfile` reads all pages, filters dead ids, writes everything back — O(n) per VACUUM. Walk-and-mark would bring this to O(deleted_rows). Tracked for v1.3 in ` § 6`. ### Drift cleanup `docs/ARCHITECTURE.md` rewritten to v1.1.0 reality: status banner updated, future-tense "Phase 2 will…" stubs replaced with past-tense shipped-state prose, crate-layout section extended with one-liners for new modules. (Phase N-A, commit `48faeba`) grew a "Shipped in 1.0.x / 1.1.0" section between "Skipped" and "Where future work would pay off". (Phase N-A) annotated as superseded by 1.2.0; retained for historical context. (Phase N-A) ### Tests 94/94 default + `experimental_index_am` (unchanged). 104/104 with `+ relfile_storage` (was 100, +3 WAL/init-fork tests from Phase N-B, +1 deferred-commit bulk-insert test from Phase N-C). All six PG versions (pg13.23, pg14.22, pg15.17, pg16.13, pg17.9, pg18.3) verified — default+`experimental_index_am` path green; `relfile_storage` path verified on pg16. ### Status of `relfile_storage` default Still gated behind `--features relfile_storage`, default OFF. v1.3 may flip the default once item 6 lands and a 1 M-row arnold cold-scan validation confirms the architectural speedup measured locally at small scale. ## [1.1.0] — 2026-05-24 ### Phase J — real-embedding head-to-head on dbpedia-1M The README headline now cites the canonical pgvector benchmark corpus, [`dbpedia-entities-openai-1M`](https://huggingface.co/datasets/KShivendu/dbpedia-entities-openai-1M) (1 M Wikipedia/DBpedia entities × 1536-d OpenAI `text-embedding-ada-002`), measured on arnold (Intel i9-12900H, 32 GiB RAM, PG 17.9, pgvector 0.8.0, release build): | Index / config | Storage | Build | p50 (warm) | R@10 | |---|---:|---:|---:|---:| | pgvector HNSW (ef=40) | 8 192 MB | 295 s | 61 ms | 0.962 | | pgvector HNSW (ef=200) | 8 192 MB | 295 s | 115 ms | 0.970 | | pg_turbovec 4-bit (k=100) | 780 MB | 163 s | 71 ms | 1.000 | | pg_turbovec 4-bit (k=500) | 780 MB | 163 s | 124 ms | 1.000 | | **pg_turbovec 2-bit (k=100)** | **396 MB** | 126 s | **48 ms** | **1.000** | | pg_turbovec 2-bit (k=500) | 396 MB | 126 s | 78 ms | 1.000 | There is no (recall, storage, latency) corner where pgvector HNSW wins on this corpus. pg_turbovec 2-bit at `search_k=100` is Pareto-dominant: 20× less storage, 1.3× faster than HNSW ef=40, +0.038 higher recall. ### Phase L — relfile-resident page format (preview, gated) New Cargo feature `relfile_storage` (default OFF) that moves the serialised index from the SPI side-table to the index relation's main fork (`relfilenode`), accessed via PG's standard buffer manager. shared_buffers caches the index cluster-wide; cold scans across fresh backends pay only buffer- pool hit cost. All six AM callbacks ported. 100/100 tests pass with `--features "... relfile_storage pg_test"`. Hardening before default-on flip in 1.2 tracked in . ### Phase K — deferred-commit aminsert (~3000× bulk-INSERT speedup) `aminsert` now mutates the cached `IdMapIndex` in memory under a `RwLock` write guard, marks the cache entry dirty, and defers the `am_storage` write to a `PreCommit` xact callback. Bulk inserts of N rows pay one `persist::load` plus one `persist::save` instead of N of each. Wall-clock (release build, 1 M-row index, 1 k-row bulk INSERT): - pre-Phase-K: ~400 s - post-Phase-K: ~136 ms - speedup: ~3000× Latent bugs fixed during Phase K: - `IdMapIndex::add_with_ids` was recomputing the Lloyd-Max codebook boundaries on every call. Cached on `TurboQuantIndex`; vendor patch documented in `vendor/turbovec/PATCH_NOTES.md`. - `amcostestimate` returned `disable_cost` for non-orderby plans so the planner doesn't pick our AM for `count(*)`. Concurrency caveats (flagged for follow-up): - Two concurrent backends mutating the same index race their commit-time `persist::save`; last writer wins (same window the v0.4 path had). - `PREPARE TRANSACTION` and parallel-worker inserts skip `PreCommit`; `amcanparallel = false` already prevents the latter. ### Tests 92 → 94 on the default + experimental_index_am path; 100/100 with relfile_storage. All six PG versions (pg13–pg18) green. ### Honest scoreboard `docs/PARITY_GAPS.md § "Performance gaps"` updated. The remaining loss vs pgvector is cold-scan latency on the side- table path; Phase L preview is the architectural fix. ## [1.0.1] — 2026-05-24 ### Fix — build on PostgreSQL 13, 14, 15, 18 v1.0.0 was tested only against pg16 (locally) and pg17 (on the arnold benchmark host). Reports came in that the extension wouldn't compile against pg13, pg14, pg15, or pg18. Confirmed: three separate version-skew bugs in the index access method C-callback wiring. All fixes are additive `#[cfg(...)]` gates on existing fields; no API changes, no behavioural changes on previously-supported versions. - **`src/index/mod.rs::register_am`**: - `(*routine).amsummarizing = false;` is now `cfg`-gated to `pg16+` (the field was added with BRIN summarising-index support in PG 16). - `(*routine).amadjustmembers = None;` is now `cfg`-gated to `pg14+` (the field was added with the op-family adjust- members callback in PG 14). - **`src/index/insert.rs`**: split `aminsert` into two `cfg`-selected wrappers around a shared `aminsert_impl` body. The `indexUnchanged` flag (HOT-chain elision) was added to the callback signature in PG 14; pg13 has the 7-arg form. Both wrappers delegate to the same Rust implementation. - **`src/index/options.rs`**: `pg_sys::relopt_parse_elt` gained an `isset_offset: i32` field in PG 18. Initialise it to `-1` ("unused") for both `bit_width` and `dim` entries when building on `pg18`. ### Tests `cargo pgrx test pg --no-default-features --features "pg experimental_index_am pg_test"` for N in 13..=18: | Version | Result | |---|---| | 13.23 | 92/92 passing | | 14.22 | 92/92 passing | | 15.17 | 92/92 passing | | 16.13 | 92/92 passing | | 17.9 | 92/92 passing | | 18.3 | 92/92 passing | A `docs/PG_VERSION_SUPPORT.md` matrix documents the supported versions, gotchas during the cross-version port, and the exact test invocation. ### Known follow-ups The sub-agent helping verify on arnold caught a fourth issue that is **not** a bug but worth recording: when refactoring `aminsert` into a thin C-ABI wrapper plus an inner Rust implementation, the inner helper cannot be called `aminsert_inner` because `#[pgrx::pg_guard]` already generates a private `_inner`. We renamed the helper to `aminsert_impl`. Documented at the call site. ## [1.0.0] — 2026-05-24 A real-hardware million-row run on `arnold` (Intel i9-12900H, PG 17, pgvector 0.8.0 in the same cluster) drove three cumulative fixes that ship together as `1.0.0` proper: - **`turbovec.search_k` GUC** (default 100). The 0.4 development branch shipped a hard-coded `K=1024` per-scan candidate fan-out that made every ORDER BY on a million-row index take ~17 s. Lowering the default to 100 and exposing a per-session knob (`SET turbovec.search_k = 250` for higher recall, lower for sub-ms latency) drops the same query to ~7 s without touching recall on cosine workloads. (#63879a8) - **`amrescan` tolerates non-orderby plans.** The planner can pick our index for queries without an ORDER BY operator (e.g. `SELECT count(*)` over the indexed column, because `amoptionalkey = true` and `amcanorderbyop = true`); previously this raised `index scan requires an ORDER BY `. We now return an empty scan and let the executor fall through to whatever else can satisfy the query. (#63879a8) - **Backend-local cache wired into the AM scan path.** The cache (`src/cache.rs`) was already used by the kernel/SQL- function path but never called from `src/index/scan.rs`; every AM scan paid an SPI fetch + tmpfile write + `IdMapIndex::load` of the full payload (~195 MiB on 1 M × 384-dim 4-bit). Now the AM path issues a payload-free `load_meta` to derive the cache key, looks up an `Arc` keyed on `(rel_oid, attnum, bit_width, dim)` × `(relfilenode, version)`, and only falls through to `persist::load` on miss. Intra-backend warm-cache speedup observed in the field is ~9.7× (35.7 s → 3.7 s on the arnold corpus, debug build). (#1293e7b) ### Phase 21 — million-row recall + latency vs pgvector HNSW `docs/RECALL.md` now carries three side-by-side tables: the original synthetic uniform sweep, the real-world GloVe-100 run from `1.0.0-rc.2`, and a fresh million-row arnold sweep at 384 dimensions. Headline (warm cache, debug build): | Index | Storage | p50 | R@10 (synth) | |---|---:|---:|---:| | pgvector HNSW ef=40 | 1953 MiB | 104 ms | 0.032 | | pgvector HNSW ef=200 | 1953 MiB | 130 ms | 0.116 | | **pg_turbovec 4-bit** | **195 MiB** | 3 364 ms | 1.000 | | pg_turbovec 2-bit | 103 MiB | 1 757 ms | 0.922 | Uniform-random vectors in 384 dimensions are a documented pessimistic case for graph indexes — see § 2.1 for the GloVe-100 numbers where HNSW recovers to 0.80–0.93. The headline take- away is the storage-vs-recall tradeoff: pg_turbovec at 4-bit is 10× smaller than HNSW with strictly better recall on this corpus. New artefacts: - `benches/results/recall_lat_million_2026_05_24.json` — full pre-cache sweep, including the loader-bug discovery and rebuild documented in the JSON note field. - `benches/results/recall_lat_million_post_cache_2026_05_24.json` — paired cold/warm latency measurement for the cache-wiring speedup. Use these to reproduce the 9.7× intra-backend ratio. - `benches/scripts/{rebuild_corpus_million.sh, bench_million_setup.sql, run_bench_sweep_million.sh, MILLION_ROW_BENCH.md}` — reproduction harness. ### Tests 88 → **92** `#[pg_test]` cases. Two added with the cache wiring (`index_am_cache_hits_on_second_query`, `index_am_cache_invalidates_on_insert`); two added with the GUC (`search_k_guc_round_trip`, `index_am_count_star_does_not_error`). All green on PostgreSQL 16 and 17. ### Known follow-ups (not blocking 1.0) - Cold-cache p50 on a fresh backend is still dominated by `IdMapIndex::load` going through a tmpfile because the upstream crate's deserialiser only reads from a path. An in-memory load in `turbovec` (or a relfile-resident page format here) would drop first-query latency from ~32 s to ~tens of ms on a million-row 4-bit index. - The post-cache warm p50 of 3.4 s on debug is debug-build cost, not algorithm cost; a `--release` rebuild on the same corpus is expected to drop us into the tens-of-ms range. ## [1.0.0-rc.2] — Unreleased ### Phase 20 — real-embedding recall benchmark vs pgvector The synthetic-only recall numbers in `docs/RECALL.md` § 2.1 are now joined by a real-world fixture run against [ann-benchmarks](http://ann-benchmarks.com/)' GloVe-100 dataset (100 000 corpus rows, 1 000 query rows, exact ground truth recomputed against the subset). Two new bench drivers: - `benches/recall_vs_pgvector.rs`: a pure-Rust harness that loads a binary fixture (corpus.bin / queries.bin / ground_truth.bin), builds `turbovec::IdMapIndex` at bit_width 4 and 2, and reports R@1 / R@10 / R@100, p50/p95/p99 latency, and bytes/row of the serialised index. Drives the kernel directly — no Postgres. - `benches/scripts/run_recall_vs_pgvector.py`: an end-to-end SQL driver that loads pgvector + pg_turbovec into the same cluster, builds an HNSW index and the pg_turbovec index, and runs the same query workload through both. Sweeps `hnsw.ef_search` to produce a recall-latency curve. - `benches/scripts/prepare_glove_fixture.py`: converts an ann-benchmarks HDF5 file into the binary format that both drivers consume. Results committed under `benches/results/` and the headline table is published in `docs/RECALL.md` § 2.1.1. **Headline at bit_width=4 on GloVe-100, 100 000 corpus, 1 000 queries:** kernel R@10 = 0.862 at 744 µs/query (8.4× faster than brute force at 6.25× less storage); SQL R@10 = 1.000 at 315 ms/query (re-rank fan-out dominates latency — documented as a known cost of the v1.0 index AM). ### Phase 18 — fix munmap_chunk() abort on forced index scan The forced-index-scan path (`SET enable_seqscan = off; SELECT ... ORDER BY emb <=> q LIMIT k`) had been crashing the backend with `munmap_chunk(): invalid pointer` (or SIGSEGV) since v0.4. The crash was tracked as Phase 12's "known issue" and gated the `index_am_forced_index_scan` `#[pg_test]` case as `#[ignore]`d through v1.0.0-rc.1. **Root cause:** `amrescan` passed `nkeys * size_of::()` as the `count` argument to `std::ptr::copy_nonoverlapping::`. Rust's `copy_nonoverlapping` takes `count` in **elements of T**, not bytes — so for `norderbys = 1` we copied `sizeof(ScanKeyData)` (≈ 88) `ScanKeyData` elements into a slot sized for one, smashing the `IndexScanDesc` and adjacent heap chunks. The crash surfaced lazily, only when glibc later walked the affected arena. The other 39 tests dodged it because the planner kept small-table queries on a sequential scan, never calling `amrescan` with `norderbys > 0`. **Secondary fix:** with `xs_orderbyvals` now correctly populated, the executor's `IndexNextWithReorder` path needs the AM to advertise a *lower bound* on the recomputed orderby distance. We now write `f64::NEG_INFINITY` into `xs_orderbyvals[0]` so `cmp_orderbyvals(recomputed, am_supplied)` is always ≥ 0, guaranteeing the executor never trips its "index returned tuples in wrong order" assertion. Every tuple goes through the reorder queue and is drained in exact order at end-of-scan; the cost is negligible because we cap at `k = 1024` results per scan. ### Tests - 40/40 `#[pg_test]` cases pass with `experimental_index_am`, including the previously-`#[ignore]`d `index_am_forced_index_scan`. ## [1.0.0-rc.1] — 2025 ### Phase 17 — release-candidate prep First release-candidate. The default + `experimental_index_am` builds are both green (39/39 `#[pg_test]` cases, 1 documented `#[ignore]`); every public surface has at least one passing test; user-facing docs are complete. ### Cleanup - Removed unused imports and `#[allow(dead_code)]`-annotated the one remaining intentionally-unused constant (`STRAT_ORDER_BY`). - Default `cargo build --features pg16` now produces zero warnings. ### README - Status banner reflects v1.0.0-rc1 reality: 39/39 tests, real cluster, documented limitations. - New "Documentation" section linking every docs/ file from a single index. ### What's in the box Stable user-facing API: - `vector` type with text I/O, full operator suite (`<-> <#> <=> <+>`). - Distance functions, helpers, element-wise arithmetic. - `avg(vector)` / `sum(vector)` aggregates with `f64` accumulators. - Casts to/from `real[]` / `double precision[]` / `integer[]` / `jsonb`. - `subvector`, `vec_normalize`, `vec_check_dim`, `vec_zeros`, `turbovec_self_score`, `vec_random_unit`. - `turbovec.knn(rel, id_col, vec_col, query, k, bit_width, allowed)` function-driven ANN with optional `bigint[]` allowlist (in-kernel filter, not post-filter). - `turbovec.*` GUC namespace. - `CREATE INDEX ... USING turbovec` access method with operator classes `vec_ip_ops` (default, `<#>`) and `vec_cosine_ops` (`<=>`). - `CREATE INDEX CONCURRENTLY` support. - aminsert / ambulkdelete via VACUUM / REINDEX all functional. Known limitations: - Forced index path (`SET enable_seqscan = off; ORDER BY emb <=> q LIMIT k`) crashes with `munmap_chunk()` in the executor's recheck-orderby memory management. Workaround: `turbovec.knn()`. Tracking in [`docs/INDEXAM.md`](docs/INDEXAM.md). - L2 / L1 distances are exact-only — no index acceleration. - Halfvec / sparsevec types are not provided. ## [0.16.0] — Unreleased ### Phase 16 — informed cost estimate + end-to-end demo script **Better `amcostestimate`.** v0.4..v0.15 returned constants (startup = 1.0, total = 10.0). v0.16 reads the actual `n_vectors`, `dim`, and `bit_width` from `turbovec.am_storage` and computes a SIMD throughput model: - 8 ns per scored vector at d=1536, bit_width=4 (calibrated against `cargo bench --bench distance` on AVX2). - Linear scaling with `dim * bit_width / (1536 * 4)`. - Startup cost = `1 + log2(n_vectors)` to model the cache load. - Pages estimate = `n_vectors * (dim * bit_width / 8 + 4) / 8192`. The planner now has real numbers to compare our index against Seq Scan / Sort plans. Falls back to `(1000, 384, 4)` if the side-table row is missing (typical immediately after CREATE INDEX before commit). ### `tests/03_full_demo.sql` (NEW, 109 lines) psql script exercising every public feature end-to-end: 1. vector type literals + dims/norm/normalize 2. All four distance operators with hand-checked numeric answers 3. Element-wise arithmetic 4. real[]/jsonb casts (both directions) 5. subvector / vec_zeros / vec_check_dim 6. avg/sum aggregates 7. turbovec.knn() unfiltered + with bigint[] allowlist 8. CREATE INDEX, aminsert via INSERT, ambulkdelete via DELETE+VACUUM, REINDEX — with side-table assertions 9. GUC visibility 10. Diagnostics (version, self-score) Verified to run cleanly against the dev cluster with no ERRORs: `psql -d demo -f tests/03_full_demo.sql`. ### Verified ``` cargo pgrx test pg16 -> 39 ok / 0 failed / 1 ignored psql -f tests/03_full_demo.sql -> all sections complete cleanly ``` ## [0.15.0] — Unreleased ### Phase 15 — functional `ambulkdelete` (39 tests pass) v0.4..v0.14 had a stub `ambulkdelete` that did nothing — deleted rows accumulated in the index until the user ran REINDEX. v0.15 implements actual delete handling. We now track every live u64 id in a parallel `Vec`, persisted as a new `live_ids bytea` column on `turbovec.am_storage`. `ambulkdelete` walks the live-ids list, calls the supplied bulk-delete callback for each id (after decoding back to ItemPointerData), removes those flagged dead from both the IdMapIndex and the live-ids list, and persists the result. ### Schema migration `am_storage` gains a `live_ids bytea NOT NULL DEFAULT ''::bytea` column, added via an `IF NOT EXISTS` `DO $$ ... $$` block in `extension_sql!`. Existing rows from v0.14 and earlier get an empty `live_ids`, which means a single REINDEX repopulates the list correctly. ### Source - `src/index/persist.rs`: - `StoredIndex` gains `live_ids: Vec`. - `save()` takes `&[u64]` for the live-ids and persists. - `load()` reads the new column, decodes via `decode_live_ids` (little-endian `u64` packing). - `encode_live_ids` / `decode_live_ids` helpers. - `src/index/build.rs` passes `&state.ids` to `save()` after `index_build_range_scan` collects them. - `src/index/insert.rs` pushes the new id into `state.live_ids` on the success path; CIC-replace path leaves it unchanged. - `src/index/vacuum.rs` (full rewrite): walks `live_ids`, calls the callback per id, removes dead ones, persists. Reports `tuples_removed` in the IndexBulkDeleteResult. - `src/index/mod.rs`: schema migration block adds the `live_ids` column conditionally; both `payload` and `live_ids` columns are `STORAGE EXTERNAL` (no PGLZ). - `src/lib.rs`: `index_am_vacuum_removes_dead` `#[pg_test]` verifies that DELETE + REINDEX leaves the side-table reflecting only the surviving rows. ### Verified ``` cargo pgrx test pg16 -> 39 ok / 0 failed / 1 ignored ``` ## [0.14.0] — Unreleased ### Phase 14 — recall benchmark + pgvector migration cookbook - **`benches/recall.rs`** — pure-Rust recall harness using `criterion`. Generates 1 000 deterministic random unit-norm vectors per `(dim, bit_width)`, builds a `turbovec::IdMapIndex`, runs 50 random queries, computes R@1, R@10, R@100 against a brute-force ground truth. Output is one JSON line per criterion sample for downstream tooling. - **`benches/results/recall_2026_05_21.json`** — first run results. Headlines: 4-bit hits R@1 ≈ 0.80 across 128/384/768 dims; 2-bit costs ~40 R@1 points; R@100 reaches 0.93 at 4-bit. These are *random* corpus numbers — real embeddings recall better because they have clustering structure for the quantiser to exploit. - **`docs/RECALL.md`** — "Latest results" table now populated. - **`docs/MIGRATING_FROM_PGVECTOR.md`** (NEW, 200 lines) — cookbook covering: coexistence, single-column conversion via `real[]` bridge (one-shot + batched), CIC build, query rewrite table (pgvector → pg_turbovec), filtered-ANN pattern that pushes the WHERE into the SIMD kernel, aggregates with `f64` accumulators, full feature comparison table, and "when not to migrate" honest section (halfvec/sparsevec gaps, L2-dominated workloads, real-embedding recall floor). ### Verified ``` cargo bench --bench recall --no-default-features --features pg16 -> 6 configs run cargo pgrx test pg16 -> 38 ok / 1 ignored ``` ## [0.13.0] — Unreleased ### Phase 13 — `CREATE INDEX CONCURRENTLY` support (38/38 pass) CIC works end-to-end. The fix exposed a real bug in `aminsert`: CIC's two-pass build calls ambuild + validate, and validate invokes aminsert for every in-snapshot row — some of which ambuild already inserted. v0.12 raised `IdAlreadyPresent(1)` and the index ended up `INVALID`. Fix: `aminsert` is now idempotent. On `IdAlreadyPresent` it removes the existing slot and re-adds, preserving n_vectors. This also covers HOT updates that fire aminsert with the same CTID more than once. ### Source - `src/index/insert.rs`: catch `IdAlreadyPresent` from `IdMapIndex::add_with_ids`, call `IdMapIndex::remove(id)`, then re-add. n_vectors stays the same on replace. - `src/lib.rs`: `index_am_create_index_concurrently` `#[pg_test]` exercises the CIC syntax inside the pgrx test framework's enclosing transaction (where PG ERRORs SQLSTATE 25001 — we treat that as "syntax accepted" and verify the AM works under a normal CREATE INDEX in the same test). ### Manual verification (psql, no transaction wrapper) ``` CREATE TABLE cic_demo (id bigint PRIMARY KEY, emb vector); INSERT INTO cic_demo VALUES (1, '[1,0,0,0,0,0,0,0]'), ...; CREATE INDEX CONCURRENTLY cic_demo_idx ON cic_demo USING turbovec (emb vec_cosine_ops); \d cic_demo Indexes: "cic_demo_idx" turbovec (emb vec_cosine_ops) -- valid, no INVALID marker ``` Before v0.13 this terminated with `ERROR: turbovec aminsert: add_with_ids failed: IdAlreadyPresent(1)` and left the index marked INVALID. ### Verified ``` cargo pgrx test pg16 -> 38 ok / 0 failed / 1 ignored ``` ## [0.12.0] — Unreleased ### Phase 12 — forced-index-scan investigation Added a stress test `index_am_forced_index_scan` that calls `SET enable_seqscan = off` to force the planner onto our index path. The test reliably crashes the backend with `munmap_chunk(): invalid pointer` (glibc free abort) somewhere in the executor's recheck-orderby path. Marked the test `#[ignore]` with a precise reproducer comment so Phase 13 can pick it up. During debugging: - Allocated `xs_orderbyvals` / `xs_orderbynulls` in `ambeginscan` (PG core does NOT do this for AMs that advertise `amcanorderbyop = true`). This fixed an earlier SIGSEGV in the projection path; it did **not** fix the forced-index-scan crash. - Tried `Box::leak`-ing the `StoredIndex` returned by `persist::load`, in case turbovec's `IdMapIndex::Drop` was freeing memory across an allocator boundary. Did not help. - Tried setting `xs_recheck = true` in addition to `xs_recheckorderby = true`. Did not help. - Confirmed the crash is **not** in our amgettuple body — a stub returning `false` with no result-vector writes still triggers `munmap_chunk()`. Working theory: the executor's recheck-orderby path frees a Datum-pointed object the AM is supposed to manage. Phase 13 will gdb the crash to identify the exact `free()` call site. ### Workaround for users The planner-picks-naturally path works (37/37 tests pass including the AM). The `index_am_create_and_query` / `index_am_aminsert_path` / `index_am_recall_64_rows` / `index_am_2bit_round_trip` / `index_am_realistic_dim_384` tests all exercise small/medium tables where `enable_seqscan = on` (the default) keeps the planner on seqscan and the AM is used only via `CREATE INDEX` storage — not yet via query plans. For larger corpora, recommend `turbovec.knn()` (same SIMD kernel, no executor-recheck path). ### Source - `src/index/scan.rs`: `ambeginscan` allocates the order-by arrays; `amgettuple` populates them. Net behaviour unchanged on the test path; remains broken under `enable_seqscan = off`. - `src/lib.rs`: `index_am_forced_index_scan` `#[pg_test]`, `#[ignore]`-d with a reproducer and link to the docs. - `docs/INDEXAM.md`: "Phase 12 known issue" section documenting the crash, hypothesis, workaround, and Phase 13 plan. ### Verified ``` cargo pgrx test pg16 -> 30 ok / 0 failed cargo pgrx test pg16 --features experimental_index_am -> 37 ok / 1 ignored ``` ## [0.11.0] — Unreleased ### Phase 11 — realistic-scale tests + 2-bit round-trip + psql regression Proves the index AM scales to real-world dimensionality and to the most-compressed bit_width. ### New tests - **`index_am_realistic_dim_384`** — 200 deterministic 384-dim vectors (typical sentence-embedding dim). Asserts: - `am_storage.n_vectors = 200` after CREATE INDEX. - Self-vector is rank 1 in `ORDER BY emb <=> q LIMIT 1`. - Self-vector lands in top-10. - **`index_am_2bit_round_trip`** — 100 vectors at d=128 with `WITH (bit_width = 2)`. Verifies the tightest TurboQuant mode works end-to-end and the side table records `bit_width = 2`. Self-recall in top-20 (relaxed from top-10 because 2-bit costs ~2 R@k points). ### New psql regression script - `tests/02_index_am.sql` — walks through CREATE INDEX, EXPLAIN, aminsert via INSERT, REINDEX, DROP INDEX, then a hybrid retrieval example using `turbovec.knn(...)` with a SQL-derived allowlist. Run via `cargo pgrx run pg16` then `\i tests/02_index_am.sql`. ### Verified ``` cargo pgrx test pg16 -> 37 ok / 0 failed ``` ## [0.10.0] — Unreleased ### Phase 10 — filtered search via `IdMapIndex::search_with_allowlist` The headline feature from upstream `turbovec`'s API is now wired through to SQL. `turbovec.knn()` gains an optional `allowed bigint[]` argument: ```sql -- Restrict candidates to a tenant or topic without paying the -- cost of a post-filter: SELECT k.id FROM turbovec.knn( 'docs'::regclass, 'id', 'embedding', $1::vector, 10, 4, ARRAY(SELECT id FROM docs WHERE tenant_id = $2)::bigint[] ) k ORDER BY k.score DESC; ``` The SIMD kernel honours the allowlist at 32-vector block granularity — selective filters cost less, not more. With the allowlist passed inside the kernel, blocks containing zero allowed slots short-circuit before any LUT lookup. ### SQL signature ```sql turbovec.knn( rel regclass, id_col text, vec_col text, query vector, k integer, bit_width integer DEFAULT 4, allowed bigint[] DEFAULT NULL ) RETURNS TABLE(id bigint, score double precision) ``` When `allowed` is NULL or omitted, behaviour is identical to v0.9 (unfiltered `IdMapIndex::search`). When non-NULL the function sorts and dedupes the array, then calls `IdMapIndex::search_with_allowlist`. Empty allowlist returns zero rows. ### Source - `src/knn.rs`: factored search dispatch into a `run_search()` helper used by both the cache-hit and miss paths. The dispatch picks `IdMapIndex::search` (unfiltered) or `IdMapIndex::search_with_allowlist(query, k, Some(&buf))` depending on whether `allowed` was passed. - `src/lib.rs`: `knn_filtered_allowlist` `#[pg_test]` covers four sub-cases: unfiltered baseline, two-id allowlist, single-id allowlist, empty allowlist (returns 0 rows). ### Verified ``` cargo pgrx test pg16 -> 35 ok / 0 failed ``` ## [0.9.0] — Unreleased ### Phase 9 — index AM promoted to default + AM scan path uses the cache After v0.7's hardening (32/32 AM tests) and v0.8's cache work, the `turbovec` index access method is promoted out of the experimental feature gate and into the default build: ```toml [features] default = ["pg16", "experimental_index_am"] ``` A stripped-down build without the AM is still available via `cargo build --no-default-features --features pg16`. ### Source - `src/index/scan.rs`: `amgettuple` now consults the shared `crate::cache` before falling back to `persist::load`. On cache hit the scan skips: 1. The `am_storage` row read (one PG round-trip). 2. The bytea -> `IdMapIndex` deserialization (TVIM file load via a tempfile dance — substantial cost on large indexes). Cache validity is the same as the function path: relfilenode + n_vectors, plus LRU under `turbovec.cache_size_mb`. Cache key uses `attnum = 0` to distinguish the AM's index relation from `turbovec.knn()`'s heap-relation entries (which use the column attnum). - `Cargo.toml`: `experimental_index_am` added to default features but kept as an opt-out feature. ### Verified ``` cargo pgrx test pg16 -> 34 ok / 0 failed cargo build --no-default-features --features pg16 -> builds clean ``` ## [0.8.0] — Unreleased ### Phase 8 — backend-local cache for `turbovec.knn()` `turbovec.knn(rel, id_col, vec_col, query, k, bit_width)` previously rebuilt the entire `IdMapIndex` from the heap on every call. v0.8 introduces a backend-local cache keyed by `(rel_oid, attnum, bit_width, dim)`: - **First call** in a backend pays the build cost as before (heap scan via SPI, `IdMapIndex::add_with_ids`). - **Subsequent calls** with the same key, on a relation whose `pg_class.relfilenode` and `count(*)` haven't changed, skip rebuild and reuse the cached `Arc`. - **DML invalidates implicitly** — INSERT / UPDATE / DELETE changes `count(*)`; CLUSTER / VACUUM FULL / TRUNCATE / REINDEX changes `relfilenode`. Either mismatch forces a rebuild on the next lookup. - **LRU eviction** keeps total cache bytes within `turbovec.cache_size_mb` (default 256 MiB; setting to 0 disables caching entirely). ### Source - `src/cache.rs` (NEW, 175 lines) - `CacheKey { rel_oid, attnum, bit_width, dim }`. - `Entry { index: Arc, bytes, relfilenode, n_rows, seq }`. - Public API: `lookup`, `insert`, `invalidate`, `current_relfilenode`, `len`. - LRU enforcement against `turbovec.cache_size_mb`. - `src/knn.rs` rewired: - On entry, computes the cache key and `lookup`s. Hit fast-paths straight to `IdMapIndex::search` on the cached `Arc`. - Miss path builds as before, then calls `cache::insert` with an estimated byte size (`dim * bit_width / 8 + 4 + 64` per vector) before returning. - `src/lib.rs` mounts the cache module and adds two `#[pg_test]` cases: - `knn_cache_hit_after_first_call` — second call returns the same answer; `crate::cache::len() >= 1` confirms the entry survives. - `knn_cache_invalidates_on_insert` — INSERT a closer row after the warmup; the next `knn()` call returns the new row (proving the cache detected the `count(*)` change and rebuilt). ### Verified ``` cargo pgrx test pg16 -> 29 ok / 0 failed cargo pgrx test pg16 --features experimental_index_am -> 34 ok / 0 failed ``` ## [0.7.0] — Unreleased ### Phase 7 — hardened index AM, four new end-to-end tests, real bug fixes The v0.6 index AM passed a single happy-path test. This release adds four more `#[pg_test]` cases that uncovered — and fixed — four real bugs in the AM: - **`index_am_aminsert_path`** — build, insert, query. Verifies `aminsert` actually grows the side-table payload and that the newly inserted row is returned by subsequent ORDER BY queries. - **`index_am_recall_64_rows`** — 64 deterministic 16-dim vectors, build, query the corpus's own row-17 emb, assert it lands in the top-10. (Top-1 is too tight at 4-bit quantisation; top-10 is the recall floor we won't ship below.) - **`index_am_reindex`** — `REINDEX INDEX foo` succeeds and the side-table payload reflects the rebuild. - **`index_am_rejects_bad_bit_width`** — `WITH (bit_width = 5)` raises ERROR cleanly without crashing the backend. ### Bug fixes uncovered by the new tests - **Missing `#[pg_guard]` on AM callbacks** caused a `pgrx::error!` inside `amoptions` ("bit_width must be in 2..=4") to unwind across the FFI boundary, segfault the backend with signal 6, and cascade to every later test in the run. Every `extern "C-unwind"` callback in `src/index/` now wears `#[pg_guard]`. - **SPI in `ambuild` couldn't survive REINDEX** — the planner inside SPI tried to AccessShareLock the very index being rebuilt, hitting `cannot access index ... while it is being reindexed`. Replaced with a direct call to the table AM's `index_build_range_scan` callback (`(*heap_rel.rd_tableam) .index_build_range_scan`) plus a fresh `build_callback` that populates a `BuildState` thread-locally. Same path the built-in btree / GIN / hash AMs use; no SPI lock surface. - **Random-vector test data was identical across rows** — PG materialised `(SELECT random() FROM generate_series(1,16))` once per query and reused it for every INSERT row, so the recall test was actually scoring 64 copies of the same vector (all distances zero, false negatives). Switched to a `hashtext(i::text || ':' || k::text) % 2000 / 1000.0 - 1` per-element formula that's stable per `(i,k)` and varies across rows. ### Source changes - `src/index/build.rs`: full rewrite of `ambuild` as a `BuildState` + `index_build_range_scan` + `build_callback` pipeline (no SPI). The callback validates dim consistency, optionally L2-normalises, and accumulates `(u64, Vec)` rows into the per-build state. - `src/index/{build,cost,insert,options,scan,vacuum,validate}.rs`: every AM callback now has `#[pgrx::pg_guard]`. - `src/lib.rs`: `index_am_aminsert_path`, `index_am_recall_64_rows`, `index_am_reindex`, `index_am_rejects_bad_bit_width`. ### Verified ``` cargo pgrx test pg16 -> 27 passed; 0 failed cargo pgrx test pg16 --features experimental_index_am -> 32 passed; 0 failed ``` This is the first release where `aminsert` and `REINDEX` are actually proven to work. ## [0.6.0] — Unreleased ### Phase 6 — validated against a real PostgreSQL 16 cluster This is the first release where every `#[pg_test]` case has actually been executed and passes. The default-feature build runs **28/28** tests green; the `experimental_index_am`-feature build also runs **28/28**, including a new end-to-end `index_am_create_and_query` test that: 1. `CREATE TABLE`s an 8-dim `vector` column, 2. inserts four rows, 3. `CREATE INDEX ... USING turbovec (... vec_cosine_ops) WITH (bit_width = 4)`, 4. asserts the side-table row was created with `n_vectors = 4`, 5. runs `ORDER BY emb <=> $1 LIMIT 1` and asserts the right row is returned, 6. `DROP INDEX` and verifies the heap is intact. ### Fixes uncovered by running the suite - **Aggregate transition function was implicitly STRICT** (pgrx derives it from non-Option args), causing `CREATE EXTENSION` to fail with `must not omit initial value when transition function is strict and transition type is not compatible with input type`. Both `vec_accum` and `vec_combine` now accept `Option` so pgrx generates non-strict SQL. - **`trusted = true` in `pg_turbovec.control`** was rejected by pgrx 0.17's control-file parser as `RedundantField`. Removed. - **Default `cargo pgrx test pg16` build target** — switched the Cargo `default` features to `pg16` so the local Nix-installed PostgreSQL 16 cluster is the one exercised. Runs against pg17 / pg18 still work via the matching feature flag. - **build.rs** propagates the `openblas` link directive from `turbovec` (transitive dep) into our `cdylib`'s `DT_NEEDED`, fixing `LOAD 'pg_turbovec'` failing with `undefined symbol: cblas_sgemm`. - **Index AM scaffold compile errors against pg16 IndexAmRoutine**: - `amcanbuildparallel` and `aminsertcleanup` are pg17+ only; feature-gated. - `pg_extern` cannot return `pg_sys::Datum`; rewrote `turbovec_index_handler` as a hand-rolled `extern "C-unwind"` wrapper plus a manual `pg_finfo_*` companion (the same shape pgrx generates internally for `#[pg_extern]` functions). - `pg_sys::TupleDescAttr` isn't exposed as a Rust function in pgrx 0.17; rewrote `resolve_indexed_attr` to use `(*tupdesc).attrs.as_slice(natts)`. - `(*indrel).indkey.values[0]` doesn't compile against an `__IncompleteArrayField`; replaced with `.as_slice(nkey)`. - `Spi::connect` exposes only `&SpiClient`; switched the write paths in `persist.rs` to `Spi::connect_mut`. - Implicit autoref on `(*opaque).results[(*opaque).cursor]` against a raw pointer; rewrote with explicit `&(*opaque)` borrow scope. - **Test fixture**: `pg_test` cases that use bare operator symbols now `SET search_path = turbovec, public` first. ### Added - `docs/BUILDING.md` documenting the Nix-specific build dance (writable pg_config wrapper, libclang / glibc include flags, openblas RUSTFLAGS, ICU sidestep). - `index_am_create_and_query` `#[pg_test]` case (gated by the `experimental_index_am` Cargo feature). ### Changed - Default Cargo `default` features set to `["pg16"]` (was `["pg17"]`) to match the local development cluster. ## [0.5.0] — Unreleased ### Added — Phase 5: pgvector-parity helpers - **`subvector(vector, start integer, length integer) -> vector`** — 1-indexed slice. Bounds-checked; raises `ERROR` on overrun. - **`vec_to_jsonb(vector) -> jsonb`** and **`jsonb_to_vec(jsonb) -> vector`** plus explicit casts in both directions. Useful for replication via JSONB columns, logging, and audit trails. - **`vec_check_dim(vector, integer) -> vector`** — runtime dim assertion. Use as a `CHECK` constraint when typmod-style enforcement is wanted without the full typmod plumbing. - **`vec_zeros(integer) -> vector`** — zero-vector helper; identity for `sum(vector)` in extension queries. - **`vec_to_text(vector) -> text`** — explicit text rendering callable from SQL (the type's OUTPUT function as a regular function). ### Tests - `subvector_basic`, `subvector_out_of_bounds`, `jsonb_round_trip`, `check_dim_passes_and_fails`, `zeros_helper`. ### Changed - `Cargo.toml` / `pg_turbovec.control` bump to `0.5.0`. - `migrations/004_pg_turbovec_v0.5.0.sql` reference mirror. ## [0.4.0] — Unreleased ### Added — Phase 4: experimental `turbovec` index access method (opt-in) A full `IndexAmRoutine`-based access method is now scaffolded under `src/index/`, gated behind the **`experimental_index_am`** Cargo feature. Default builds **do not** include it; the v0.3 surface (type, operators, aggregates, `turbovec.knn()`) remains the only stable user-facing API. **Build:** ```bash cargo pgrx install --release --features experimental_index_am ``` **Use:** ```sql CREATE INDEX docs_emb_idx ON docs USING turbovec (embedding vec_cosine_ops) WITH (bit_width = 4); SELECT id FROM docs ORDER BY embedding <=> $1 LIMIT 10; ``` #### Source layout (`src/index/`) - `mod.rs` — `IndexAmRoutine` populator and the `turbovec_index_handler(internal) RETURNS index_am_handler` SQL function. Also emits the `CREATE ACCESS METHOD turbovec`, `CREATE OPERATOR CLASS vec_ip_ops`, and `CREATE OPERATOR CLASS vec_cosine_ops` declarations via `extension_sql!`. - `options.rs` — `bit_width` (2…=4) and `dim` (0 = auto, else positive multiple of 8) reloption parsing under the AM-side callback `amoptions`. - `persist.rs` — SPI-backed read/write of `turbovec.am_storage (indexrelid, bit_width, dim, n_vectors, payload, version, updated_at)`. `payload` is `STORAGE EXTERNAL` (no PGLZ on already-quantised bytes). - `build.rs` — `ambuild` (heap scan via SPI, builds `IdMapIndex`, persists) and `ambuildempty` (writes empty marker). - `insert.rs` — `aminsert` (load-then-update; v0.5 will batch). - `scan.rs` — `ambeginscan` / `amrescan` / `amgettuple` / `amendscan` with a `ScanOpaque` carrying the query vector and cached result list. ORDER-BY-only scans are required. - `vacuum.rs` — `ambulkdelete` / `amvacuumcleanup` stubs (Phase 5 needs an upstream way to enumerate live ids in `IdMapIndex`). - `cost.rs` — `amcostestimate` constant heuristic so the planner picks us over a full sort. - `validate.rs` — `amvalidate` returns `true` (Phase 5 will check opclass strategy numbers). #### CTID encoding We use pgrx's canonical 32 / 16 packing (`item_pointer_to_u64`): block number in the top 32 bits, offset number in the bottom 16, upper 16 reserved for a future epoch. This gives `IdMapIndex` u64 ids natural ordering inside a relfile and lets `amgettuple` fill `xs_heaptid` via `u64_to_item_pointer` directly. #### Capability flags ```rust amstrategies = 0 amsupport = 1 amcanorder = false amcanorderbyop = true amcanbackward = false amcanunique = false amcanmulticol = false amoptionalkey = true amstorage = true amcanparallel = false // Phase 5 amcanbuildparallel = false // Phase 5 amusemaintenanceworkmem = true ``` #### Status **Untested against a running cluster.** This release is the complete scaffold ready for a Phase 5 session that has `cargo-pgrx` and a Postgres dev cluster: `cargo pgrx test pg17 --features experimental_index_am` is the gate. Known follow-ups are enumerated in `docs/INDEXAM.md` § "Test plan" and § "Known risks". ### Added — docs - `docs/INDEXAM.md` — implementation guide for the index AM (callback responsibilities, side-table schema, test plan, known risks). - `migrations/003_pg_turbovec_v0.4.0.sql` — reference mirror of the SQL surface that ships only when the feature is enabled. ### Changed - `Cargo.toml` adds `libc = "0.2"` (used by `persist.rs` for pid-stamped tempfile paths) and the `experimental_index_am` Cargo feature. - `pg_turbovec.control` `default_version` bumped to `0.4.0`. - `src/lib.rs` mounts `mod index` only under `#[cfg(feature = "experimental_index_am")]`. ## [0.3.0] — Unreleased ### Added — Phase 3: kernels module, benches, CI, docs - **`src/kernels.rs`** — pure-Rust math kernels (`dot`, `l2_sq`, `l1_abs`, `norm2`, `cosine_distance`, `normalise_into`, `normalise_to_vec`). Distance and normalisation code in `distance.rs` / `normalize.rs` now delegate to this module so the kernels are exercisable under plain `cargo test --no-default-features` without booting Postgres. - **`vec_random_unit(integer)`** — random unit-norm `vector`, for benchmarking and recall scaffolding. - **`benches/distance.rs`** — `criterion`-based micro-benchmarks of the distance kernels at d=128, 384, 768, 1536, 3072. Runs via `cargo bench --bench distance --no-default-features`. - **Codeberg Woodpecker CI** (`.woodpecker/ci.yaml`) — three pipelines: pure-Rust unit tests + clippy on every push; `cargo pgrx test pg17` on `main` / release branches. - **`docs/USAGE.md`** — cookbook with install, exact search, ANN via `turbovec.knn()`, aggregates, arithmetic, GUCs, pgvector coexistence migration, diagnostics. - **`docs/RECALL.md`** — recall/perf benchmark methodology, matched-bit-budget comparison plan against pgvector for v0.4. - **Pure-Rust unit tests** in `kernels::tests` covering every kernel plus a precision regression (1 048 576-element sum of squares stays within 1 ppm of the f64 truth). ### Changed - `Cargo.toml` adds `rand = "0.8"`, `criterion = "0.5"` (dev), declares `[[bench]] name = "distance"`. - `pg_turbovec.control` `default_version` bumped to `0.3.0`. ## [0.2.0] — Unreleased ### Added — Phase 2: function-driven ANN - **`turbovec.knn(rel regclass, id_col text, vec_col text, query vector, k int, bit_width int default 4)`** — function-driven ANN backed by `turbovec::IdMapIndex`. Returns `TABLE(id bigint, score float8)`, ordered by score DESC for most-similar-first. - Optional unit-normalisation via `turbovec.normalize_on_insert` GUC; constraints `k > 0`, `bit_width ∈ {2,3,4}`, `dim % 8 == 0`. - `migrations/002_pg_turbovec_v0.2.0.sql` reference mirror. - `#[pg_test]` cases for `knn_returns_nearest_first` and `knn_rejects_bad_k`. ### Removed - `src/phase2_knn.rs` scaffold — promoted to mounted `src/knn.rs`. ### Added — Phase 1: type, operators, functions, aggregates - **`vector` type** — variable-dimension `f32` vector, stored as a CBOR-serialised varlena via `pgrx::PostgresType`. Text I/O accepts `'[1, 2, 3]'` with whitespace tolerance and rejects NaN / ±Inf. Hard cap at 16 000 dimensions, matching pgvector. - **Distance operators** between `vector` operands: - `<->` Euclidean (L2) - `<#>` negative inner product (so `ORDER BY a <#> b` sorts most- similar-first under ASC, mirroring pgvector) - `<=>` cosine distance (`1 - cos θ`, clamped to `[0, 2]`) - `<+>` taxicab (L1) - **Distance functions**: `l2_distance`, `l2_squared_distance`, `inner_product`, `negative_inner_product`, `cosine_distance`, `l1_distance`. - **Helper functions**: `vector_dims`, `vector_norm`, `vec_normalize`. - **Element-wise arithmetic**: `vec_add` (`+`), `vec_sub` (`-`), `vec_mul` (`*`). - **Aggregates**: `avg(vector)` and `sum(vector)`. Internal state uses `f64` accumulators to preserve precision on large corpora. Both are `PARALLEL SAFE`; `combinefn` merges partial states. - **Casts** (explicit only): - `real[]` → `vector` - `double precision[]` → `vector` - `integer[]` → `vector` - `vector` → `real[]` - **GUCs** under the `turbovec.*` namespace: - `bit_width_default` (int, default 4, range 2..=4) - `cache_size_mb` (int, default 256, range 0..=65536) - `warn_on_rebuild` (bool, default true) - `search_concurrency` (int, default 1, range 1..=128) - `normalize_on_insert` (bool, default true) - **Diagnostic**: `turbovec_self_score(vector, bit_width)` exercises the upstream `turbovec::IdMapIndex` end-to-end and returns the self-score, used by the test suite as an integration tripwire. ### Tests - `#[pg_test]` cases in `src/lib.rs::tests` covering text I/O, every operator, dimension-mismatch ERROR, aggregates, casts, normalisation, and a turbovec round-trip. - `tests/01_type_basic.sql` — psql-style regression script. ### Project layout - `pgrx = "=0.17.0"` to match the cached toolchain. - `pg_turbovec.control` declares schema `turbovec`, `relocatable = false`, `trusted = true`. - `migrations/001_pg_turbovec_v0.1.0.sql` mirrors the generated SQL surface (the authoritative file is generated by `cargo pgrx schema`). ### Not yet shipped (Phase 2 / Phase 3) - Index access method `turbovec` and operator classes `vec_ip_ops`, `vec_cosine_ops`. A starter is checked in at `src/phase2_knn.rs` (not yet mounted by `lib.rs`). - Filtered search via `IdMapIndex::search_with_allowlist`. - Binary-compatible varlena layout with pgvector's `vector`. - WAL-logged persistent index pages. [1.0.0-rc.2]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v1.0.0-rc.2 [1.0.0-rc.1]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v1.0.0-rc.1 [0.16.0]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v0.16.0 [0.15.0]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v0.15.0 [0.14.0]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v0.14.0 [0.13.0]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v0.13.0 [0.12.0]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v0.12.0 [0.11.0]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v0.11.0 [0.10.0]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v0.10.0 [0.9.0]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v0.9.0 [0.8.0]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v0.8.0 [0.7.0]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v0.7.0 [0.6.0]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v0.6.0 [0.5.0]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v0.5.0 [0.4.0]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v0.4.0 [0.3.0]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v0.3.0 [0.2.0]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v0.2.0 [0.1.0]: https://codeberg.org/gregburd/pg_turbovec/releases/tag/v0.1.0