# pg_turbovec Open-source vector similarity search for PostgreSQL, backed by Google Research's [TurboQuant](https://arxiv.org/abs/2504.19874) algorithm via the [`turbovec`](https://crates.io/crates/turbovec) Rust crate. Store compact vector indexes alongside your PostgreSQL data: **1-bit sign-BQ with full-precision reranking**, or **2/3/4-bit TurboQuant**. The original vectors stay in the table for reranking. Supports: - approximate nearest-neighbour search with full-precision candidate reranking; exact search through PostgreSQL's distance operators and a sequential scan - **1-bit + rerank**, including IVF: `WITH (bit_width = 1, lists = N)`; see [the 1-bit example](#1-bit-search-with-full-precision-reranking) - **three index kinds**: a **flat** quantized scan (exact-recall- capable), an opt-in **IVF** layer (`WITH (lists = N)`) that is out-of-core end-to-end so a larger-than-RAM index can be built and queried (a **Vamana navigable-graph** kind, `WITH (graph = true)`, also exists but is **deprecated** — see below) - single-precision (`f32`) vectors with on-disk compression to ~16× smaller than `pgvector` at 4-bit - L2 distance (`<->`), inner product (`<#>`), cosine distance (`<=>`), L1 distance (`<+>`) - filtered / hybrid ANN — partial index, **in-kernel allowlist** (a selective filter gets *cheaper*, not more expensive), or iterative scan ([guide](docs/FILTERING.md)) - multivector & dense+sparse hybrid — ColBERT-style MaxSim re-rank (`max_sim`) + reciprocal rank fusion (`rrf_score`) + named-vector schema pattern ([guide](docs/HYBRID_SEARCH.md)) - pgvector-compatible function names (`to_vector`, `array_to_vector`, `subvector`, `vector_dims`, `vector_norm`, `inner_product`, `l2_distance`, `cosine_distance`, `l1_distance`) - any [language](https://www.postgresql.org/docs/current/external-pl.html) with a Postgres client Plus [ACID](https://en.wikipedia.org/wiki/ACID) compliance, point-in- time recovery, JOINs, GUCs, parallel-safe aggregates, and all of the [other great features](https://www.postgresql.org/about/) of Postgres. [![Rust 1.96+](https://img.shields.io/badge/rust-1.96+-93450a)](https://www.rust-lang.org/) [![PostgreSQL 13-19](https://img.shields.io/badge/postgres-13--19-336791)](https://www.postgresql.org/) [![Apache 2.0](https://img.shields.io/badge/license-Apache_2.0-blue.svg)](LICENSE) > **Status:** v2.11.0 - built on upstream **turbovec 1.1.1** (wire > format v8; staged 2/4-bit search on aarch64 and AVX-512 VBMI+VNNI hosts, > 1.14-1.17x faster 4-bit end-to-end on Graviton4 at 1M x 1024-d -- see > [CHANGELOG](CHANGELOG.md)). The full `#[pg_test]` suite passes against > PostgreSQL 13, 14, 15, 16, 17, and 18 (and 19beta1, experimentally). > **v2.0.0 is a MAJOR wire-format break (v7 → v8): upgrading from any > 1.x requires `ALTER EXTENSION ... UPDATE` then a one-time `REINDEX > INDEX` per turbovec index** (a pre-v8 index ERRORs at first scan with > a REINDEX hint — never silent). It is materially faster than the 1.29 > line at the same storage. See [`docs/UPGRADING.md`](docs/UPGRADING.md) > for the migration and [`docs/PARITY_GAPS.md`](docs/PARITY_GAPS.md) for > the honest scoreboard vs pgvector. ## Why pg_turbovec? **On 1 M × 1536-d real OpenAI embeddings, pg_turbovec matches pgvector HNSW's recall at ~10–20× less on-disk storage, with exact re-ranking against the heap.** Storage efficiency and exact/near-exact recall are what pg_turbovec does better than anything else — not raw latency (see the honest latency note below). Head-to-head, warm cache, release build. Storage + recall (measured on real embeddings): | metric (1 M × 1536-d, real OpenAI) | pg_turbovec 2-bit | pgvector HNSW | |---|---:|---:| | **On-disk index / vector** | **≈ 412 B (measured)** | ≈ 8 192 B | | **Storage @ 1 M** | **≈ 412 MB** | ≈ 8 GB (~20× larger) | | Build (1 M) | ~minutes | ~5 min | | Recall@10 (IVF, tuned) | ~0.90 (R@100 → 0.99) | 0.96–0.99 | | Exact re-ranking vs heap | ✓ (`xs_recheckorderby`) | ✓ | ### Common objections, answered with measurements Four things evaluators say about pg_turbovec. Two are misunderstandings we caused, one is true-and-fixed, one is true-and-stands. #### "pg_turbovec doesn't support ANN" It does. The default (`lists = 0`) scans all **quantized codes**, selects a candidate set, and reranks candidates using the original vectors. Quantization can exclude a true neighbour from that set, so even this flat index is an approximate search. IVF additionally restricts candidate search to selected cells: ```sql -- Flat quantized search with full-precision candidate reranking (DEFAULT). CREATE INDEX ON items USING turbovec (embedding vec_cosine_ops); -- Approximate, cell-pruned IVF scan. This is the ANN path. CREATE INDEX ON items USING turbovec (embedding vec_cosine_ops) WITH (lists = 1000); ``` `lists` defaults to `0`, which means flat. Nothing in a default build tells you IVF exists, which is a documentation failure on our side, not a misreading on yours. Since v2.10.2 a flat build over 100 k rows emits a `NOTICE` pointing at the ANN option. See [Choosing `lists`](#choosing-lists--the-ann-tuning-knob) for how to pick a value — and note that on our own measurements, **flat is the right choice more often than you would expect** (that section says when). #### "HNSW has a high memory footprint and is slow to build" **Agreed — that is why we did not build on HNSW.** Measured head-to-head on the same host, 10 M × 1536-d real embeddings ([`docs/RECALL.md`](docs/RECALL.md)): | | pgvector HNSW (m=16, ef_construction=64) | pg_turbovec 4-bit | |---|---:|---:| | Index size | 65.5 GiB | **14.9 GiB** (4.4× smaller) | | Build time | 3 h 38 min | **1 h 24 min** (2.6× faster) | Our build log also shows HNSW slowing **super-linearly past 5 M rows** even with the whole corpus in page cache, which matches the common complaint. #### "pg_turbovec's own build memory was worse than HNSW's" **This was true, and it is now fixed.** The same 10 M × 1536-d benchmark measured pg_turbovec's build at **121 GiB peak + 60 GiB swap** against HNSW's 16.9 GiB — we were ~7× *worse*. That figure is from **v1.5.1 (May 2026)** and is still quoted in `docs/RECALL.md` for the run it belongs to. Root cause (found and fixed in **v2.10.1**): `vector` is stored as CBOR, so `FromDatum` palloc'd a decoded buffer per row, and the build callback had **no memory-context management at all** — every decoded row accumulated for the whole scan. The spill was streaming the corpus to disk while PostgreSQL held a decoded copy of *all of it* in RAM. Measured after the fix: | build | before | after | |---|---|---| | 2 M × 1024-d peak | 12.16 GiB | **3.45 GiB** (3.5× less, 16 % faster) | | 10 M × 1024-d on a 61 GiB host | **OOM-killed** (`anon-rss` 58.5 GiB) | **completes at 11.20 GiB** in 70 min | **Scope, stated plainly:** the post-fix numbers are 1024-d. We have **not** re-measured 10 M × **1536-d**, which is what the 121 GiB figure was, so treat that specific number as superseded-in-mechanism but not yet re-measured at its original dimension. #### "HNSW is faster on query latency" **True, and we don't dispute it.** A flat pg_turbovec scan is `O(n·dim)`; HNSW is sublinear. If p50 latency is your binding constraint and your corpus fits in RAM, **pgvector HNSW is the better choice today** — see the honest latency note immediately below, which we keep at the top of this README on purpose. Where we are measurably better: **storage** (4.4–20× smaller), **build time** (2.6× faster), **larger-than-RAM corpora** (IVF is out-of-core end-to-end), and **exact recall** when you need R@10 = 1.000 rather than 0.96. --- ## Choosing `lists` — the ANN tuning knob `lists` is the IVF coarse-cell count (`nlist`). **The default is `0`, which means a flat quantized scan, not guaranteed exact top-k.** We keep it at `0` deliberately: on our own measurements, enabling IVF at the default `bit_width = 4` makes things *worse* up to at least 1 M rows. ### First: do you want IVF at all? At 1 M × 1024-d with the default `bit_width = 4` ([`docs/BQ_RECALL_BENCH.md`](docs/BQ_RECALL_BENCH.md) § 0.6g): | config | latency | recall@10 | |---|---:|---:| | **flat** (`lists = 0`, the default) | **6.08 ms** | **1.000** | | `lists = 1024` (≈ √n) | 16.04 ms (62 % slower) | **capped at 0.959** | That recall cap is the important half: at `probes = 128`, widening the rerank window from 32 to 2000 left recall at **exactly 0.959** across all eight windows. It is a per-probe ceiling, it is CPU-independent, and **no amount of tuning removes it.** So IVF is not a free "make it faster" switch — it trades a recall ceiling for a smaller scan. Our measured decision rule: | your situation | use | |---|---| | `bit_width ≥ 2` (incl. the default 4), n ≲ 1 M | **`lists = 0`** (flat) — faster *and* exact | | `bit_width = 1`, n ≳ 1 M, target recall ≲ 0.95 | **`lists ≈ √n`** — measured 38–47 % faster | | `bit_width = 1`, target recall ≳ 0.98 | **`lists = 0`** — IVF cannot reach it | | index larger than RAM | **`lists ≈ √n`** — IVF is the only out-of-core query path | | n ≫ 1 M with `bit_width ≥ 2` | **measure it** — we have no data above 1 M here | ### Second: if you do want it, pick `lists ≈ √n` ```sql -- 1 M rows -> lists = 1000; 10 M -> 3162; 100 M -> 10000. SELECT round(sqrt(count(*)))::int AS suggested_lists FROM items; CREATE INDEX ON items USING turbovec (embedding vec_cosine_ops) WITH (lists = 1000); -- substitute the value above ``` **`√n` is a ceiling, not a target.** At 1 M we measured `lists = 4096` (4× √n) as worse than `lists = 1024` on *every* axis: **11× the build time** and ~50 % higher latency. Each cell holds proportionally fewer rows, so a given recall needs proportionally more `probes` — you pay twice. If in doubt, go *under* `√n`. ### Third: tune `probes` at query time, not `lists` `lists` is baked in at build time; `probes` is a per-session GUC, so this is the knob to sweep: ```sql SET turbovec.probes = 16; -- cells scanned per query (default 16) SET turbovec.oversample = 2.0; -- widen the exact-rerank window ``` Recall rises with `probes` and saturates at the per-probe ceiling described above. Sweep `probes` (8, 16, 32, 64, …) against **your own** ground truth and stop at the first value that meets your recall target; everything beyond that is latency you are paying for nothing. ### Fourth: measure on your own data Our 1 M and 250 k figures come from *different corpora*, so cross-scale deltas in our docs are suggestive rather than measured. Your corpus is the only authority for your workload. [`docs/PRODUCTION.md`](docs/PRODUCTION.md) has a 20-minute experiment that settles flat-vs-IVF on your data, comparing **at matched recall** (comparing at matched `probes` instead is the classic way to get a flattering, wrong answer). ## Honest latency note (read this) > **v2.11.0 (turbovec 1.1.1):** on aarch64 (Graviton, Ampere, Apple) and x86 > with AVX-512 VBMI+VNNI, flat 4-bit queries are **1.14–1.17× faster > end-to-end** at 1M × 1024-d (Graviton4: 7.88 → 6.72 ms at `search_k=32`), > recall unchanged; the scan kernel itself is 3–6.5× faster, so the > per-candidate heap recheck is now most of a query's cost. AVX2-only x86 is > unchanged. [docs/BENCHMARKS.md](docs/BENCHMARKS.md#turbovec-111-staged-search-on-graviton4-v2110-2026-10-05). > v2.11.0 also fixes a **silent index-entry corruption under concurrent writes > + VACUUM** present since v1.29.1 -- upgrade, then `REINDEX` indexes that took > writes from long-lived connections. See [CHANGELOG](CHANGELOG.md). pg_turbovec's **flat** kind is an `O(n·dim)` quantized full scan — at 1 M rows its warm p50 is **~2.5 s on AVX2**, and pgvector HNSW (a sublinear graph) is **~490× faster** on latency. **We do not beat HNSW on latency**, and any earlier README claim to the contrary was a pre-AVX2-fix measurement artifact and is retracted (see [`docs/PARITY_GAPS.md`](docs/PARITY_GAPS.md)). The latency paths are: - **IVF** (`WITH (lists = N)`) — cell-pruned scan, out-of-core capable. Measured ~17–24 ms warm p50 at 1 M × 1536-d (R@10 0.84–0.89), the practical latency config. - **Vamana graph** (`WITH (graph = true)`) — **DEPRECATED, scheduled for removal.** It was added to chase HNSW's query latency while keeping TurboQuant's compression, but measured at *matched recall* it never delivers: its apparent sublinearity holds only at iso-*beam*, and once recall is held equal the curves diverge rather than cross. Use the default flat index below ~1 M vectors, or `WITH (lists = N)` (IVF) at scale. | corpus / target | flat | IVF | graph | |------------------------|-----------------|---------------------|--------------| | SIFT-1M/128d, R@10 ≥0.95 | 0.98 ms, qps@8 1380 | 1.8 ms, qps@8 2039 | 26.2 ms, qps@8 299 | | GIST-1M/960d, R@10 ≥0.95 | 5.88 ms, qps@8 279 | 11.3 ms, qps@8 480 | **unreachable** | | GIST-10M/960d, R@10 ≥0.98 | 34.2 ms, qps@8 31 | 28.4 ms, qps@8 161 | **unreachable** | It also loses on build time (57–90×), storage, and has no out-of-core path. It is **IVF, not the graph**, that beats flat's O(n) wall. Use pg_turbovec when your workload is **cosine / inner-product semantic search that is storage-constrained** and you want full-precision re-ranking and Postgres-native ACID/joins — and you use the IVF kind (`WITH (lists = N)`, not a bare flat scan) for latency at scale. If you need raw HNSW latency or ANN over halfvec/sparse representations, pgvector + HNSW is the right pick today. Feature breakdown: | Feature | pg_turbovec | pgvector + HNSW | |----------------------------------|-------------------|-------------------| | Storage / 1536-dim row (2-bit) | **≈ 412 B (measured)** | 8 192 B (measured) | | Latency at 1 M+ (raw) | slower (flat O(n); IVF/graph prune) | **faster (sublinear HNSW)** | | Index kinds | flat, IVF (out-of-core), ColBERT, Vamana graph | HNSW, IVFFlat | | Filtered search | In-kernel SIMD allowlist | Post-filter | | Index AM lifecycle | CREATE / CIC / aminsert / ambulkdelete / VACUUM / REINDEX | same | | Distance ops indexed | `<#>` `<=>` `<->` `<+>` (turbovec kernel) | `<->` `<#>` `<=>` `<+>` (HNSW + IVF) | | L2 / L1 distance ANN | indexed (`vec_l2_ops` / `vec_l1_ops`) | indexed via HNSW | | Halfvec, sparsevec, bitvec types | ✓ (exact ops; not AM-indexable) | ✓ | | License | Apache-2.0 | PostgreSQL | *Methodology: recall numbers use exact top-k ground truth; see [`docs/RECALL.md`](docs/RECALL.md). Latency numbers are contention- controlled AVX2 warm p50 (`docs/PARITY_GAPS.md`, `docs/BENCHMARKS.md`).* ## Why pg_turbovec instead of `binary_quantize() + bit_hamming_ops`? If you've reached for pgvector's `binary_quantize()` + `bitvec` + `bit_hamming_ops` HNSW index, the reason is almost always memory pressure: "I have 100 M × 1536-dim embeddings, FP32 doesn't fit, I'll trade recall for 32× compression." **At the same byte budget, pg_turbovec's 2-bit mode wins on recall.** Measured numbers from [`benches/results/recall_dbpedia_1M_2026_05_24.json`](benches/results/recall_dbpedia_1M_2026_05_24.json) on 1 M × 1536-d OpenAI ada-002 embeddings; the 1-bit Hamming line is the upper bound from the upstream pgvector docs since we don't have a direct 1 M-row measurement on the same corpus. | Approach | Bytes / 1536-dim row | R@10 (real OpenAI ada-002 embeddings) | |---|---:|---:| | FP32 (raw `vector`) | 6 144 | 1.00 (ground truth) | | FP16 (`halfvec`) | 3 072 | ≈ 1.00 | | TurboQuant 4-bit (`turbovec` index, default) | 780 (measured payload / 1 M rows) | **1.000 (search_k = 100)** | | TurboQuant 2-bit (`turbovec` index) | 396 (measured payload / 1 M rows) | **1.000 (search_k = 100)** | | 1-bit + Hamming HNSW (pgvector `bit_hamming_ops`) | 192 | ≈ 0.65-0.75 (literature) | The 4-bit / 2-bit numbers come from the 50-query head-to-head sweep in [`docs/RECALL.md § 2.2`](docs/RECALL.md); the synthetic random-vector measurements in [`docs/RECALL.md § 2.1`](docs/RECALL.md) are deliberately pessimistic because random points have no clustering structure to exploit - the dbpedia run shows what real embedding geometry buys you. **Why does Lloyd-Max scalar quantization beat 1-bit thresholding at the same byte count?** TurboQuant first rotates the input by a fixed orthogonal matrix so that, after rotation, each coordinate independently follows a known Beta distribution that converges to N(0, 1/d). It then assigns buckets via Lloyd-Max scalar quantization - provably the *distortion-rate-optimal* scalar code for that distribution - and packs them at 2, 3, or 4 bits per coordinate. 1-bit thresholding (pgvector's `binary_quantize()`) is the same idea pinned to `bit_width = 1`: it keeps the sign and throws the magnitude away. At 2 bits, Lloyd-Max with 4 reconstruction levels lands materially closer to the Shannon distortion-rate lower bound than a 2-bucket sign threshold can - so pg_turbovec at `bit_width = 2` occupies essentially the same byte budget as 1-bit Hamming with strictly higher recall. See the [TurboQuant paper, arXiv:2504.19874](https://arxiv.org/abs/2504.19874) for the full distortion analysis. **When is `bit_hamming_ops` still the right tool?** When the bit vector is *the data*, not a compression of an `f32` vector - i.e. native binary embeddings (Cohere's binary mode), perceptual / image fingerprints (pHash, dHash for near-duplicate detection), and SimHash / MinHash signatures over text shingles. For those workloads pgvector's HNSW on `bit_hamming_ops` is the right tool and there is no reason to use pg_turbovec instead. ## Choose your `bit_width` ### 1-bit search with full-precision reranking **pg_turbovec supports 1-bit binary quantization followed by reranking.** `bit_width = 1` uses centered sign-BQ in pg_turbovec, separate from upstream TurboQuant's 2/3/4-bit codec. Hamming distance selects candidates; PostgreSQL fetches their original vectors from the heap and recomputes the requested distance. Reranking is automatic on the index-AM query path: no separate binary-vector column or manual reranking CTE is required. ```sql -- Assuming items(id, embedding) contains turbovec.vector values. CREATE INDEX items_embedding_bq ON items USING turbovec (embedding turbovec.vec_cosine_ops) WITH (bit_width = 1, lists = 0); BEGIN; SET LOCAL turbovec.search_k = 800; -- example candidate budget; tune for your data SET LOCAL turbovec.oversample = 1.0; SELECT id FROM items ORDER BY embedding OPERATOR(turbovec.<=>) (SELECT embedding FROM items WHERE id = 42) LIMIT 10; COMMIT; ``` Use `EXPLAIN` to confirm index use. `search_k` controls the initial candidate budget, `oversample` multiplies it, and `hi_dim_rerank` may raise its floor. **Exact candidate distances do not guarantee exact top-k recall:** reranking cannot recover a neighbour omitted by Hamming selection. `WITH (bit_width = 1, lists = 1000)` also supports IVF with reranking. Tune `probes` for cell coverage as well as the candidate budget; increasing only `search_k` cannot recover neighbours in unprobed cells. Keep `lists = 0` unless measurements justify IVF. See the storage/recall trade-offs below. | Workload | Recommended | Storage / 1536-dim (measured) | R@10 (1 M dbpedia) | |---|---|---:|---:| | Want pgvector-equivalent recall, halve storage | `halfvec` (no quantization) | 3 072 B | ≈ 1.0 | | RAG / semantic search, R@10 ≥ 0.95 acceptable | **`bit_width = 4` (default)** | 780 B | 1.000 | | Memory pressure dominates, R@10 ≥ 0.85 acceptable | `bit_width = 2` | 396 B | 1.000 | | Replacing `binary_quantize() + bit_hamming_ops` | `bit_width = 2` (strictly better) | 396 B | 1.000 (vs 0.65-0.75) | | Absolute minimum storage, latency has slack | `bit_width = 1` (sign BQ, opt-in) | 192 B | 0.967 @ w=256 (measured, 1024-d) | **`bit_width = 1`** (new in v2.6.0) is a different scheme from the 2/3/4-bit TurboQuant path: it keeps only each coordinate's sign (after subtracting the corpus mean), scores by Hamming, and relies on the exact re-rank for accuracy. It is the smallest option — `dim/8` bytes and no per-vector scale — and deliberately **opt-in, never a default**: it is lossy enough that `turbovec.hi_dim_rerank` widens the exact-rerank window for it at any dimension. A corpus whose vectors all share one sign pattern even after centering is rejected at build rather than silently returning arbitrary rows. **Measured** on 250k x 1024-d Cohere-wiki (AVX2 host, 100 held-out queries, exact ground truth): storage is **3.98x smaller** than 4-bit and **2.02x** smaller than 2-bit, but at matched recall it costs **2.7-6.1x the latency** and needs a **25x wider** exact-rerank window than 2-bit to clear R@10 >= 0.99 (window 800 vs 32). Use it where storage is the binding constraint and latency has slack -- never as a default. **1-bit is a HIGH-DIMENSION technique.** A dim sweep (256/512/1024-d, same corpus) shows the penalty collapsing as dimension rises: the rerank window it needs versus 2-bit for R@10 >= 0.95 goes **125x (256-d) -> 25x (512-d) -> 8x (1024-d)**, and its storage edge improves too (1.90x -> 1.97x vs 2-bit). At 256-d it needs to rerank 6.4% of the corpus to reach R@10 >= 0.99, which makes it **effectively unusable at 256-d and below**. Prefer it at 768-d and up. The **storage, recall and latency ratios** are all confirmed: the original timings were taken on a host that could not reach the harness's load gate, so the sweep was re-run after that was fixed. Recall reproduced exactly, the contended p50s proved uniformly 14-16% pessimistic, and **the ratios held to two decimal places** (2.70 -> 2.75x and 6.13 -> 6.09x). Absolute milliseconds quoted below are therefore ~15% conservative. Full curve, the resolution, and the pre-registered predictions: [`docs/BQ_RECALL_BENCH.md`](docs/BQ_RECALL_BENCH.md) § 0. It composes with `lists = N` (IVF): `WITH (lists = N, bit_width = 1)` stores the sign codes cell-contiguous and probes only `turbovec.probes` cells, combining the storage win with the scan win. **When IVF actually pays for 1-bit** (measured on a real 1M x 1024-d Cohere corpus, AVX-512): at **1M rows** IVF beats flat by **47% at R@10 >= 0.90** and **38% at R@10 >= 0.95** -- but it CANNOT reach R@10 >= 0.99 at all, because cell-restricted search caps recall per probe count (0.986 max at probes=128). At 250k rows flat wins at every target. And for `bit_width >= 2` flat wins everywhere up to 1M, because 2-bit only needs a 32-wide rerank window so its full scan is already cheap. So: **1-bit + n >= ~1M + target <= ~0.95 -> use lists = N; otherwise flat.** > That is guidance about **when IVF is worth enabling**, not about what is > supported. `WITH (lists = N)` composes with **every** `bit_width` and always > has -- 4-bit IVF is the original IVF path, out-of-core end-to-end since > v1.13.0. The only combination the code rejects is `bit_width = 1` with > `graph = true`. If you are on 4-bit and want IVF, nothing is blocking you; > see [`docs/BQ_RECALL_BENCH.md`](docs/BQ_RECALL_BENCH.md) 0.6g for what to > expect and what has actually been measured. Note an `INSERT` into an existing IVF+BQ index appends rather than placing the row in its cell, which degrades that index to a flat Hamming scan until the next `REINDEX` — reportable via `turbovec.index_is_degraded()`. Measured storage and recall come from the head-to-head sweep on 1 M × 1536-d OpenAI ada-002 embeddings; methodology and the synthetic random-vector numbers (§ 2.1) live in [`docs/RECALL.md`](docs/RECALL.md). For dimensions other than 1536, multiply storage through by `dim / 1536`. ## Installation `pg_turbovec` requires PostgreSQL 13–18 (19beta1 experimental) and a Rust toolchain ≥ 1.96 (pgrx 0.19's MSRV; the default stable toolchain works). PostgreSQL 16 is the reference development platform; 13/14/15/17/18/19 are tested in CI. ```bash # One-time setup. cargo install --locked cargo-pgrx --version 0.19.1 cargo pgrx init # bootstraps a private PostgreSQL cluster # Build & install into the dev cluster. git clone https://codeberg.org/gregburd/pg_turbovec cd pg_turbovec cargo pgrx install --release # default features include the index AM # Or build a stripped-down variant without the index AM: cargo pgrx install --release --no-default-features --features pg16 ``` ### Nix flake The repo ships a flake with one package per supported PostgreSQL major (`pg_turbovec_13` … `pg_turbovec_19`; `default` = PG18): ```bash # Build the PG18 extension: nix build github:gburd/pg_turbovec#pg_turbovec_18 # result/lib/pg_turbovec.so # result/share/postgresql/extension/pg_turbovec--2.10.2.sql + .control ``` On PostgreSQL 18+ you can point a stock server at the store path without copying files (PG18 added `extension_control_path`): ``` extension_control_path = '$system:/nix/store/…-pg_turbovec-X.Y.Z/share/postgresql' dynamic_library_path = '$libdir:/nix/store/…-pg_turbovec-X.Y.Z/lib' ``` For PG13–17, copy (or symlink) `lib/pg_turbovec.so` into `pg_config --pkglibdir` and `share/postgresql/extension/*` into `pg_config --sharedir`/extension, or overlay the extension into nixpkgs' `postgresql_NN.pkgs`. A dev shell with the full toolchain (rust, cargo-pgrx 0.19.1, clang/bindgen, openblas, bison/flex/readline/icu) is available via `nix develop`. For a Nix-based build (the dev environment for this project) see [`docs/BUILDING.md`](docs/BUILDING.md). For migrating from pgvector see [`docs/MIGRATING_FROM_PGVECTOR.md`](docs/MIGRATING_FROM_PGVECTOR.md). ## Getting Started Enable the extension (do this once per database): ```sql CREATE EXTENSION pg_turbovec; SET search_path = public, turbovec; ``` > **Add `pg_turbovec` to `shared_preload_libraries`.** The `turbovec.*` GUCs > (`probes`, `search_k`, `iterative_scan`, `out_of_core`, …) are registered > when the library loads. Without preloading, a session has **none** of them: > `SET turbovec.probes = 16` is silently accepted and does nothing, and every > query runs at the defaults no matter what you tune. > > ```sql > ALTER SYSTEM SET shared_preload_libraries = 'pg_turbovec'; -- then restart > ``` > > Verify (expect a non-zero count, not `0`): > > ```sql > SELECT count(*) FROM pg_settings WHERE name LIKE 'turbovec.%'; > ``` > > Indexes and queries work without preloading — only the tuning knobs > disappear, which makes the failure quiet and easy to misdiagnose. This cost > a full benchmark round on 2026-09-22 before it was spotted. Create a table with a `vector` column: ```sql CREATE TABLE items ( id bigserial PRIMARY KEY, body text, embedding vector ); ``` Insert vectors: ```sql INSERT INTO items (body, embedding) VALUES ('hello', '[0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8]'), ('world', '[0.2, 0.1, 0.4, 0.3, 0.6, 0.5, 0.8, 0.7]'); -- Or via array cast: INSERT INTO items (body, embedding) VALUES ('greeting', ARRAY[0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8]::real[]::vector); -- Or via the pgvector-style function: INSERT INTO items (body, embedding) VALUES ('hi', to_vector('[0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8]', 8, false)); ``` Get the nearest neighbours by cosine distance: ```sql SELECT id, body, embedding <=> '[0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8]'::vector AS dist FROM items ORDER BY embedding <=> '[0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8]'::vector LIMIT 5; ``` Also supports inner product (`<#>`), L2 (`<->`), and L1 (`<+>`). `<#>` returns the *negative* inner product so `ORDER BY ... ASC` returns most-similar-first - same convention as pgvector. ## Storing ```sql -- Variable dimension: CREATE TABLE items (id bigserial PRIMARY KEY, embedding vector); -- Or with a runtime dim assertion via a CHECK constraint -- (typmod-style enforcement without typmod plumbing): CREATE TABLE items ( id bigserial PRIMARY KEY, embedding vector, CHECK (turbovec.vec_check_dim(embedding, 1536)) ); -- Or add to an existing table: ALTER TABLE items ADD COLUMN embedding vector; ``` `vector` accepts dim 1..16000 (matching pgvector's cap). The TurboQuant kernel additionally requires dim be a multiple of 8 when used with the index AM; pad your embeddings if your model emits something awkward (e.g. 384 = 48 × 8 ✓, 1536 = 192 × 8 ✓). Insert vectors in bulk via `COPY`: ```sql COPY items (embedding) FROM STDIN WITH (FORMAT TEXT); [1,2,3,4,5,6,7,8] [0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8] \. ``` Upsert via `ON CONFLICT`: ```sql INSERT INTO items (id, embedding) VALUES (1, '[1,2,3,4,5,6,7,8]') ON CONFLICT (id) DO UPDATE SET embedding = EXCLUDED.embedding; ``` ## Querying ### Distance operators | Op | Meaning | Indexed (turbovec AM)? | |-------|------------------------------------|------------------------------------------| | `<->` | Euclidean (L2) distance | yes - `vec_l2_ops` | | `<#>` | negative inner product | yes - `vec_ip_ops` (default) | | `<=>` | cosine distance (`1 - cos θ`) | yes - `vec_cosine_ops` | | `<+>` | taxicab (L1) distance | yes - `vec_l1_ops` | Distances are returned as `double precision`. Distance accumulators are `f64` internally - `pg_turbovec`'s `avg(vector)` and `sum(vector)` preserve more precision than pgvector's `f32` accumulators on million-row corpora. ### Named functions (pgvector-compatible) ```sql l2_distance(a vector, b vector) RETURNS double precision inner_product(a vector, b vector) RETURNS double precision cosine_distance(a vector, b vector) RETURNS double precision l1_distance(a vector, b vector) RETURNS double precision vector_dims(v vector) RETURNS integer vector_norm(v vector) RETURNS double precision to_vector(text) RETURNS vector to_vector(text, integer, boolean) RETURNS vector -- with dim check array_to_vector(real[]) RETURNS vector array_to_vector(real[], integer, boolean) RETURNS vector -- with dim check subvector(v vector, start integer, length integer) RETURNS vector vec_normalize(vector) RETURNS vector vec_zeros(integer) RETURNS vector vec_check_dim(vector, integer) RETURNS vector -- assertion ``` ### Aggregates ```sql SELECT avg(embedding) FROM items WHERE topic = 'cats'; -- centroid SELECT sum(embedding) FROM items; ``` Both are `PARALLEL SAFE` and use `f64` accumulators. ### JSONB I/O ```sql SELECT '[1, 2.5, -3]'::vector::jsonb; -- → [1, 2.5, -3] SELECT '[1, 2.5, -3]'::jsonb::vector; -- → vector ``` ## Indexing `pg_turbovec` ships an index access method named `turbovec`, included in the default build (the `experimental_index_am` and `relfile_storage` Cargo features were retired in v1.3.0; the relfile-resident storage path is now the only strategy). ```sql -- Cosine-distance ordering (most common for semantic search): CREATE INDEX items_emb_idx ON items USING turbovec (embedding vec_cosine_ops) WITH (bit_width = 4); -- Inner-product ordering: CREATE INDEX items_emb_ip_idx ON items USING turbovec (embedding vec_ip_ops) WITH (bit_width = 2); -- 32× compression vs FP32; some recall loss -- Online (non-blocking) build: CREATE INDEX CONCURRENTLY items_emb_idx ON items USING turbovec (embedding vec_cosine_ops); ``` Reloptions: | Option | Default | Range | Notes | |-------------|----------------|--------|-------| | `bit_width` | `turbovec.bit_width_default` (4) | 1, 2, 3, 4 | Lower = smaller index, lower recall. `2/3/4` = TurboQuant. `1` = centered sign binary quantization with full-precision reranking (shipped in v2.6.0; see [1-bit search](#1-bit-search-with-full-precision-reranking)). | ### Index AM lifecycle The `turbovec` AM supports the full PostgreSQL index lifecycle: - `CREATE INDEX [CONCURRENTLY]` (parallel build is roadmap) - `INSERT` → `aminsert` (idempotent on `IdAlreadyPresent` - handles CIC's two-pass build and HOT updates) - `DELETE` + `VACUUM` → `ambulkdelete` (we track every live `u64` id in a parallel `Vec` so dead rows are removed correctly) - `REINDEX [CONCURRENTLY]` - `DROP INDEX` - Order-by-op scans (`ORDER BY emb <=> q LIMIT k`) via the executor's recheck-orderby path ### Filtered (hybrid) search Three patterns, one decision matrix — full guide in [`docs/FILTERING.md`](docs/FILTERING.md): a **partial index** (`CREATE INDEX ... WHERE tenant_id = X`) for known filter values, the **in-kernel allowlist** `turbovec.knn(..., allowed)` for selective per-query id sets, and **iterative scan** for the normal `ORDER BY ... LIMIT` ergonomics. The allowlist (shown below) pushes the id set into the SIMD scoring loop — a *selective* filter gets cheaper, not more expensive (the kernel short-circuits 32-vector blocks whose allowed-slot mask is empty before any LUT lookup; measured crossover at ~7–10% selectivity). ```sql -- Top-10 nearest, restricted to a tenant or topic: SELECT k.id, d.body FROM turbovec.knn( 'items'::regclass, 'id', 'embedding', '[...]'::vector, 10, 4, ARRAY(SELECT id FROM items WHERE tenant_id = $1)::bigint[] ) k JOIN items d USING (id) ORDER BY k.score DESC; ``` The function-driven `turbovec.knn(...)` API and the `turbovec` index AM share the same TurboQuant kernel and the same backend-local cache; pick whichever fits your query shape (the AM integrates with `ORDER BY ... LIMIT`; `knn(...)` lets you pass an allowlist for hybrid retrieval). ## Performance > **Operations note: `shared_buffers`.** As of v1.19.0, every byte of > a pg_turbovec index is read through PostgreSQL's buffer manager > (`ReadBufferExtended`) — there is no relfile mmap. `shared_buffers` > sizing matters: size it to hold the hot (compressed) index for best > cold-fill latency. pg_turbovec's 7–15× compression vs fp32 HNSW is > what makes "the index fits `shared_buffers`" achievable at corpus > sizes where an uncompressed index could not. A cold-fill on an > index larger than `shared_buffers` pays the buffer-manager's > per-page pin/lock/copy cost the first time each backend touches it; > warm (already-cached) scans are unaffected by index size. > > Out-of-core (>RAM) serving for IVF indexes doesn't need > `shared_buffers` to hold the whole index either — `turbovec. > out_of_core` (default `auto`) keeps the per-backend resident set at > O(probes·cell_size) by gathering only the probed cells' pages > through the buffer manager. > > Architecture: [`docs/ARCHITECTURE.md` § 8.1 "Index AM · > buffer-cache-only reads"](docs/ARCHITECTURE.md#81-index-am--buffer-cache-only-reads). > Design rationale for the (now-shipped) buffer-cache-only read path: > [`docs/BUFFER_CACHE_ONLY_DESIGN.md`](docs/BUFFER_CACHE_ONLY_DESIGN.md). > Historical mmap-era benchmarks (v1.5.0–v1.18.x, since removed): > [`docs/RECALL.md § 2.5–2.6`](docs/RECALL.md). > **Performance methodology.** The headline numbers in the > table at the top of this README come from a real > head-to-head against pgvector 0.8.0 HNSW on the > [`dbpedia-entities-openai-1M`](https://huggingface.co/datasets/KShivendu/dbpedia-entities-openai-1M) > corpus (1 M Wikipedia/DBpedia entities × 1536-d OpenAI > `text-embedding-ada-002` embeddings) running on a single Intel > i9-12900H box with 32 GiB RAM and PG 17.9 from the pgrx-managed > install tree. Full methodology, query set, ground-truth > generation, and reproduction scripts in > [`docs/RECALL.md § 2.2`](docs/RECALL.md) and > [`benches/scripts/`](benches/scripts/). The synthetic-uniform > tables below are from a pure-Rust kernel bench and are kept for > historical comparison - they understate real-world recall because > uniform-random vectors have no clustering structure for > quantization to exploit. ### Recall (synthetic, 1 000 random unit-norm vectors, 50 queries) | dim | bit_width | R@1 | R@10 | R@100 | |----:|----------:|-----:|-----:|------:| | 128 | 2 | 0.40 | 0.65 | 0.76 | | 128 | 4 | 0.80 | 0.89 | 0.93 | | 384 | 2 | 0.34 | 0.62 | 0.76 | | 384 | 4 | 0.78 | 0.89 | 0.93 | | 768 | 2 | 0.50 | 0.62 | 0.76 | | 768 | 4 | 0.82 | 0.88 | 0.92 | Random vectors have no clustering structure for the quantiser to exploit - real embeddings (GloVe, OpenAI ada-002) recall meaningfully better. Reproduction: ```bash cargo bench --bench recall --no-default-features --features pg16 ``` Real-world fixtures via the `TURBOVEC_FIXTURE_PATH` env var; format documented in [`docs/RECALL.md`](docs/RECALL.md) § 6.1. ### Compression (from the TurboQuant paper) | dim | FP32 / vector | TurboQuant 4-bit / vector | TurboQuant 2-bit / vector | |-------|-----------:|-----------------------:|-----------------------:| | 128 | 512 B | 68 B | 36 B | | 384 | 1 536 B | 196 B | 100 B | | 768 | 3 072 B | 388 B | 196 B | | 1536 | 6 144 B | 772 B | 388 B | | 3072 | 12 288 B | 1 540 B | 772 B | A 10 M-row × 1536-dim corpus that needs ~62 GiB of RAM as FP32 fits in ~7.7 GiB at 4-bit and ~3.9 GiB at 2-bit - without any data-dependent codebook training. ### Search speed (from the TurboQuant paper, x86 AVX-512BW) 100 K vectors, 1 K queries, k=64, single-threaded: - TurboQuant **matches or beats** FAISS `IndexPQFastScan` at every 4-bit configuration tested (d=384, 768, 1536, 3072). - TurboQuant runs **within ±1%** of FAISS at 2-bit single-threaded. - On ARM (Apple M3 Max), TurboQuant **beats** FAISS by 12-20% at every config the paper measured. We have not yet run pg_turbovec end-to-end against pgvector + HNSW or pgvectorscale + StreamingDiskANN. That comparison is the next item on the v1.0.0 roadmap; see [`docs/RECALL.md`](docs/RECALL.md). ### How it works TurboQuant compresses each vector to 2/3/4 bits per coordinate using: 1. **Normalize** - strip the L2 norm; store as a single `f32` scale. 2. **Random rotation** - multiply by a fixed orthogonal matrix so each coordinate independently follows a known Beta distribution. 3. **Lloyd-Max scalar quantisation** - bucket each coordinate into 2/3/4-bit codes optimal for the known distribution. 4. **Bit-pack** - `dim` coordinates → `dim * bit_width / 8` bytes. 5. **Length-renormalised scoring** - one extra scalar per vector removes the inner-product downward bias the quantiser introduces. No codebook training, no data passes - adding vectors is `O(dim)` per vector with no rebuild as the corpus grows. Search rotates the query once and scores directly against the bit-packed codes via SIMD nibble-LUT kernels (NEON, AVX2, AVX-512BW). ## Configuration `pg_turbovec` exposes 20 GUCs under the `turbovec.*` namespace (all USERSET — settable per session). The full reference with tuning guidance is in [docs/PRODUCTION.md](docs/PRODUCTION.md); the most commonly-tuned ones: | GUC | Type | Default | Range | |----------------------------------|------|---------|----------------| | `turbovec.bit_width_default` | int | `4` | `2..=4` | | `turbovec.probes` | int | `16` | `1..=65536` | | `turbovec.search_k` | int | `32` | `1..=100000` | | `turbovec.oversample` | float| `1.0` | `1.0..=100.0` | | `turbovec.hi_dim_rerank` | enum | `auto` | off, auto, on | | `turbovec.iterative_scan` | enum | `off` | off, relaxed_order | | `turbovec.max_probes` | int | `64` | `1..=65536` | | `turbovec.max_scan_tuples` | int | (see docs) | | | `turbovec.out_of_core` | enum | `auto` | off, auto, on | | `turbovec.coarse_graph` | enum | `auto` | off, auto, on | | `turbovec.graph_ef` | int | `0` | 0 auto (=512) / `1..=1000000` | | `turbovec.graph_build_partitions`| int | `-1` | -1 auto / 0-1 single-pass / N | | `turbovec.build_parallelism` | int | `0` | `0..=128` | | `turbovec.scan_parallelism` | int | `0` | `0..=128` | | `turbovec.cache_size_mb` | int | `256` | `0..=65536` | | `turbovec.normalize_on_insert` | bool | `true` | - | | `turbovec.warn_on_rebuild` | bool | `true` | - | | `turbovec.allowlist` | str | `""` | CSV of heap-TID bigints | ```sql -- Compress harder during this session: SET turbovec.bit_width_default = 2; -- Disable the backend-local index cache: SET turbovec.cache_size_mb = 0; ``` ## Migrating from pgvector `pg_turbovec` and `pgvector` coexist cleanly - different schema, type name, and operator-dispatch table. See [`docs/MIGRATING_FROM_PGVECTOR.md`](docs/MIGRATING_FROM_PGVECTOR.md) for the full cookbook. TL;DR: ```sql ALTER TABLE docs ADD COLUMN embedding_tv turbovec.vector; UPDATE docs SET embedding_tv = embedding::real[]::turbovec.vector; CREATE INDEX CONCURRENTLY docs_emb_tv_idx ON docs USING turbovec (embedding_tv vec_cosine_ops) WITH (bit_width = 4); ``` A binary-compatible `vector` varlena layout (zero-copy cast to/from pgvector's `vector`) is on the v1.0 roadmap; until then the `real[]` bridge is the supported interop path. ## Reference - **Type:** `vector` (variable dimension, `f32` coordinates, 1..16000) - **Schema:** `turbovec` (set on the search_path or fully qualify) - **Operator classes:** `vec_ip_ops` (default, `<#>`), `vec_cosine_ops` (`<=>`), `vec_l2_ops` (`<->`), `vec_l1_ops` (`<+>`), and `vec_colbert_ops` (multivector MaxSim) - **Index AM:** `turbovec` (build with `WITH (bit_width = 1|2|3|4)`) - **Aggregates:** `avg(vector)`, `sum(vector)` - **Full surface listing:** [`docs/USAGE.md`](docs/USAGE.md) and the generated `sql/pg_turbovec--.sql` after `cargo pgrx schema`. ## Documentation - [`docs/ARCHITECTURE.md`](docs/ARCHITECTURE.md) - module map, type / operator / aggregate signatures, index AM contract, GUC semantics, phased roadmap. - [`docs/USAGE.md`](docs/USAGE.md) - cookbook covering install, exact + ANN search, aggregates, arithmetic, tuning. - [`docs/PRODUCTION.md`](docs/PRODUCTION.md) - **deployment guide**: install, GUC tuning, replication, monitoring, troubleshooting. - [`docs/MIGRATING_FROM_PGVECTOR.md`](docs/MIGRATING_FROM_PGVECTOR.md) - hands-on migration with query rewrite tables and a feature comparison. - [`docs/FILTERING.md`](docs/FILTERING.md) - **filtering & hybrid search guide**: partial index vs in-kernel allowlist `knn()` vs iterative scan, a decision matrix, and the measured allowlist selectivity crossover. - [`docs/HYBRID_SEARCH.md`](docs/HYBRID_SEARCH.md) - **multivector & hybrid breadth guide**: ColBERT MaxSim re-rank (`max_sim`), dense+sparse reciprocal rank fusion (`rrf_score` + CTE recipe), and the named-vector multi-column pattern. - [`docs/INDEXAM.md`](docs/INDEXAM.md) - implementation guide for the `turbovec` index access method. - [`docs/RECALL.md`](docs/RECALL.md) - recall benchmark methodology and the latest measured numbers. - [`docs/PARITY_GAPS.md`](docs/PARITY_GAPS.md) - feature-by-feature comparison against pgvector. - [`docs/PARTITIONED_SCALE.md`](docs/PARTITIONED_SCALE.md) - scaling past PostgreSQL's single-table ceiling to billions of vectors via hash-partitioning + per-partition indexes + native `Merge Append`. - [`docs/DEPLOYING_ON_MANAGED_POSTGRES.md`](docs/DEPLOYING_ON_MANAGED_POSTGRES.md) - durability/replication, restricted-superuser compatibility, read replicas, cancellation, and the memory / insert-cost model for managed or hosted PostgreSQL. - [`docs/PG_VERSION_SUPPORT.md`](docs/PG_VERSION_SUPPORT.md) - the per-major support matrix (PG 13-18, 19beta1 experimental). - [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md) - the published head-to-head benchmark (Cohere-wiki 1M vs pgvector HNSW), incl. the AVX2 latency frontier and the honest "flat-scan loses on latency, wins on storage + exact recall" finding. - [`docs/BUILDING.md`](docs/BUILDING.md) - Nix-specific build recipe (writable `pg_config` wrapper, `BINDGEN_EXTRA_CLANG_ARGS`, `RUSTFLAGS` for openblas). - [`RELEASING.md`](RELEASING.md) - release process, version-bump checklist, Codeberg release flow. - [`CHANGELOG.md`](CHANGELOG.md) - phase-by-phase release notes. - [`tests/`](tests/) - psql regression scripts you can run yourself (`cargo pgrx run pg16`, then `\i tests/03_full_demo.sql`). ## FAQ **Is `pg_turbovec` a drop-in replacement for `pgvector`?** No, by design. We coexist: type name `vector`, schema `turbovec`, operator dispatch by argument type. Pgvector users have years of `vector(1536)` columns and tooling - pretending to be a drop-in would silently change semantics around normalisation and recall. The [migration cookbook](docs/MIGRATING_FROM_PGVECTOR.md) shows the explicit `real[]` bridge. **What about `halfvec`, `sparsevec`, `bitvec`?** The **types and their distance/arithmetic operators exist** (`halfvec`, `sparsevec`, `bitvec` with `<->`, `<#>`, `<=>`, `<+>`, and Hamming/Jaccard `<~>`/`<%>` for `bitvec`) and coexist with pgvector's. What they are **not** is indexable by the turbovec index AM: the AM quantises full-precision `f32` input, and half-precision, sparse, and bit representations don't map onto the TurboQuant kernel. Use them as column types and with the exact operators; for ANN over those representations, pgvector's HNSW is the right tool. **What about L2 / L1 ANN?** Both are indexable by the turbovec AM — `vec_l2_ops` (`<->`) and `vec_l1_ops` (`<+>`) drive the kernel and rerank exactly, the same as `vec_ip_ops` / `vec_cosine_ops`. `l2_distance` / `l1_distance` are also available as exact functions. **What's not in 1.0?** The two items most likely to be asked about: - **Binary-compatible varlena layout for `vector`.** We use a CBOR-derived varlena rather than pgvector's `[vl_len_, dim, unused, f32[dim]]` byte layout. The cross-extension migration via `::real[]::vector` is one-shot and finishes in seconds on a million rows; the 16× quantization savings dominate the per-row layout overhead, so binary-compat is a nice-to-have, not a 1.0 blocker. - **Indexed Hamming / Jaccard ANN on `bitvec`.** The TurboQuant kernel is a scalar quantizer for dense `f32` vectors; it doesn't fit Hamming-space ANN. And the workload that motivates `bit_hamming_ops` (memory-pressured semantic search) is already covered better by `bit_width = 2` - same byte budget, materially higher recall. **Why two crates: `pg_turbovec` and `turbovec`?** [`turbovec`](https://crates.io/crates/turbovec) is the upstream TurboQuant implementation in pure Rust by Ryan Codrai. `pg_turbovec` is the PostgreSQL extension built with [pgrx](https://github.com/pgcentralfoundation/pgrx) on top of it. pg_turbovec currently pins a small [fork](https://github.com/gburd/turbovec) at the upstream **0.9.0** line plus an additive integration layer (a borrowed/skip-prepare reconstruction path used to rebuild a search-ready index from PostgreSQL's own relation pages without a turbovec file). Upstream has since released **1.0.0**, which stabilises the on-disk format and reorganises the encode/rotation API (the v5 block-Hadamard rotation changed every encoded byte); adopting it is a planned, wire-format-aware major migration on the pg_turbovec side rather than a routine bump. The additive pieces we carry are tracked for upstreaming (turbovec issue #70's `from_parts` + in-memory I/O asks have already landed upstream). ## Ecosystem Vector search on PostgreSQL has three serious open-source options. The honest comparison: - **[pgvector](https://github.com/pgvector/pgvector)** - production- tested at scale, larger feature surface (HNSW for L2 *and* inner product *and* L1, plus `halfvec`, `sparsevec`, `bitvec`), and an ecosystem of clients that already speak its types and operators. Choose pgvector if you don't have memory pressure and you value maturity. - **[pgvectorscale](https://github.com/timescale/pgvectorscale)** - SOTA published latency on 50 M+ row corpora via StreamingDiskANN, layered on top of pgvector. Choose pgvectorscale if your corpus is in the tens-of-millions-of-rows range and you can run TimescaleDB. - **pg_turbovec** - smallest on-disk footprint, in-kernel filtered ANN (selective `WHERE` clauses make scans *cheaper*, not more expensive), zero codebook training. Choose pg_turbovec if memory dominates your cost equation. The three coexist cleanly in the same database - separate schemas, separate type oids, separate operator dispatch. You can A/B them on your own data without committing to one. ## Contributing Issues and patches: . ```bash # Run the full test suite (boots a private PG cluster): cargo pgrx test pg16 # Run pure-Rust kernel + recall benches (no Postgres): cargo bench --bench distance --no-default-features --features pg16 cargo bench --bench recall --no-default-features --features pg16 # Lints + format: cargo clippy --features pg16 --tests -- -D warnings cargo fmt --all -- --check ``` See [`CONTRIBUTING.md`](CONTRIBUTING.md). ## Acknowledgements - **Ryan Codrai** for the [`turbovec`](https://github.com/RyanCodrai/turbovec) Rust crate and the SIMD kernels that do all the actual work. - **Google Research** for [TurboQuant](https://arxiv.org/abs/2504.19874) (ICLR 2026) - the algorithm. - **The pgvector authors** for setting the API conventions (`<-> <#> <=> <+>`, `to_vector`, `array_to_vector`, `subvector`, `vector_dims`, `vector_norm`) we mirror. - **The pgrx maintainers** for making PostgreSQL extension development in Rust possible. ## License Apache-2.0 © Greg Burd. See [`LICENSE`](LICENSE).