--- title: Indexing Vectors description: Vectors live alongside text and filters in the ParadeDB index canonical: https://www.paradedb.com/docs/reference/indexing/indexing-vectors --- This is a beta feature available in versions `0.25.0` and above. See [How Vector Search Works](/concepts/vector/overview) for background. Make sure the [pgvector](https://github.com/pgvector/pgvector) extension is installed first. ParadeDB uses pgvector's vector types, but not its HNSW or IVF indexes. The ParadeDB index can index pgvector's `vector` type alongside your text and other columns. This lets you combine vector search with full text search and filters in a single index, which can significantly improve latency/recall for selective queries. ## Creating the Index The `mock_items` table comes with an `embedding` column of type `vector(8)` populated with sample embeddings. In this example, `embedding` is added to the ParadeDB index with cosine similarity as the distance function. ```sql SQL CREATE INDEX search_idx ON mock_items USING paradedb (id, description, category, embedding vector_cosine_ops) WITH (key_field='id'); ``` ```ts Drizzle import { indexing } from "@paradedb/drizzle-paradedb"; // In the pgTable definition: (table) => [ indexing .paradedbIndex("search_idx") .on( table.id, table.description, table.category, indexing.vectorField(table.embedding, "cosine"), ), ]; ``` ```python Django from django.db import connection from paradedb.indexes import ParadeDBIndex with connection.schema_editor() as schema_editor: schema_editor.add_index( MockItem, ParadeDBIndex( fields={ "id": {}, "description": {}, "category": {}, "embedding": {"metric": "cosine"}, }, key_field="id", name="search_idx", ), ) ``` ```python SQLAlchemy from sqlalchemy import Index from paradedb.sqlalchemy import indexing idx = Index( "search_idx", indexing.ParadeDBField(MockItem.id), indexing.ParadeDBField(MockItem.description), indexing.ParadeDBField(MockItem.category), indexing.VectorField(MockItem.embedding, metric="cosine"), postgresql_using="paradedb", postgresql_with={"key_field": "id"}, ) with engine.begin() as conn: idx.create(conn) ``` ```ruby Rails ActiveRecord::Base.connection.add_paradedb_index( :mock_items, fields: { id: {}, description: {}, category: {}, embedding: { metric: :cosine } }, key_field: :id, name: :search_idx ) ``` ```cs EF Core modelBuilder.Entity() .HasParadeDbIndex("search_idx", e => e.Id) .HasField(e => e.Description) .HasField(e => e.Category) .HasField(e => e.Embedding, VectorMetric.Cosine); ``` The desired distance function is encoded into the index definition, and cannot be changed without reindexing. ```sql SQL embedding vector_l2_ops -- L2 (default) embedding vector_cosine_ops -- cosine embedding vector_ip_ops -- inner product ``` ```ts Drizzle indexing.vectorField(mockItems.embedding, "l2"); // L2 (default) indexing.vectorField(mockItems.embedding, "cosine"); // cosine indexing.vectorField(mockItems.embedding, "ip"); // inner product ``` ```python Django "embedding": {"metric": "l2"} # L2 (default) "embedding": {"metric": "cosine"} # cosine "embedding": {"metric": "ip"} # inner product ``` ```python SQLAlchemy indexing.VectorField(MockItem.embedding, metric="l2") # L2 (default) indexing.VectorField(MockItem.embedding, metric="cosine") # cosine indexing.VectorField(MockItem.embedding, metric="ip") # inner product ``` ```ruby Rails embedding: { metric: :l2 } # L2 (default) embedding: { metric: :cosine } # cosine embedding: { metric: :ip } # inner product ``` ```cs EF Core .HasField(e => e.Embedding, VectorMetric.L2) // L2 (default) .HasField(e => e.Embedding, VectorMetric.Cosine) // cosine .HasField(e => e.Embedding, VectorMetric.InnerProduct) // inner product ``` Only pgvector's `vector` type is supported. The `halfvec`, `sparsevec`, and `bit` types are not yet indexable. If you track index build progress with [`pg_stat_progress_create_index`](https://www.postgresql.org/docs/current/progress-reporting.html#CREATE-INDEX-PROGRESS-REPORTING), you may notice progress appear to "stop" at intervals. This is expected: vectors are clustered with k-means at these points, which is computationally expensive. ## Index Options ParadeDB uses a SPANN-style vector index, which is similar to an IVF index but with additional structures to improve recall and latency, especially over large datasets. The following `WITH` options control how vectors are clustered and indexed. All are set at index build time and apply to every vector field in the index. An example: ```sql CREATE INDEX search_idx ON mock_items USING paradedb (id, description, category, embedding vector_cosine_ops) WITH (training_sample_ratio=0.32, max_leaf_size=100); ``` The fraction of indexed vectors sampled to train the clustering at build time. Must be between `0.000001` and `1.0`; `1.0` trains on every vector. The default samples 32% of the vectors. Larger samples can improve centroid quality at the cost of a slower, more memory-intensive build. The maximum number of sampled training vectors in a leaf of the hierarchical clustering tree. Must be an integer between `1` and `2147483647`. Smaller values generally produce more centroids and smaller clusters; the actual centroid count depends on the clustering result. This setting is independent of `training_sample_ratio` and does not cap the number of indexed vectors assigned to a cluster. These options replace `centroid_ratio` and `training_samples_per_centroid`. The old names are no longer accepted. Recreate indexes that explicitly store either old option using the new options; `REINDEX` alone does not remove obsolete index options. The new defaults sample 32% of vectors and use a maximum training leaf size of 100.