Your pgvector Search Gets Slower as You Add Data. Here’s the Setting Everyone Misses.

Your semantic search was instant at ten thousand rows. At two million it’s 800 milliseconds and climbing, and you never touched the query. Almost always the cause is the same: pgvector is doing an exact, brute-force scan of every vector because its index was never built — or was built with defaults that don’t fit your data.

Exact search doesn’t scale, and it’s the default

Without an approximate index, pgvector compares your query vector against every row. That’s fine at ten thousand and fatal at two million. The fix is an ANN index — but an ANN index has two knobs that decide everything, and both have quietly wrong defaults for a large table.

-- lists ≈ rows / 1000 up to ~1M rows, then ≈ sqrt(rows).
-- 2,000,000 rows -> ~1,414 lists, NOT the tiny number you'll get by guessing.
CREATE INDEX ON docs
  USING ivfflat (embedding vector_cosine_ops)
  WITH (lists = 1414);

-- probes trades recall for speed AT QUERY TIME. The default of 1 is far too low.
-- A sane starting point is ~sqrt(lists); here sqrt(1414) ≈ 38. Then tune to a recall target.
SET ivfflat.probes = 38;

Two failure modes, opposite symptoms. Too few lists and each partition is huge, so every probe scans a lot — slow. Too few probes and you scan too few partitions — fast, but you silently miss relevant results, which in a RAG system means confidently answering from the wrong chunks. You cannot tune one without measuring the other.

Build the index after the data is loaded

IVFFlat clusters your existing vectors to define its lists. Build it on an empty or tiny table and the clusters are meaningless; every later insert lands in an ill-fitting partition and recall degrades. Load first, then index — and rebuild after any large ingest.

-- HNSW: slower to build, larger on disk, but better recall/latency
-- and no clusters to go stale as data changes.
CREATE INDEX ON docs USING hnsw (embedding vector_cosine_ops)
  WITH (m = 16, ef_construction = 64);
SET hnsw.ef_search = 40;   -- the recall/speed dial at query time

If your data grows or churns continuously, HNSW usually ages better than IVFFlat because it has no clusters to go stale. It costs more to build and store; that’s the trade you’re making.

The lesson

Measure recall, not just latency. A vector search that got ten times faster by missing a third of the right answers isn’t faster — it’s broken with a good p99. Pick a fixed set of queries with known-good results, and every time you touch lists, probes, or the index type, confirm recall held before you celebrate the speed.

The benchmark harness and recall test are on GitHub: github.com/waghmaredb/vexpose-labs. Tuning vector search at scale? Compare notes on LinkedIn or X.

Comments

Leave a Reply

Discover more from {{ vExpose }}.Blog

Subscribe now to keep reading and get access to the full archive.

Continue reading