By the end of this post you’ll be able to take a set of text documents, embed them, load them into a vector database, run a similarity query against that store, and understand which knobs matter when results come back wrong. The goal is not a survey of every product on the market — it’s a working mental model plus a concrete implementation path you can adapt to your own project.

A vector database stores high-dimensional numeric vectors and answers the question “which stored vectors are closest to this query vector?” at scale. That is the whole product category in one sentence. Everything else — chunking strategy, distance metrics, index type, filtering — is detail on top of that core operation.

This post is organized as a troubleshooting checklist. Each section starts with a symptom you might see when working with embeddings and vector search, explains the likely cause, and gives a fix. If your system already works and you just want to tune it, skip to the section that matches your current problem.


Symptom: You stored vectors and every query returns the same document

You built a small pipeline, inserted a few hundred documents, ran three different queries, and the top result is identical every time. Or the top results look random rather than ranked by relevance.

Cause: The most common explanation is a vector normalization mismatch combined with a distance metric that assumes normalized inputs. Many embedding models output unit-length vectors when configured a certain way, but the SDK you used to compute embeddings may not normalize by default. If your collection was created with cosine similarity but the vectors were inserted unnormalized, the ranking math does the wrong thing.

A second common cause is a mismatched embedding model. If the query was embedded with text-embedding-3-small but the stored documents were embedded with a different model, the vector spaces are unrelated. Cosine similarity between two vectors from different models is meaningless — you’re comparing coordinates in two different coordinate systems.

Fix: Embed a test string, and a query string that should match it, using the same function. Compute similarity yourself and confirm it exceeds a known threshold. If it doesn’t, the model or normalization is the problem, not the database.

import numpy as np
from openai import OpenAI

client = OpenAI()

def embed(text: str, model: str = "text-embedding-3-small") -> np.ndarray:
    resp = client.embeddings.create(model=model, input=text)
    v = np.array(resp.data[0].embedding, dtype=np.float32)
    return v / np.linalg.norm(v)  # normalize to unit length

doc = "Refunds are processed within five business days."
query = "How long does a refund take?"

d, q = embed(doc), embed(query)
print("cosine similarity:", float(np.dot(d, q)))

A score above roughly 0.7 for a matching pair is common with this model family. If a clearly related pair scores near zero, your embedding function or model identifier is wrong. Confirm the exact model string your SDK uses, and confirm the same string is stored per collection so you can detect drift later.


Symptom: The query is fast but the results are only approximately right

You’re getting reasonable results, but occasionally a clearly more relevant document ranks below a less relevant one. Latency is fine; accuracy is “close enough” until a user notices.

Cause: Approximate nearest neighbor (ANN) indexing. Almost every production vector database does not compute exact distance against every stored vector. It builds an index — HNSW, IVF, or a product quantizer variant — that trades a small amount of recall for a large gain in query speed. The tradeoff is intentional and usually worth it, but the default parameters are tuned for a generic workload, not yours.

Fix: Measure recall against a brute-force baseline before you accept approximate results. Brute force over 10,000 vectors is trivially fast and gives you the ground truth ranking. Compare your ANN results to it, then tune two or three parameters and re-measure.

# Ground truth: exact top-k by inner product (vectors already normalized)
exact_scores = corpus_matrix @ query_vector          # shape (N,)
exact_top_k = np.argsort(-exact_scores)[:10]

# ANN result, e.g. from an HNSW index query
ann_ids = [hit.id for hit in index.query(vector=query_vector.tolist(), top_k=10)]
ann_top_k = np.array(ann_ids)

# Recall@10 against the exact baseline
recall = len(set(exact_top_k.tolist()) & set(ann_top_k.tolist())) / 10
print(f"recall@10 = {recall:.2f}")

A recall@10 below about 0.8 for a small corpus usually means the ANN parameters are too aggressive. In HNSW, raising ef_search (sometimes exposed as ef or efConstruction depending on the library) increases recall at the cost of latency. In IVF-based indexes, raising nprobe does the same. Start by increasing the search-time parameter, not the build-time one — rebuilds are expensive, per-query tuning is cheap.

When NOT to use ANN: if your collection is under roughly 50,000 vectors and latency budget is generous, exact search is often simpler and removes an entire class of recall bugs. Databases like pgvector support both modes on the same table; use exact during development and switch to an index once you have a reason to.


Search for “how do I reset my password” and you get every document that contains the literal phrase “reset my password.” Search for “I forgot my login credentials” — same intent — and you get nothing useful.

Cause: The embedding model is weak for your domain, or your documents are too long to embed into a single meaningful vector. Embedding models compress a passage into one fixed-length vector. A 3,000-word document embedded as a single vector has to represent everything in it, and the resulting vector tends to sit near the center of the semantic space — close to nothing in particular.

Fix: Chunk first, embed second. Split long documents into overlapping windows sized to the model’s effective context — typically 200 to 500 tokens per chunk with 10 to 20 percent overlap so a sentence that straddles a boundary appears in both neighbors. Store the parent document ID alongside each chunk so you can deduplicate results at retrieval time.

from langchain_text_splitters import RecursiveCharacterTextSplitter

splitter = RecursiveCharacterTextSplitter(
    chunk_size=400,        # characters or tokens depending on your tokenizer
    chunk_overlap=60,      # preserves sentences that straddle boundaries
    separators=["\n\n", "\n", ". ", " ", ""],
)

chunks = splitter.split_text(long_document)
for i, chunk in enumerate(chunks):
    vector = embed(chunk)
    collection.insert({
        "id": f"{doc_id}::{i}",
        "vector": vector.tolist(),
        "parent_id": doc_id,
        "text": chunk,
    })

The separator list matters more than people expect. Putting "\n\n" first means paragraph boundaries are respected before sentence boundaries are. If your source is structured prose, this single change often improves ranking more than swapping embedding models.

Do not reuse one chunk size across every content type. API documentation, chat logs, and legal contracts have very different natural units. Chunk at the level where a unit is self-contained — a function signature plus its description is one chunk; a paragraph is one chunk; a whole policy document is not.


Symptom: Results are semantically right but violate a hard filter (tenant, date, permission)

You want the ten most relevant chunks belonging to tenant A. You get a mix of tenants A and B.

Cause: Two very different failure modes live under this symptom, and they require different fixes.

The first is a post-filtering problem: you query the vector index for the top 100 results, then discard those that fail the tenant filter. If tenant A is a small fraction of the collection, most of your 100 candidates belong to other tenants and you end up with five usable results.

The second is embedding drift by metadata: the query is embedded as a single string, but the tenant context was never included in the embedded text, so the model can’t distinguish otherwise-identical chunks across tenants.

Fix: Push the filter into the index query, and include the distinguishing context in the text you embed.

query_vector = embed("how do I rotate an API key")

results = collection.query(
    vector=query_vector.tolist(),
    top_k=10,
    filter={"tenant_id": {"$eq": "tenant_abc"}},   # pre-filter, not post
    include=["text", "parent_id", "distance"],
)

Most vector databases support a metadata filter parameter that is applied before or during the ANN traversal — some engines implement this efficiently with filtered HNSW variants, others resort to scanning and filtering. Check which mode your database uses for the filter shape you’re passing. A $eq on a single low-cardinality field behaves very differently from a range filter on a timestamp column.

For drift, embed a prefixed string: f"[tenant={tenant_id}] {chunk_text}". This is a small change that costs nothing at query time and makes tenant-distinct chunks separate in the vector space rather than overlapping.

When NOT to use metadata filtering for access control: if the security model is strict, do not rely on the vector index alone. Filtering at the index level is a performance optimization, not an authorization boundary. Enforce permissions at the data-access layer that returns results to the user, so a database misconfiguration doesn’t become a data leak.


Symptom: Insertions are slow or the index keeps rebuilding

You insert a batch of 10,000 chunks and the operation stalls, or your query latency spikes during ingestion.

Cause: HNSW builds the graph incrementally but degrades when insertions arrive one at a time. Each insert triggers a graph update that may rewrite portions of the structure. Batching helps enormously; more importantly, most engines allow you to insert without index maintenance and build the index once at the end.

Fix: Separate ingestion from indexing when the workload is bulk. Different products expose this differently — some have an explicit “create index” call, others defer index construction automatically when a collection is empty.

# pgvector example: load rows first, build the index once
psql "$DATABASE_URL" -c "
  CREATE TABLE docs (id bigserial primary key, parent_id text, content text, embedding vector(1536));
"

# Bulk load with COPY for throughput
psql "$DATABASE_URL" -c "\copy docs (parent_id, content, embedding) FROM 'embeddings.csv' WITH (FORMAT csv)"

# Build the index after the table is populated
psql "$DATABASE_URL" -c "
  CREATE INDEX ON docs
  USING hnsw (embedding vector_cosine_ops)
  WITH (m = 16, ef_construction = 64);
"

The m parameter controls how many neighbors each node in the HNSW graph keeps; higher values improve recall and increase memory. ef_construction controls the search depth during graph construction; higher values improve index quality and slow the build. Defaults around m=16, ef_construction=64 are a reasonable starting point; tune them only after you have measured recall against a baseline.

The dimension of the vector column must match your embedding model exactly. text-embedding-3-small outputs 1536 dimensions; text-embedding-3-large outputs 3072. A mismatch will either fail at insert or silently truncate depending on the driver, and silent truncation is the worse outcome because it produces plausible-looking garbage rankings.


Symptom: The system works on a demo dataset but degrades on real traffic

Your prototype with 500 chunks returned good results. In production with 500,000 chunks, the same queries return worse, slower answers.

Cause: Two things change at scale: the number of near-duplicate vectors increases, and the fraction of relevant results in the top-k shrinks. In a small collection, the k=10 results are almost all relevant because the collection itself is small. In a large one, ten results may include eight near-misses because there are simply more documents that are approximately on-topic.

Fix: Retrieval quality at scale depends more on ranking than on vector search alone. Introduce a reranking step: retrieve a larger candidate set (say 50) with the vector index, then rescore those candidates with a cross-encoder or a small reranking model that reads the query and each candidate together. The reranker is slower per item but runs on only 50 items, so total latency stays manageable.

Budget for this step from the start. A pipeline that goes embedding → vector top-k → LLM context works for demos but rarely survives contact with a real corpus. The standard production pattern is embedding → vector top-k (k larger than needed) → rerank → top-n → LLM. Each stage has a clear job and a clear failure mode.


A Checklist for Debugging a Vector Search Pipeline

SymptomLikely CauseFix
Same top result for every queryModel mismatch or missing normalizationVerify embed function consistency; unit-normalize
Fast but slightly wrong rankingsAggressive ANN parametersCompare to exact baseline; raise ef_search or nprobe
Semantically related docs missingChunks too large, or weak modelChunk to 200–500 tokens with overlap; try a stronger model
Filters ignored or leakingPost-filtering or missing context in embedded textPre-filter in the query; prefix metadata into the text
Slow ingestion, latency spikesRow-by-row inserts with live index updatesBatch inserts; build index after bulk load
Good demo, bad productionNo reranking stageAdd cross-encoder rerank over a larger candidate set

The single most useful habit when building with vector databases is to keep an exact-search baseline around for as long as possible. It costs almost nothing at small scale, it gives you a ground truth to compare against, and it turns vague “the results seem off” complaints into measurable recall numbers you can iterate on. Approximate search is a performance tool, and performance tools deserve a reference point.

The second habit is to treat embedding model choice as a system decision, not a library decision. Changing the model means re-embedding everything and invalidating your index. Version the model identifier in your schema, store it alongside the vectors, and treat a model upgrade as a migration with a backfill and a rollback plan — not a config change.