Deep Dive · Jul 14, 2026 · 7 min read
Production Document Q&A: Chunking, Hybrid Search, Citations, and Evals
The engineering choices that decide whether a document assistant is trusted: how you split, search, cite, and test.
Document Q&A fails quietly: the answer looks fine and is wrong. Trust comes from retrieval quality, enforced citations, and a test set. This post walks through each, with code.
Chunking
Split on document structure first (headings, clauses, table rows), then by size. Keep a small overlap and carry the heading path into each chunk so it makes sense on its own.
def chunk_section(heading_path, text, max_tokens=350, overlap=40):
words = text.split()
chunks, start = [], 0
while start < len(words):
end = min(start + max_tokens, len(words))
body = " ".join(words[start:end])
chunks.append({"heading": " > ".join(heading_path), "text": f"{' > '.join(heading_path)}\n{body}"})
if end == len(words):
break
start = end - overlap
return chunksToken counts here are approximated with words. Use your embedding model's tokenizer in production, and treat tables and scanned pages as separate cases that need their own extraction.
Hybrid search
Vector search finds meaning; keyword search finds exact terms such as part numbers and clause ids. Run both and merge the ranked lists with reciprocal rank fusion.
def rrf(rank_lists, k=60):
"""Reciprocal rank fusion over lists of chunk ids, best first."""
scores = {}
for ranking in rank_lists:
for rank, chunk_id in enumerate(ranking, start=1):
scores[chunk_id] = scores.get(chunk_id, 0.0) + 1.0 / (k + rank)
return sorted(scores, key=scores.get, reverse=True)
def search(query, user, index, top_k=8):
allowed = index.allowed_doc_ids(user) # permissions before ranking
vec = index.vector_search(query, filter_ids=allowed, limit=30)
kw = index.keyword_search(query, filter_ids=allowed, limit=30)
return rrf([vec, kw])[:top_k]Apply permission filters before ranking, not after. Filtering afterward leaks existence through gaps in results and wastes top-k slots.
Enforced citations
Ask the model for structured output: an answer and the chunk ids supporting each claim. Then verify in code that every cited id was actually retrieved, and optionally that the quoted text appears in the chunk.
def verify(answer_json, retrieved):
by_id = {c.id: c.text for c in retrieved}
for claim in answer_json["claims"]:
for cid in claim["chunk_ids"]:
if cid not in by_id:
return False, f"cites unknown chunk {cid}"
if claim.get("quote") and not any(claim["quote"] in by_id[c] for c in claim["chunk_ids"]):
return False, "quote not found in cited chunk"
return True, "ok"If verification fails, retry once with the error message, then fall back to "I could not find a reliable answer" and show the top passages.
Evaluation
- Collect 50 to 100 real questions with the correct answer and source passage.
- Measure retrieval recall@8: was the gold passage in the top eight?
- Measure answer correctness separately, with human review at first.
- Track abstention: it should decline when the answer is not in the documents.
- Re-run on every change to chunking, embeddings, or prompts.
Operational notes
- Re-index on document change with a content hash, so unchanged files are skipped.
- Store the document version with each chunk so answers can cite the version used.
- Log queries with no good results; they are your content gap list.
Choosing an embedding model
Embedding quality bounds retrieval quality. Compare two or three candidate models on your own questions: embed your chunks, run your golden questions, and measure recall at eight. Consider the dimension and its storage cost, the maximum input length, multilingual support, and whether you can run the model privately. Re-run the comparison when you change chunking, since the two interact.
def recall_at_k(search, golden, k=8):
hits = 0
for q in golden:
ids = search(q["question"], top_k=k)
hits += any(g in ids for g in q["gold_chunk_ids"])
return hits / len(golden)Reranking
Hybrid search gets the right passage into the top thirty. A reranker, which scores each query and passage pair jointly, then picks the best eight. It is slower per item than vector search, so apply it only to the shortlist. Measure whether it improves recall at eight on your data before paying the extra latency.
Tables, forms, and scans
- Tables: keep rows with their headers, and serialize as text with column names, so a row makes sense on its own.
- Scanned pages: run OCR, record the confidence, and flag low-confidence pages for review instead of indexing garbage.
- Forms: extract fields as key-value pairs and index both the pairs and the raw text.
- Slides and diagrams: extract text and notes, and describe images only where they carry information you need.
Freshness and versioning
People ask about the current policy, not last year's. Store an effective date and a version with each chunk, prefer the latest version by default, and when the answer depends on a version, say which one was used. Keep old versions searchable for audit, but require the user to ask for them.
Latency and cost
A typical request embeds the question, runs two searches, reranks, and calls a model with several thousand tokens of context. Cache embeddings for repeated questions, trim the context to the best passages, and stream the answer so that people see progress. Track tokens per answer and set an alert when the average rises.
Failure analysis routine
Each week, take the thumbs-down answers and classify the cause: the answer was not in the corpus, retrieval missed it, the chunk lacked context, the model misread it, or the document was outdated. The distribution tells you where to invest. In most deployments, content and retrieval fixes beat prompt changes.
- Not in corpus: add or correct the document.
- Retrieval miss: adjust chunking, add keyword boosts, or tune the reranker.
- Model error: tighten the instructions and the verification step.
- Stale content: fix the source and add an expiry date.
How we can help
We build document assistants and audit existing ones. An audit covers chunking, retrieval recall, citation enforcement, and permissions, and ends with a prioritized fix list. Contact us to scope one.
Related reading
Need help implementing this?
Our consultants run architecture reviews and build production pilots. Book a free scoping call to talk through your design.
Book a Free Scoping Callor email us at hello@deepvero.com