Deep Dive · Jun 16, 2026 · 7 min read
Engineering a Support Copilot: Retrieval, Drafting, and Review
A technical walk through ticket intake, hybrid retrieval, grounded drafting, confidence gating, and the feedback loop.
A support copilot looks simple in a demo and gets hard in production. This post covers the pieces that decide whether agents trust it: retrieval quality, grounded drafting, confidence gating, and feedback capture.
Pipeline
- Intake: a webhook from the helpdesk sends the ticket text, customer tier, and product area.
- Retrieve: find help articles and past resolved tickets relevant to the ticket.
- Draft: generate a reply that uses only the retrieved passages.
- Gate: score the draft and either suggest it or flag it for manual handling.
- Review: the agent edits and sends, and the outcome is stored for learning.
Knowledge store schema
A relational database with a vector extension is enough for most teams. Store text, metadata, and an embedding side by side so filters and similarity search combine cleanly.
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE kb_chunk (
id BIGSERIAL PRIMARY KEY,
source_type TEXT NOT NULL, -- 'article' | 'resolved_ticket'
source_id TEXT NOT NULL,
product_area TEXT,
updated_at TIMESTAMPTZ NOT NULL,
body TEXT NOT NULL,
embedding VECTOR(1536) NOT NULL -- dimension must match your embedding model
);
CREATE INDEX ON kb_chunk USING hnsw (embedding vector_cosine_ops);
CREATE INDEX ON kb_chunk (product_area);Retrieval with filters
Filter by product area and recency before ranking. This removes most irrelevant matches cheaply.
SELECT id, source_type, source_id, body,
1 - (embedding <=> %(query_embedding)s) AS similarity
FROM kb_chunk
WHERE product_area = %(area)s
AND updated_at > now() - interval '24 months'
ORDER BY embedding <=> %(query_embedding)s
LIMIT 8;Grounded drafting
Tell the model to answer only from the supplied passages, to cite them, and to say when the passages are insufficient. Pass the passage ids so citations can be verified afterward.
SYSTEM = """You draft replies for support agents.
Use ONLY the numbered passages. Cite passage numbers like [2].
If the passages do not answer the question, reply exactly: NEEDS_HUMAN.
Match the company tone: concise, polite, no promises about refunds or dates."""
def draft(ticket, passages, gateway):
numbered = "\n\n".join(f"[{i+1}] {p.body}" for i, p in enumerate(passages))
messages = [
{"role": "system", "content": SYSTEM},
{"role": "user", "content": f"Passages:\n{numbered}\n\nTicket:\n{ticket.text}"},
]
return gateway.complete("support_copilot", "large", messages, temperature=0.2, max_tokens=500)Confidence gating
Never show a weak draft. Combine simple signals: the top retrieval similarity, whether the model returned NEEDS_HUMAN, whether every citation maps to a real passage, and whether the topic is on a deny list such as billing disputes or legal threats.
DENY_TOPICS = {"chargeback", "legal", "gdpr request", "cancel contract"}
def gate(ticket, passages, draft_text):
if draft_text.strip() == "NEEDS_HUMAN":
return False, "model_unsure"
if not passages or passages[0].similarity < 0.55:
return False, "weak_retrieval"
if any(t in ticket.text.lower() for t in DENY_TOPICS):
return False, "restricted_topic"
cited = {int(n) for n in re.findall(r"\[(\d+)\]", draft_text)}
if not cited or any(n < 1 or n > len(passages) for n in cited):
return False, "bad_citations"
return True, "ok"The 0.55 threshold is an example. Tune it on your own labeled tickets, and revisit it whenever the embedding model or content changes.
Feedback loop
Store the draft, the final sent reply, and an edit-distance score. Heavily edited drafts point to missing or stale knowledge. Weekly, review the worst ten and fix the source articles rather than the prompt.
Evaluation
- Build a set of 100 to 200 historical tickets with the reply an expert would send.
- Score each draft for factual correctness, completeness, and tone, using human review at first and an automated judge once calibrated.
- Track retrieval separately: did the right passage appear in the top eight?
Handling the hard cases
Multi-turn and attachments
Real tickets are threads, not single messages. Summarize the thread into the current question and the facts already established, and retrieve against that summary rather than the whole history. For attachments such as screenshots and logs, extract text first, and pass only the relevant lines, since long logs can crowd out the knowledge passages.
Multiple languages
Use a multilingual embedding model so that a ticket in one language can match articles in another, and ask the model to reply in the customer's language while citing the source passage numbers. Evaluate each supported language separately, because quality is not uniform.
Conflicting sources
Old articles and recent resolved tickets often disagree. Rank by recency and by source authority, with official articles above ticket history, and show the agent when two retrieved passages conflict so that a person resolves it, and the conflict becomes a content fix.
Latency budget
Agents will not wait. Budget the pipeline: roughly a second for retrieval, a few seconds for drafting, and run both while the agent opens the ticket. Stream the draft if the helpdesk supports it, and cache retrieval results per ticket so that regenerating with a different tone does not repeat the search.
async def prepare_suggestion(ticket):
summary = await summarize_thread(ticket) # small, fast model
passages = await retrieve(summary, area=ticket.area) # ~1s target
draft = await draft_reply(ticket, passages) # larger model, ~3-5s target
ok, reason = gate(ticket, passages, draft.text)
return Suggestion(text=draft.text if ok else None, passages=passages, reason=reason)Rollout in four stages
- Shadow: generate drafts silently for two weeks, and compare with what agents actually sent.
- Pilot: show suggestions to five to ten volunteer agents on one queue, with a one-click feedback control.
- Expand: add queues one at a time, and watch edit ratio and customer satisfaction on each.
- Optimize: tune the gate thresholds, retire stale articles, and add the top missing topics to the knowledge base.
What to report to leadership
Report outcomes the business recognizes: median time to first response, time to resolution, tickets per agent per day, and customer satisfaction, each compared with the pre-launch baseline and with a control group where possible. Also report the share of tickets where the copilot stayed silent. A system that knows when not to suggest is easier to trust than one that always does.
Common failure modes
- Drafts that are fluent but ignore a policy change from last month: add effective dates to articles and prefer newer ones.
- Over-confident replies on billing: keep the deny list and gate by topic.
- Agents ignoring the tool because it is slow or clumsy: measure usage and fix the interface before tuning the model.
- Feedback never reviewed: assign an owner and a weekly slot.
How we can help
We build support copilots end to end and, for teams that already have one, run a retrieval and evaluation audit that finds where drafts go wrong. Book a scoping call to discuss your helpdesk and knowledge sources.
Related reading
Need help implementing this?
Our consultants run architecture reviews and build production pilots. Book a free scoping call to talk through your design.
Book a Free Scoping Callor email us at hello@deepvero.com