← Software Services

Deep Dive · Jun 16, 2026 · 7 min read

Engineering a Support Copilot: Retrieval, Drafting, and Review

A technical walk through ticket intake, hybrid retrieval, grounded drafting, confidence gating, and the feedback loop.

Swimlane of a support copilot from ticket creation through retrieval, drafting, a safety gate and agent reviewShort version← Building a Customer Support Copilot: What It Takes

A support copilot looks simple in a demo and gets hard in production. This post covers the pieces that decide whether agents trust it: retrieval quality, grounded drafting, confidence gating, and feedback capture.

Two distributions of drafts by retrieval score with thresholds that route to no suggestion, flagged review, or show draft
Calibrate thresholds on labeled tickets, not on intuition.

Pipeline

  • Intake: a webhook from the helpdesk sends the ticket text, customer tier, and product area.
  • Retrieve: find help articles and past resolved tickets relevant to the ticket.
  • Draft: generate a reply that uses only the retrieved passages.
  • Gate: score the draft and either suggest it or flag it for manual handling.
  • Review: the agent edits and sends, and the outcome is stored for learning.

Knowledge store schema

A relational database with a vector extension is enough for most teams. Store text, metadata, and an embedding side by side so filters and similarity search combine cleanly.

schema.sql
CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE kb_chunk (
    id            BIGSERIAL PRIMARY KEY,
    source_type   TEXT NOT NULL,          -- 'article' | 'resolved_ticket'
    source_id     TEXT NOT NULL,
    product_area  TEXT,
    updated_at    TIMESTAMPTZ NOT NULL,
    body          TEXT NOT NULL,
    embedding     VECTOR(1536) NOT NULL   -- dimension must match your embedding model
);

CREATE INDEX ON kb_chunk USING hnsw (embedding vector_cosine_ops);
CREATE INDEX ON kb_chunk (product_area);

Retrieval with filters

Filter by product area and recency before ranking. This removes most irrelevant matches cheaply.

retrieve.sql
SELECT id, source_type, source_id, body,
       1 - (embedding <=> %(query_embedding)s) AS similarity
FROM kb_chunk
WHERE product_area = %(area)s
  AND updated_at > now() - interval '24 months'
ORDER BY embedding <=> %(query_embedding)s
LIMIT 8;

Grounded drafting

Tell the model to answer only from the supplied passages, to cite them, and to say when the passages are insufficient. Pass the passage ids so citations can be verified afterward.

draft.py
SYSTEM = """You draft replies for support agents.
Use ONLY the numbered passages. Cite passage numbers like [2].
If the passages do not answer the question, reply exactly: NEEDS_HUMAN.
Match the company tone: concise, polite, no promises about refunds or dates."""

def draft(ticket, passages, gateway):
    numbered = "\n\n".join(f"[{i+1}] {p.body}" for i, p in enumerate(passages))
    messages = [
        {"role": "system", "content": SYSTEM},
        {"role": "user", "content": f"Passages:\n{numbered}\n\nTicket:\n{ticket.text}"},
    ]
    return gateway.complete("support_copilot", "large", messages, temperature=0.2, max_tokens=500)

Confidence gating

Never show a weak draft. Combine simple signals: the top retrieval similarity, whether the model returned NEEDS_HUMAN, whether every citation maps to a real passage, and whether the topic is on a deny list such as billing disputes or legal threats.

gate.py
DENY_TOPICS = {"chargeback", "legal", "gdpr request", "cancel contract"}

def gate(ticket, passages, draft_text):
    if draft_text.strip() == "NEEDS_HUMAN":
        return False, "model_unsure"
    if not passages or passages[0].similarity < 0.55:
        return False, "weak_retrieval"
    if any(t in ticket.text.lower() for t in DENY_TOPICS):
        return False, "restricted_topic"
    cited = {int(n) for n in re.findall(r"\[(\d+)\]", draft_text)}
    if not cited or any(n < 1 or n > len(passages) for n in cited):
        return False, "bad_citations"
    return True, "ok"

The 0.55 threshold is an example. Tune it on your own labeled tickets, and revisit it whenever the embedding model or content changes.

Feedback loop

Store the draft, the final sent reply, and an edit-distance score. Heavily edited drafts point to missing or stale knowledge. Weekly, review the worst ten and fix the source articles rather than the prompt.

Evaluation

  • Build a set of 100 to 200 historical tickets with the reply an expert would send.
  • Score each draft for factual correctness, completeness, and tone, using human review at first and an automated judge once calibrated.
  • Track retrieval separately: did the right passage appear in the top eight?

Handling the hard cases

Multi-turn and attachments

Real tickets are threads, not single messages. Summarize the thread into the current question and the facts already established, and retrieve against that summary rather than the whole history. For attachments such as screenshots and logs, extract text first, and pass only the relevant lines, since long logs can crowd out the knowledge passages.

Multiple languages

Use a multilingual embedding model so that a ticket in one language can match articles in another, and ask the model to reply in the customer's language while citing the source passage numbers. Evaluate each supported language separately, because quality is not uniform.

Conflicting sources

Old articles and recent resolved tickets often disagree. Rank by recency and by source authority, with official articles above ticket history, and show the agent when two retrieved passages conflict so that a person resolves it, and the conflict becomes a content fix.

Latency budget

Agents will not wait. Budget the pipeline: roughly a second for retrieval, a few seconds for drafting, and run both while the agent opens the ticket. Stream the draft if the helpdesk supports it, and cache retrieval results per ticket so that regenerating with a different tone does not repeat the search.

prepare.py
async def prepare_suggestion(ticket):
    summary = await summarize_thread(ticket)                      # small, fast model
    passages = await retrieve(summary, area=ticket.area)          # ~1s target
    draft = await draft_reply(ticket, passages)                   # larger model, ~3-5s target
    ok, reason = gate(ticket, passages, draft.text)
    return Suggestion(text=draft.text if ok else None, passages=passages, reason=reason)

Rollout in four stages

  • Shadow: generate drafts silently for two weeks, and compare with what agents actually sent.
  • Pilot: show suggestions to five to ten volunteer agents on one queue, with a one-click feedback control.
  • Expand: add queues one at a time, and watch edit ratio and customer satisfaction on each.
  • Optimize: tune the gate thresholds, retire stale articles, and add the top missing topics to the knowledge base.

What to report to leadership

Report outcomes the business recognizes: median time to first response, time to resolution, tickets per agent per day, and customer satisfaction, each compared with the pre-launch baseline and with a control group where possible. Also report the share of tickets where the copilot stayed silent. A system that knows when not to suggest is easier to trust than one that always does.

Common failure modes

  • Drafts that are fluent but ignore a policy change from last month: add effective dates to articles and prefer newer ones.
  • Over-confident replies on billing: keep the deny list and gate by topic.
  • Agents ignoring the tool because it is slow or clumsy: measure usage and fix the interface before tuning the model.
  • Feedback never reviewed: assign an owner and a weekly slot.

How we can help

We build support copilots end to end and, for teams that already have one, run a retrieval and evaluation audit that finds where drafts go wrong. Book a scoping call to discuss your helpdesk and knowledge sources.

Related reading

Need help implementing this?

Our consultants run architecture reviews and build production pilots. Book a free scoping call to talk through your design.

Book a Free Scoping Call

or email us at hello@deepvero.com