← Software Services

Deep Dive · Jul 28, 2026 · 7 min read

Our GenAI Delivery Framework: Eval-Driven Development in 30 Days

How we structure a month-long build around an evaluation set, so quality is measured from week one.

Heatmap of evaluation pass rate by category over five versionsShort version← How We Run a GenAI Project, From Kickoff to Launch

The short post describes our process in plain terms. This version shows the engineering practice underneath: we build the test set before the system, and every change must pass it.

Three cards: development set, held-out set and CI gate
Quality is a number the whole team can see.

Week 1: scope and evaluation set

  • Write the success criteria as numbers, such as a target accuracy and a latency limit.
  • Collect 50 to 150 real examples with expected outputs, covering normal, edge, and adversarial cases.
  • Split into a development set for iterating and a held-out set that is only run before release.
eval/cases.jsonl
{"id": "case-0042",
 "input": "Can I return an opened item after 30 days?",
 "expected": {"must_include": ["30 days", "unopened"], "must_not_include": ["guarantee"]},
 "tags": ["returns", "edge-case"],
 "split": "dev"}

Weeks 2-3: iterate against the dev set

Change one thing at a time: the prompt, retrieval settings, or model. After each change run the dev set and compare against the previous run. Keep every run's results so regressions are visible.

eval/run.py
def run_eval(system, cases):
    results = []
    for case in cases:
        output = system(case["input"])
        ok_inc = all(s.lower() in output.lower() for s in case["expected"].get("must_include", []))
        ok_exc = not any(s.lower() in output.lower() for s in case["expected"].get("must_not_include", []))
        results.append({"id": case["id"], "pass": ok_inc and ok_exc, "tags": case["tags"]})
    passed = sum(r["pass"] for r in results)
    return {"pass_rate": passed / len(results), "failures": [r["id"] for r in results if not r["pass"]]}

String checks are a starting point. For open-ended outputs, add a rubric-based model judge, but calibrate it against human scores on a sample first.

Week 4: release gate and launch

Run the held-out set once. Release only if it meets the thresholds from week 1. Wire the same check into CI so later changes cannot silently lower quality.

ci-gate.yml
# .github/workflows/eval.yml (excerpt)
- name: Run evaluation
  run: python eval/run.py --split heldout --min-pass-rate 0.90 --max-p95-latency 4.0

The thresholds above are examples. Set yours from the business need, not from what the first prototype achieved.

What this gives you

  • Quality is a number the whole team can see.
  • Changes are safe to make, because regressions are caught.
  • Stakeholders get evidence at launch, not opinions.

How to build a good evaluation set

The quality of the set decides the quality of the project. Draw cases from real usage whenever possible: historical tickets, real questions from a shared inbox, or logged searches. Add cases for the failures people worry about, for known edge cases, and for adversarial inputs. Keep each case independent, and write the expected result before you see the system's output, to avoid fitting the test to the answer.

  • Cover the realistic spread of inputs, not only the easy ones.
  • Include cases where the right behavior is to decline or escalate.
  • Label each case with tags so results can be broken down by category.
  • Review the set with a domain expert, and fix disputed expectations.

Metrics that fit the task

  • Classification and extraction: exact match, precision, and recall per field.
  • Question answering: correctness against a reference, groundedness in sources, and abstention behavior.
  • Drafting: a rubric score for accuracy, completeness, and tone, plus the edit ratio from real users.
  • All tasks: latency percentiles and cost per request.

Model judges, carefully

For open-ended outputs, a model judge scales evaluation, but it needs validation. Have two humans grade 50 outputs, measure how often the judge agrees with them, and read the disagreements. Where the judge is unreliable, fix the rubric, or keep humans in the loop for that dimension.

Regressions and flakiness

Generative systems vary between runs. For each case, run several samples when a single failure would be noisy, and report pass rates instead of single outcomes. Track the suite over time, and treat a drop of more than the normal run-to-run variation as a regression to investigate.

samples.py
def pass_rate_with_samples(system, case, grader, n=5):
    results = [grader(case, system(case["input"])) for _ in range(n)]
    return sum(results) / n

What a good weekly demo looks like

Each week, show the system running on real cases, show the evaluation results next to last week's, and show the top five failures with the plan for each. Stakeholders see progress as a trend, and disagreements about quality are settled by looking at the cases instead of by opinion.

Definition of done

  • The held-out set passes the agreed thresholds.
  • Monitoring and alerting are live, and an owner is named.
  • Security and access reviews are complete.
  • Users are trained, and a feedback route exists.
  • The evaluation set and runner are handed over, so the client can keep testing.

How we can help

We run this framework as a fixed-scope engagement: scope, evaluation set, build, and release gate in about 30 days. We also set it up for teams that already have a prototype and no way to measure it.

Related reading

Need help implementing this?

Our consultants run architecture reviews and build production pilots. Book a free scoping call to talk through your design.

Book a Free Scoping Call

or email us at hello@deepvero.com