Deep Dive · Aug 25, 2026 · 7 min read
From Pilot to Production: A GenAI Readiness Checklist
The engineering gaps between a convincing demo and a system people rely on, with a checklist to close them.
Pilots stall for organizational reasons, but there is also a concrete technical gap. A demo works on the examples its builder tried. Production must work on everything users throw at it, every day, within budget.
Readiness checklist
Quality
- An evaluation set drawn from real usage, with a pass threshold agreed in advance.
- Known failure categories and a plan for each.
- A regression run on every prompt, model, or retrieval change.
Reliability
- Timeouts, retries with backoff, and fallbacks when a provider is down.
- Rate limits and queueing for bursts.
- Graceful behavior when the model returns malformed output.
import random, time
def call_with_retry(fn, attempts=4, base=0.5, retry_on=(TimeoutError, ConnectionError)):
for i in range(attempts):
try:
return fn()
except retry_on:
if i == attempts - 1:
raise
time.sleep(base * (2 ** i) + random.uniform(0, 0.25)) # exponential backoff with jitterSecurity and access
- Authentication through your existing identity provider.
- Per-user data access enforced in retrieval, not just in the UI.
- Redaction of sensitive fields before they reach external models.
Operations
- Tracing for each request: prompt, retrieved context, output, latency, and cost.
- Dashboards and alerts for error rate, latency, and spend.
- A named owner and an on-call route.
People
- Training for users and a feedback button inside the tool.
- A rollout plan: pilot group, then wider release.
- A rollback plan.
Cutover pattern
Run the new system in shadow mode first: it produces outputs that are logged but not shown. Compare with what humans did, fix the gaps, then enable it for a small group behind a feature flag.
def handle(request, flags, legacy, copilot, log):
result = legacy(request)
if flags.enabled("copilot_shadow", request.user):
log.record(request, copilot(request)) # not shown to the user
if flags.enabled("copilot_live", request.user):
return copilot(request)
return resultScoring readiness
A checklist is more useful when it produces a decision. Score each area from zero to three: zero means missing, one means partial, two means adequate for a limited release, and three means ready for general availability. Release to a limited group when every area is at least two and the critical ones, such as security and quality, are at least two with a plan to reach three.
AREAS = ["quality", "reliability", "security", "operations", "people", "rollback"]
CRITICAL = {"quality", "security", "rollback"}
def decide(scores):
if any(scores[a] < 2 for a in CRITICAL):
return "not ready"
if all(scores[a] >= 3 for a in AREAS):
return "general availability"
if all(scores[a] >= 2 for a in AREAS):
return "limited release"
return "close the gaps first"Load and failure testing
Test the system under realistic peak load, not only the demo path. Simulate a provider outage, a slow response, malformed output, and an unavailable data source, and confirm that the fallback behaves as designed. Check that the system fails safely: a clear message and a route to a person, not a silent wrong answer.
Support model
- Who answers user questions in the first month?
- Who triages quality complaints, and what is the response time?
- How are incidents declared, and who can switch the system off?
- How do changes get approved and released?
Data protection review
Before launch, confirm what data the system sees, where it is sent, what is stored, and for how long. Check that logs do not retain sensitive content longer than allowed, and that deleting a source document removes it from the index. Document the review, because auditors and customers will ask.
Launch communication
Tell users what the tool is for, what it is not for, and how to give feedback. Share examples of good use. Be candid about limits, because trust is built when the tool behaves as described.
Post-launch plan
- A two-week hypercare period with daily review of feedback and failures.
- A monthly review of quality, cost, adoption, and the improvement backlog.
- A scheduled re-run of the readiness assessment after major changes.
The checklist is a tool for having the conversation at the right time. Run it early, when gaps are cheap to close, and again just before launch.
How we can help
We run a one-week production readiness review against this checklist and deliver a prioritized remediation plan. We also build the monitoring and rollout tooling when your team is stretched.
Related reading
Need help implementing this?
Our consultants run architecture reviews and build production pilots. Book a free scoping call to talk through your design.
Book a Free Scoping Callor email us at hello@deepvero.com