Deep Dive · Jun 2, 2026 · 7 min read
A GenAI Reference Architecture and Delivery Roadmap
The platform pieces you need before the second GenAI project, and the order to build them in.
The short version of this topic says: pilot one use case, prove value, scale. This deep dive covers what "scale" means technically. The goal is a thin shared platform, so that the second and third use cases cost a fraction of the first.
The four layers
- Applications: the user-facing surface, such as a helpdesk plugin, web app, or chat tool.
- Orchestration: prompt assembly, tool calls, retries, and control flow for each use case.
- Model gateway and retrieval: one place that calls models and one place that searches your data.
- Foundations: identity, data access, logging, evaluation, and cost tracking.
Build order
Resist building everything first. Build each layer only when a real use case needs it, in this order:
- Weeks 1-4: the first use case end to end, with a hard-coded model call and a small evaluation set.
- Weeks 5-8: extract a model gateway, so model choice, keys, and cost tracking live in one module.
- Weeks 9-12: extract retrieval as a service with access control, then add tracing.
- After that: add use cases by reusing the gateway, retrieval, and evals.
The model gateway
A gateway gives you one interface to swap models, enforce limits, and record usage. Keep it small. Its job is to normalize requests, apply policy, and log.
from dataclasses import dataclass
import time
@dataclass
class LLMResult:
text: str
input_tokens: int
output_tokens: int
latency_s: float
model: str
class Gateway:
def __init__(self, providers, policy, logger):
self.providers = providers # {"small": client, "large": client}
self.policy = policy # per-use-case limits and allowed models
self.logger = logger
def complete(self, use_case: str, tier: str, messages: list[dict], **params) -> LLMResult:
rules = self.policy[use_case]
if tier not in rules["allowed_tiers"]:
raise PermissionError(f"{use_case} may not use tier {tier}")
params["max_tokens"] = min(params.get("max_tokens", 1024), rules["max_output_tokens"])
start = time.monotonic()
raw = self.providers[tier].chat(messages=messages, **params)
result = LLMResult(
text=raw.text, input_tokens=raw.input_tokens, output_tokens=raw.output_tokens,
latency_s=time.monotonic() - start, model=raw.model,
)
self.logger.record(use_case=use_case, tier=tier, result=result)
return resultArchitecture decision records
Write down each decision that is hard to reverse, with the alternatives considered. A short record per decision saves long debates later.
id: ADR-004
title: Use a model gateway instead of direct provider SDK calls
status: accepted
context: Three use cases call models directly. Cost and model choice are not visible.
decision: All calls go through Gateway.complete(); use cases never import provider SDKs.
consequences:
- one place to add caching, rate limits, and redaction
- slight latency overhead (single hop)
- provider-specific features need an escape hatch via paramsWhere teams go wrong
- Building a platform before any use case proves value.
- Skipping the evaluation set, so no change can be measured.
- Letting each team pick its own model and vendor, which fragments cost and security review.
- Treating security and access control as a later phase.
Worked example: three use cases on one platform
Consider a mid-sized company that wants a support copilot, an internal knowledge search, and a contract review assistant. Built separately, each team would choose its own model, write its own retrieval code, and build its own evaluation scripts. Built on a shared platform, the work splits cleanly.
- The model gateway is written once and used by all three, so cost reporting, key management, and redaction rules apply everywhere.
- The retrieval service is built for the support copilot first, then extended with new sources and access rules for the other two.
- The evaluation harness is generic: each use case supplies its own golden set and rubric, and the runner, reports, and CI integration are shared.
The first use case carries the cost of building the platform pieces. The second typically reuses most of them and needs only new connectors and prompts. The third mostly needs content and evaluation work. This is the reason to extract shared components only after the first use case works, so that the abstractions are shaped by real needs.
Data and access design
Decide early which data classes the platform may touch, and encode that decision in code rather than in a policy document. A simple classification scheme, such as public, internal, confidential, and restricted, can drive routing: restricted data never leaves your network and is served only by a self-hosted model, while internal data may use an approved external provider under contract.
ALLOWED_TIERS_BY_CLASS = {
"public": {"external_large", "external_small", "self_hosted"},
"internal": {"external_large", "external_small", "self_hosted"},
"confidential": {"external_small", "self_hosted"}, # only providers approved for this class
"restricted": {"self_hosted"},
}
def pick_tier(data_class, requested):
if requested not in ALLOWED_TIERS_BY_CLASS[data_class]:
raise PermissionError(f"{requested} not allowed for {data_class} data")
return requestedThe mapping above is an example; the real one comes from your security and legal teams. The point is that the rule is enforced at the gateway, so no individual project can bypass it by accident.
Operating model
Architecture fails without ownership. Assign a small platform team that owns the gateway, retrieval service, and evaluation tooling, and embed use-case teams that own prompts, content, and acceptance criteria. Hold a short weekly review of cost, quality, and incidents across use cases, so that problems in one are fixed in all.
Risks and mitigations
- Platform sprawl: ship each component only when two use cases need it.
- Vendor lock-in: keep prompts, evaluation sets, and data exports in your own repositories, and hide providers behind the gateway.
- Quality regressions from model upgrades: pin versions, and gate every upgrade on the evaluation suite.
- Cost surprises: set per-use-case budgets in the gateway policy and alert at 80 percent.
A 90-day plan
- Days 1-30: scope, evaluation set, and the first use case end to end, with direct model calls.
- Days 31-60: extract the gateway and tracing, add cost reporting, and harden the first use case for production.
- Days 61-90: extract retrieval with access control, launch the second use case on the shared pieces, and review the platform backlog.
By day 90 you should know, with data, how much a new use case costs to deliver and run, and which platform investments are worth making next.
How we can help
We run a two-week architecture review that maps your existing systems to this reference design and produces a sequenced build plan. We can also build the first use case and the shared gateway together, so the platform grows from real demand.
Related reading
Need help implementing this?
Our consultants run architecture reviews and build production pilots. Book a free scoping call to talk through your design.
Book a Free Scoping Callor email us at hello@deepvero.com