Deep Dive · Aug 18, 2026 · 7 min read
A Total Cost of Ownership Model for GenAI Systems
A worked cost model: build, run, and operate, with the levers that move each line.
Model price per token is the most visible cost and rarely the largest. A useful estimate covers build effort, inference, infrastructure, and the people who keep the system good.
Cost components
- Build: discovery, integration, evaluation set creation, security review.
- Inference: input and output tokens times request volume.
- Infrastructure: vector store, queues, hosting, monitoring.
- Operations: on-call, content updates, model upgrades, evaluation runs.
- Human review: time spent approving or correcting outputs.
A simple model
def tco_12_months(build_cost, requests_per_month, tokens_in, tokens_out,
price_in_per_m, price_out_per_m, infra_per_month,
ops_hours_per_month, hourly_rate, review_minutes_per_100_requests, reviewer_rate):
inference = requests_per_month * 12 * (tokens_in * price_in_per_m + tokens_out * price_out_per_m) / 1_000_000
infra = infra_per_month * 12
ops = ops_hours_per_month * 12 * hourly_rate
review = requests_per_month * 12 / 100 * review_minutes_per_100_requests / 60 * reviewer_rate
return {"build": build_cost, "inference": inference, "infra": infra, "ops": ops, "review": review,
"total": build_cost + inference + infra + ops + review}Every input is yours to fill in. The structure matters more than any default: it makes the assumptions explicit and lets you test sensitivity.
Sensitivity analysis
Vary each input by plus and minus 50 percent and see which moves the total most. Often it is review time or request volume, not model price.
def sensitivity(base_inputs, fn, keys, delta=0.5):
base = fn(**base_inputs)["total"]
out = {}
for k in keys:
lo = fn(**{**base_inputs, k: base_inputs[k] * (1 - delta)})["total"]
hi = fn(**{**base_inputs, k: base_inputs[k] * (1 + delta)})["total"]
out[k] = (hi - lo) / base # relative swing of the total
return dict(sorted(out.items(), key=lambda kv: -kv[1]))Levers
- Route simple requests to a smaller model.
- Cache repeated prompts and shared context.
- Retrieve fewer, better passages to shorten prompts.
- Batch non-urgent work, where providers offer discounted batch processing.
- Automate review for low-risk outputs once quality is proven.
Compare with the benefit
Estimate benefit in the same units: hours saved times loaded hourly cost, plus error reduction. A pilot is worth it when the benefit comfortably exceeds the 12-month total under pessimistic assumptions.
A worked example with placeholder numbers
Suppose a support copilot handles 20,000 requests a month, with about 3,000 input tokens and 400 output tokens per request. At your provider's prices, inference is a straightforward multiplication, and for many workloads it turns out to be a modest fraction of the total. The larger lines are usually the one-time build, the engineering time to keep content and integrations current, and the human review time. This is why the model above separates them.
inputs = dict(build_cost=BUILD, requests_per_month=20_000, tokens_in=3_000, tokens_out=400,
price_in_per_m=PRICE_IN, price_out_per_m=PRICE_OUT, infra_per_month=INFRA,
ops_hours_per_month=24, hourly_rate=RATE, review_minutes_per_100_requests=30, reviewer_rate=REVIEW_RATE)
result = tco_12_months(**inputs)
print({k: round(v) for k, v in result.items()})
print(sensitivity(inputs, tco_12_months, ["requests_per_month", "tokens_in", "review_minutes_per_100_requests", "ops_hours_per_month"]))All the capitalized names are values you supply. Run the sensitivity analysis and read it as a ranked list of what to measure in the pilot.
Fixed versus variable costs
- Mostly fixed: build, integration, platform tooling, and ongoing engineering.
- Mostly variable: inference and review time, which scale with usage.
- Step costs: new regions, new languages, or new data sources that require additional engineering.
Because variable costs scale, success can raise your bill. Model the high-adoption scenario as well as the base case, and check that unit economics still work.
Cost control tactics, in order of effort
- Low effort: shorter prompts, lower output caps, and removing unneeded context.
- Medium effort: caching, batching of non-urgent work, and routing easy requests to a smaller model.
- Higher effort: fine-tuning a small model for a narrow task, or self-hosting where volume justifies the operations cost.
Budgeting and monitoring
Put a budget in the gateway policy for each use case, and alert at defined percentages. Report cost per successful outcome, such as cost per resolved ticket, not just tokens, because that is the number that connects to value. Review monthly and after any model or prompt change.
What teams forget
- The cost of building and maintaining the evaluation set.
- Security review, legal review, and procurement time.
- Change management and training.
- Model migrations when a provider retires a version.
Presenting the model to finance
Share the assumptions, not just the totals. Show the base, pessimistic, and optimistic cases, list what you will measure in the pilot to replace each assumption, and agree the decision rule in advance. Finance teams trust a model they can interrogate far more than a single optimistic figure.
How we can help
We build this model with you before any code is written, using your real volumes and workflows, and we tell you honestly when a use case does not pay back. Ask for a cost and value assessment.
Related reading
Need help implementing this?
Our consultants run architecture reviews and build production pilots. Book a free scoping call to talk through your design.
Book a Free Scoping Callor email us at hello@deepvero.com