Deep Dive · Aug 4, 2026 · 7 min read
How LLMs Work for Practitioners: Tokens, Context, Sampling, and Cost
The mechanics that explain price, latency, and quirks, in the depth an engineering lead needs.
To estimate cost, debug odd behavior, or design prompts, you need a working model of what an LLM does. This is that model, without the math.
Tokens
Models do not read characters or words. A tokenizer splits text into tokens, often word fragments. A rough rule for English is about four characters per token, but other languages and code often use more. You pay per token, and the context window is measured in tokens.
Next-token prediction
At each step the model produces a probability distribution over the next token, picks one, appends it, and repeats. Everything else, such as answering questions or following instructions, emerges from this loop plus training.
Sampling parameters
- Temperature: higher values flatten the distribution and increase variety; low values make output more repeatable.
- Top-p: sample only from the smallest set of tokens whose probabilities add up to p.
- Max tokens: a hard cap on output length, which also caps cost.
- Even at low temperature, outputs may vary slightly across calls, so test for consistency rather than assuming identical results.
The context window
Everything the model sees, including instructions, retrieved passages, chat history, and its own output, must fit in the window. Long inputs cost more, run slower, and can dilute attention to key facts. Budget the window deliberately.
def budget(window=128_000, reserve_output=2_000, system=800, history=6_000):
"""Tokens left for retrieved context in one request (illustrative numbers)."""
return window - reserve_output - system - history
def fits(passages_tokens, window_budget):
kept, used = [], 0
for t in passages_tokens: # already ranked best first
if used + t > window_budget:
break
kept.append(t); used += t
return keptEstimating cost
Cost is input tokens times the input price plus output tokens times the output price. Output tokens usually cost more. Multiply by request volume to get a monthly figure.
def monthly_cost(requests_per_day, in_tokens, out_tokens, in_price_per_m, out_price_per_m, days=30):
per_request = (in_tokens * in_price_per_m + out_tokens * out_price_per_m) / 1_000_000
return per_request * requests_per_day * days
# Prices are inputs: take them from your provider's current price list.
print(monthly_cost(2000, 3500, 400, in_price_per_m=PRICE_IN, out_price_per_m=PRICE_OUT))Levers that reduce cost without hurting quality: shorter prompts, caching repeated context, retrieving fewer but better passages, and routing easy requests to a smaller model.
Practical consequences
- Quirks with counting, spelling, and exact arithmetic come from tokenization; use code tools for exact work.
- Models have a training cutoff; give them current facts through retrieval.
- Instructions placed clearly and repeated near the end of long prompts are followed more reliably.
Why outputs vary, and how to control it
Even with the same prompt, outputs can differ. Sampling is one source. Provider-side changes, batching, and floating-point effects are others. Design for this: test with multiple runs, avoid depending on exact wording, and use structured output with validation when downstream code needs a stable format. Where you need repeatability, lower the temperature, fix model versions, and store the outputs you rely on.
Structured output
Most business integrations need JSON, not prose. Many providers support schema-constrained output. Whatever the mechanism, always parse and validate in code, and have a bounded retry that passes the validation error back to the model.
def get_structured(gateway, messages, model_cls, retries=2):
for attempt in range(retries + 1):
text = gateway.complete("extract", "small", messages, temperature=0).text
try:
return model_cls.model_validate_json(text)
except Exception as e:
messages = messages + [
{"role": "assistant", "content": text},
{"role": "user", "content": f"That was invalid: {e}. Return corrected JSON only."},
]
raise ValueError("could not obtain valid output")Latency
Latency has two parts: time to first token, which grows with prompt length, and generation time, which grows with output length. Shorten prompts, cap the output, stream results to the user, and parallelize independent calls. For interactive tools, perceived speed from streaming often matters more than total time.
Caching
- Exact-match cache for identical requests, keyed on model, prompt, and parameters.
- Prompt caching offered by some providers for repeated prefixes, which makes long shared instructions cheaper; check your provider's current rules.
- Embedding cache so unchanged text is not re-embedded.
Limits to plan around
- Knowledge cutoff: models do not know recent events unless you supply them.
- Context limits: long documents need chunking or summarization.
- Reasoning errors: multi-step arithmetic and logic can fail; give the model tools for exact computation.
- Bias and variability: test across the groups and inputs you serve.
A mental model for design
Treat the model as a fast, fallible reader and writer that follows instructions approximately. Give it the facts, constrain the output, check what it returns, and keep people responsible for consequential decisions. Systems built with that mental model are easier to test, cheaper to run, and safer to trust.
How we can help
We help teams size cost and latency before building, and we tune prompts and routing on existing systems to cut spend. Ask us for a short cost and architecture review.
Related reading
Need help implementing this?
Our consultants run architecture reviews and build production pilots. Book a free scoping call to talk through your design.
Book a Free Scoping Callor email us at hello@deepvero.com