← Software Services

Deep Dive · Aug 4, 2026 · 7 min read

How LLMs Work for Practitioners: Tokens, Context, Sampling, and Cost

The mechanics that explain price, latency, and quirks, in the depth an engineering lead needs.

Two bar charts of next-token probabilities at low and high temperatureShort version← What Is Generative AI? A Plain-English Guide for Business Leaders

To estimate cost, debug odd behavior, or design prompts, you need a working model of what an LLM does. This is that model, without the math.

Line chart showing request cost rising with token counts for input and output
Shorter prompts and capped outputs are the cheapest levers.

Tokens

Models do not read characters or words. A tokenizer splits text into tokens, often word fragments. A rough rule for English is about four characters per token, but other languages and code often use more. You pay per token, and the context window is measured in tokens.

Next-token prediction

At each step the model produces a probability distribution over the next token, picks one, appends it, and repeats. Everything else, such as answering questions or following instructions, emerges from this loop plus training.

Sampling parameters

  • Temperature: higher values flatten the distribution and increase variety; low values make output more repeatable.
  • Top-p: sample only from the smallest set of tokens whose probabilities add up to p.
  • Max tokens: a hard cap on output length, which also caps cost.
  • Even at low temperature, outputs may vary slightly across calls, so test for consistency rather than assuming identical results.

The context window

Everything the model sees, including instructions, retrieved passages, chat history, and its own output, must fit in the window. Long inputs cost more, run slower, and can dilute attention to key facts. Budget the window deliberately.

context_budget.py
def budget(window=128_000, reserve_output=2_000, system=800, history=6_000):
    """Tokens left for retrieved context in one request (illustrative numbers)."""
    return window - reserve_output - system - history

def fits(passages_tokens, window_budget):
    kept, used = [], 0
    for t in passages_tokens:            # already ranked best first
        if used + t > window_budget:
            break
        kept.append(t); used += t
    return kept

Estimating cost

Cost is input tokens times the input price plus output tokens times the output price. Output tokens usually cost more. Multiply by request volume to get a monthly figure.

cost.py
def monthly_cost(requests_per_day, in_tokens, out_tokens, in_price_per_m, out_price_per_m, days=30):
    per_request = (in_tokens * in_price_per_m + out_tokens * out_price_per_m) / 1_000_000
    return per_request * requests_per_day * days

# Prices are inputs: take them from your provider's current price list.
print(monthly_cost(2000, 3500, 400, in_price_per_m=PRICE_IN, out_price_per_m=PRICE_OUT))

Levers that reduce cost without hurting quality: shorter prompts, caching repeated context, retrieving fewer but better passages, and routing easy requests to a smaller model.

Practical consequences

  • Quirks with counting, spelling, and exact arithmetic come from tokenization; use code tools for exact work.
  • Models have a training cutoff; give them current facts through retrieval.
  • Instructions placed clearly and repeated near the end of long prompts are followed more reliably.

Why outputs vary, and how to control it

Even with the same prompt, outputs can differ. Sampling is one source. Provider-side changes, batching, and floating-point effects are others. Design for this: test with multiple runs, avoid depending on exact wording, and use structured output with validation when downstream code needs a stable format. Where you need repeatability, lower the temperature, fix model versions, and store the outputs you rely on.

Structured output

Most business integrations need JSON, not prose. Many providers support schema-constrained output. Whatever the mechanism, always parse and validate in code, and have a bounded retry that passes the validation error back to the model.

structured.py
def get_structured(gateway, messages, model_cls, retries=2):
    for attempt in range(retries + 1):
        text = gateway.complete("extract", "small", messages, temperature=0).text
        try:
            return model_cls.model_validate_json(text)
        except Exception as e:
            messages = messages + [
                {"role": "assistant", "content": text},
                {"role": "user", "content": f"That was invalid: {e}. Return corrected JSON only."},
            ]
    raise ValueError("could not obtain valid output")

Latency

Latency has two parts: time to first token, which grows with prompt length, and generation time, which grows with output length. Shorten prompts, cap the output, stream results to the user, and parallelize independent calls. For interactive tools, perceived speed from streaming often matters more than total time.

Caching

  • Exact-match cache for identical requests, keyed on model, prompt, and parameters.
  • Prompt caching offered by some providers for repeated prefixes, which makes long shared instructions cheaper; check your provider's current rules.
  • Embedding cache so unchanged text is not re-embedded.

Limits to plan around

  • Knowledge cutoff: models do not know recent events unless you supply them.
  • Context limits: long documents need chunking or summarization.
  • Reasoning errors: multi-step arithmetic and logic can fail; give the model tools for exact computation.
  • Bias and variability: test across the groups and inputs you serve.

A mental model for design

Treat the model as a fast, fallible reader and writer that follows instructions approximately. Give it the facts, constrain the output, check what it returns, and keep people responsible for consequential decisions. Systems built with that mental model are easier to test, cheaper to run, and safer to trust.

How we can help

We help teams size cost and latency before building, and we tune prompts and routing on existing systems to cut spend. Ask us for a short cost and architecture review.

Related reading

Need help implementing this?

Our consultants run architecture reviews and build production pilots. Book a free scoping call to talk through your design.

Book a Free Scoping Call

or email us at hello@deepvero.com