← Software Services

Deep Dive · Oct 8, 2026 · 7 min read

Testing GenAI Myths With Quick Experiments

Instead of debating claims about GenAI, run small experiments. Here is how to test five common assumptions.

Grouped bars comparing prompt-only and retrieval-based systems on accuracy, abstention and cost efficiencyShort version← 5 GenAI Myths Business Leaders Should Drop

Strong opinions about GenAI, positive and negative, are cheap. Each can be turned into a hypothesis and tested in hours or days on your own data.

Horizontal bars of abstention rates for three systems
Test questions whose answer is not in your documents.

Myth: the model is always right

Test: run 50 real questions with known answers, and record accuracy with and without retrieval. Also test questions whose answer is not in your documents, and measure whether it abstains.

abstention.py
def abstention_rate(system, unanswerable_cases):
    abstained = sum(1 for c in unanswerable_cases if is_abstention(system(c["input"])))
    return abstained / len(unanswerable_cases)

def is_abstention(text):
    return any(p in text.lower() for p in ("could not find", "not in the", "i do not know", "needs_human"))

Myth: we need perfect data first

Test: build the system on one clean document set, then add a noisy one, and compare accuracy. The result tells you which data work is worth doing.

Myth: bigger models are always better

Test: run the same cases through a small and a large model and compare quality, latency, and cost. Pick the cheapest that meets your threshold.

model-comparison.py
for model in ("small", "large"):
    r = run_eval(lambda q: gateway.complete("uc", model, prompt(q)).text, cases)
    print(model, f"pass={r['pass_rate']:.2f}", f"p95={r['p95_latency']:.1f}s", f"cost=${r['cost']:.2f}")

Myth: it will replace the team

Test: time-and-motion study. Measure how long humans take per task with and without the tool, and what fraction of outputs need human correction. This shows augmentation versus replacement for that task.

Myth: it is only for tech companies

Test: list five document-heavy or repetitive workflows in an operations team and run a baseline prompt on real samples. Feasibility depends on the task, not the industry.

Run experiments like an engineer

  • State the hypothesis and the pass criterion first.
  • Use real data, not cherry-picked demos.
  • Record versions of prompts and models.
  • Share negative results too.

Designing a fair test

A fair test uses real inputs, a fixed rubric, and the same conditions for each option. Avoid demos selected to impress. Randomize the order in which options are graded and, where possible, hide which system produced each output.

blind.py
import random

def blind_pairs(cases, system_a, system_b, seed=1):
    rnd = random.Random(seed)
    items = []
    for c in cases:
        outs = [("A", system_a(c["input"])), ("B", system_b(c["input"]))]
        rnd.shuffle(outs)
        items.append({"case": c["id"], "first": outs[0][1], "second": outs[1][1],
                      "key": [outs[0][0], outs[1][0]]})      # keep the key away from graders
    return items

Interpreting small samples

Fifty cases give a rough picture, not a precise one. Report results as ranges, and look at the failures, which are often more informative than the average. If two options differ by a few cases, treat them as equivalent and choose on cost, risk, or ease of integration.

More myths to test

  • It cannot handle our jargon: test with a glossary in the prompt and a handful of real examples.
  • It is too slow: measure latency with streaming and realistic prompts.
  • It will leak our data: review the provider terms, test redaction, and check what your logs store.
  • It is too expensive: compute cost per task from measured token counts.

Recording and sharing results

Keep a short experiment log: the hypothesis, the setup, the data, the result, and the decision. Share it, including the experiments that disproved a favorite idea. Over time it becomes the organization's evidence base, and it shortens future debates.

From experiment to decision

Agree before the test what result would change your mind. If the answer is nothing, the experiment is for show. When results are in, decide, document the decision and the evidence, and set a date to revisit it, because models and prices change quickly.

How we can help

We run two-week assessment sprints that test your most important assumptions on your own data and give leadership a clear go or no-go. Ask us about one.

Related reading

Need help implementing this?

Our consultants run architecture reviews and build production pilots. Book a free scoping call to talk through your design.

Book a Free Scoping Call

or email us at hello@deepvero.com