Deep Dive · Oct 8, 2026 · 7 min read
Testing GenAI Myths With Quick Experiments
Instead of debating claims about GenAI, run small experiments. Here is how to test five common assumptions.
Strong opinions about GenAI, positive and negative, are cheap. Each can be turned into a hypothesis and tested in hours or days on your own data.
Myth: the model is always right
Test: run 50 real questions with known answers, and record accuracy with and without retrieval. Also test questions whose answer is not in your documents, and measure whether it abstains.
def abstention_rate(system, unanswerable_cases):
abstained = sum(1 for c in unanswerable_cases if is_abstention(system(c["input"])))
return abstained / len(unanswerable_cases)
def is_abstention(text):
return any(p in text.lower() for p in ("could not find", "not in the", "i do not know", "needs_human"))Myth: we need perfect data first
Test: build the system on one clean document set, then add a noisy one, and compare accuracy. The result tells you which data work is worth doing.
Myth: bigger models are always better
Test: run the same cases through a small and a large model and compare quality, latency, and cost. Pick the cheapest that meets your threshold.
for model in ("small", "large"):
r = run_eval(lambda q: gateway.complete("uc", model, prompt(q)).text, cases)
print(model, f"pass={r['pass_rate']:.2f}", f"p95={r['p95_latency']:.1f}s", f"cost=${r['cost']:.2f}")Myth: it will replace the team
Test: time-and-motion study. Measure how long humans take per task with and without the tool, and what fraction of outputs need human correction. This shows augmentation versus replacement for that task.
Myth: it is only for tech companies
Test: list five document-heavy or repetitive workflows in an operations team and run a baseline prompt on real samples. Feasibility depends on the task, not the industry.
Run experiments like an engineer
- State the hypothesis and the pass criterion first.
- Use real data, not cherry-picked demos.
- Record versions of prompts and models.
- Share negative results too.
Designing a fair test
A fair test uses real inputs, a fixed rubric, and the same conditions for each option. Avoid demos selected to impress. Randomize the order in which options are graded and, where possible, hide which system produced each output.
import random
def blind_pairs(cases, system_a, system_b, seed=1):
rnd = random.Random(seed)
items = []
for c in cases:
outs = [("A", system_a(c["input"])), ("B", system_b(c["input"]))]
rnd.shuffle(outs)
items.append({"case": c["id"], "first": outs[0][1], "second": outs[1][1],
"key": [outs[0][0], outs[1][0]]}) # keep the key away from graders
return itemsInterpreting small samples
Fifty cases give a rough picture, not a precise one. Report results as ranges, and look at the failures, which are often more informative than the average. If two options differ by a few cases, treat them as equivalent and choose on cost, risk, or ease of integration.
More myths to test
- It cannot handle our jargon: test with a glossary in the prompt and a handful of real examples.
- It is too slow: measure latency with streaming and realistic prompts.
- It will leak our data: review the provider terms, test redaction, and check what your logs store.
- It is too expensive: compute cost per task from measured token counts.
Recording and sharing results
Keep a short experiment log: the hypothesis, the setup, the data, the result, and the decision. Share it, including the experiments that disproved a favorite idea. Over time it becomes the organization's evidence base, and it shortens future debates.
From experiment to decision
Agree before the test what result would change your mind. If the answer is nothing, the experiment is for show. When results are in, decide, document the decision and the evidence, and set a date to revisit it, because models and prices change quickly.
How we can help
We run two-week assessment sprints that test your most important assumptions on your own data and give leadership a clear go or no-go. Ask us about one.
Related reading
Need help implementing this?
Our consultants run architecture reviews and build production pilots. Book a free scoping call to talk through your design.
Book a Free Scoping Callor email us at hello@deepvero.com