Deep Dive · Sep 1, 2026 · 7 min read
Scoring GenAI Use Cases With a Weighted Model
A transparent scoring method for choosing between candidate projects, with a small script you can reuse.
Picking a first project by enthusiasm leads to arguments. A weighted scoring model makes the trade-offs explicit and lets stakeholders debate the weights instead of the favorites.
Criteria
- Value: hours saved, revenue enabled, or risk reduced.
- Frequency: how often the task occurs.
- Data readiness: whether the needed data exists and is accessible.
- Feasibility: whether current models can do it reliably.
- Risk: cost of a wrong output, with lower risk scoring higher.
- Measurability: whether before and after can be measured.
Score and rank
WEIGHTS = {"value": 0.25, "frequency": 0.15, "data": 0.20, "feasibility": 0.20, "risk": 0.10, "measurable": 0.10}
candidates = {
"support reply drafts": {"value": 4, "frequency": 5, "data": 4, "feasibility": 4, "risk": 3, "measurable": 5},
"contract clause review": {"value": 5, "frequency": 2, "data": 3, "feasibility": 3, "risk": 2, "measurable": 3},
"meeting summaries": {"value": 2, "frequency": 5, "data": 5, "feasibility": 5, "risk": 5, "measurable": 3},
}
def score(c):
return sum(WEIGHTS[k] * c[k] for k in WEIGHTS)
for name, c in sorted(candidates.items(), key=lambda kv: -score(kv[1])):
print(f"{score(c):.2f} {name}")The scores above are illustrative. Each team scores its own candidates on a 1 to 5 scale, and the weights should reflect strategy.
Check the sensitivity
Re-run with a few alternative weight sets. If the top choice changes, the ranking is fragile and the discussion should focus on the criteria that flip it.
Probe the top two
Before committing, spend two or three days on a feasibility probe for each: run 20 real examples through a baseline prompt and look at the failures. This often reorders the list, because data and feasibility scores were guesses.
def probe(system, examples, grader):
results = [grader(ex, system(ex["input"])) for ex in examples] # grader returns 0..1
return {"mean": sum(results) / len(results), "worst": sorted(results)[:3]}Calibrating the scores
Scores are only useful if people apply them consistently. Write a one-line anchor for each level of each criterion, so that a three means the same thing to everyone. For example, for data readiness: one means the data is in people's heads, three means it exists but needs cleaning, and five means it is clean, accessible, and permissioned.
ANCHORS = {
"data": {1: "not written down", 2: "scattered across personal files", 3: "exists, needs cleaning",
4: "accessible with minor work", 5: "clean, accessible, permissioned"},
"risk": {1: "wrong output could cause serious harm", 3: "wrong output is costly but reviewable",
5: "wrong output is low impact and easy to catch"},
}Group scoring
Have three to five people score independently first, then compare. Large disagreements are the useful part: they reveal different assumptions about the data, the risk, or the value. Discuss those, adjust the anchors, and rescore. Average the final scores, and record the reasoning.
Estimating value credibly
Avoid inflated value numbers. Use a time study on a small sample: have three people time the task as they do it today. Multiply the average by the frequency and the loaded hourly cost, then apply a conservative share for how much the tool can handle. Present a range.
Using the ranking
- Pick the top candidate that also has a committed business owner and available data.
- Keep the runner-up as the next project.
- Park the highest-risk candidates until the team has built evaluation and review capability.
- Revisit the ranking each quarter as the platform and the team's skills improve.
Red flags that override the score
- No one owns the process and can decide what good looks like.
- The data cannot legally be used with the intended tools.
- Success cannot be measured, even roughly.
- The task needs perfect accuracy with no review.
From choice to charter
Turn the chosen use case into a one-page charter: the problem, the users, the success measure and threshold, the data sources, the risks, the timeline, and the named owner. This is the document that keeps the project from drifting, and the baseline for the evaluation set.
How we can help
We facilitate a half-day workshop that produces a scored shortlist, then run the feasibility probes and bring evidence to your decision. Get in touch to book one.
Related reading
Need help implementing this?
Our consultants run architecture reviews and build production pilots. Book a free scoping call to talk through your design.
Book a Free Scoping Callor email us at hello@deepvero.com