Prioritising experiments: a simple scoring model
A backlog of 50 ideas and no way to choose is why experimentation stalls. A simple impact-effort-confidence scoring model to prioritise the tests that matter.

The real bottleneck isn’t ideas — it’s choosing
Most growth teams have no shortage of experiment ideas. What they lack is a way to choose between them, so they default to whatever’s loudest, easiest, or most recent. That’s how programmes stall: effort scatters across low-impact tests while the high-impact ones wait. A simple, shared scoring model fixes this — not because the maths is precise, but because it forces an honest, comparable conversation about what to run next.
A simple scoring model
Score each idea on three dimensions, each 1–10:
- Impact — if this works, how much does it move the metric that matters right now?
- Confidence — how sure are we it’ll work, based on evidence, precedent or logic?
- Ease — how quick and cheap is it to run?
Combine them (a simple average or product) into one score, and rank the backlog. Ideas with high impact, reasonable confidence and low effort rise to the top; pet projects with low scores stay honest at the bottom.
Why it works even though it’s rough
The scores are estimates, not truth — and that’s fine. The value isn’t precision; it’s that scoring forces the team to articulate why an idea matters, exposes low-value pet projects, and creates a shared, defensible order. It turns “what should we test?” from an argument into a ranked queue. Tie the scoring to your current growth constraint so “impact” always means impact on the thing that matters this quarter.
Keep it lightweight
The model earns its keep only if it’s fast. A shared sheet, three quick scores per idea, re-ranked weekly, is enough. Don’t turn prioritisation into a project — the goal is to decide quickly and get back to testing. (This feeds directly into running 3–4 experiments a week.)
Frequently asked questions
Isn’t scoring subjective?
Yes — deliberately. The point is a shared, comparable conversation, not false precision. Rough scores beat loudest-voice prioritisation.
Which framework — ICE, PIE, RICE?
Any consistent one works. Impact/Confidence/Ease is the simplest; use what your team will actually maintain.
How often should we re-score?
Weekly, lightly — as results come in, confidence and impact estimates change, so the ranking should too.
Backlog full, direction unclear? A Growth Diagnostic helps you prioritise what to test first. Request a Growth Diagnostic →
Leave a Reply