ICE prioritisation: choosing the experiment that deserves capacity
ICE prioritisation makes growth work deliberate when a team has far more ideas than it can run. It is not a claim that a score is perfectly objective; it is a shared discipline for deciding which idea should consume scarce time first.
Score every dimension from 1 to 10, then rank ideas in descending order. The usefulness lies less in the arithmetic than in making the evidence, constraint, effort and disagreement explicit.
| Dimension | Question | High score means | Discipline required |
|---|---|---|---|
| Impact | How much will this move the metric that matters? | Direct, material movement in the primary growth constraint | Do not award a high score to a downstream or unrelated metric |
| Confidence | How sure are we that the predicted effect will occur? | Specific, direct evidence | State the evidence in one sentence |
| Ease | How quickly can we build, measure and interpret the test? | Fast path to a usable signal | Include measurement time, not only build time |
Score impact against the constraint, not against excitement
Impact is bounded by the current primary constraint. If activation is the constraint, a homepage test that improves acquisition may be worthwhile, but it does not deserve the same impact score as an intervention that directly improves activation.
| Impact score | Interpretation |
|---|---|
| 9–10 | Could shift the primary constraint by about 20% or more; rare |
| 6–8 | Meaningful movement in the constraint, or strong movement in a secondary metric |
| 3–5 | Modest movement in a supporting metric |
| 1–2 | Mainly changes a vanity metric |
For Clairo, only 34% of free sign-ups record a meeting. A pricing-page test may lift free-to-paid conversion, but it affects a thin downstream slice while activation remains leaky. Reducing onboarding friction deserves the higher impact score.
Earn confidence with evidence
Confidence is the most easily inflated score: liking an idea is not evidence for it. A useful ceiling rule is: if the team cannot name supporting evidence in one sentence, confidence cannot exceed 5.
| Confidence score | Evidence standard |
|---|---|
| 9–10 | Strong direct evidence, such as an earlier equivalent test in the same product/metric |
| 6–8 | Indirect but relevant evidence: comparable-company pattern plus own-user qualitative signal |
| 3–5 | Sound hypothesis without direct evidence |
| 1–2 | Guess prompted by brainstorming or fashion |
An earlier mobile-onboarding shortcut that lifted activation by 11% is direct evidence. A comparable SaaS case study showing a 15% lift after cutting onboarding steps, combined with own-user feedback that onboarding is long, is useful indirect evidence.
Ease is time to trustworthy learning
Ease includes building, instrumentation, waiting for data and analysis. A two-day build that needs six weeks to generate a credible signal is not a 9.
| Ease score | Typical total effort |
|---|---|
| 9–10 | Under a day; measurement already exists (copy, subject line, simple price change) |
| 6–8 | One to two weeks with light instrumentation |
| 3–5 | Three to six weeks; several teams, new metrics, referral mechanic or feature |
| 1–2 | Two months or more; major infrastructure/new platform bet |
Exam tip: An expensive feature can be valuable, but it often belongs in a product-bet roadmap rather than an experimental backlog.
ICE in use
At Clairo, shortening onboarding from four steps to two scores , , , hence . A fast homepage-hero test can have but only when it targets acquisition rather than the activation constraint, so it can rank below slower activation tests. ICE prevents the team from confusing speed with importance.
At Zoko, 62% of first-time starter-kit buyers do not repurchase within 90 days. Day 7, 14 and 21 WhatsApp progress-photo prompts directly address the observed habit/retention problem, whereas influencer content, TikTok partnerships and subscription toggles primarily target acquisition or revenue and should not crowd out the constraint.
Key takeaways
- ICE is a transparent prioritisation habit, not a precision machine.
- Score impact against the primary constraint, not a flattering adjacent metric.
- Confidence above 5 needs named evidence; ease includes time to usable learning.
- Rank the backlog, then let capacity decide how many high-ranking ideas can run.
The growth backlog: a system, not an idea graveyard
A list merely stores ideas. A growth backlog determines what runs, what stops and what the team learned. It needs stages, fields, owners, limits and a recurring operating ritual.
Five irreversible lifecycle stages
| Stage | Meaning |
|---|---|
| Idea | Thought recorded, but not yet evaluated |
| Scored | Impact, confidence and ease completed; rank is known |
| Active | Running now; exactly one owner and a written decision rule |
| Deciding | Time/sample bound reached; apply the rule |
| Closed | Scale/kill/iterate decision and reusable learning captured |
Every row must have a stage. An experiment that succeeds still closes after it is made permanent; a failed idea closes with a learning, not by returning silently to “idea”.
The eight essential fields
| Field | Why it exists |
|---|---|
| Experiment | One-line description of what changes |
| Hypothesis | “If we do X, Y will change because Z” — makes the causal claim falsifiable |
| Metric | One decisive number |
| ICE | Inputs and live average, not an unexplained final score |
| Stage | Shows operational state |
| Owner | One named person accountable for status |
| Decision rule | Scale, kill, iterate thresholds and time/sample bound, written before launch |
| Learning | Mandatory when closed; prevents repeat waste |
WIP limits and weekly cadence
Work in progress (WIP) is capped to protect signal quality and attention. Growth teams commonly use three to five simultaneous active experiments. More overlapping treatments make causal attribution difficult, diffuse ownership and reduce learning per experiment.
| Ritual | Purpose |
|---|---|
| Monday decision meeting | Apply completed rules; write learning; rescore non-active ideas on new evidence; fill vacant WIP slots from the ranking |
| Friday check-in | Each owner gives a concise status; flag an experiment that has reached a decision bound |
Confidence can be rescored as new evidence arrives, but not mid-test: an active experiment must run to its already-written rule.
Exam tip: “Team A and Team B” is not ownership. If every active row does not name one human, nobody is accountable.
Key takeaways
- A backlog has defined stages, a single owner, rules and learning — not just ideas.
- A closed experiment without a learning statement is not truly closed.
- WIP limits protect causality and focus; active rows are not rescored midstream.
- A lightweight Monday/Friday ritual keeps the system alive.
Decision rules: pre-commit before looking at results
A decision rule is a contract with the team’s future self. Without it, a middling result permits endless post-hoc debate. With it, every outcome becomes scale, kill or iterate.
| Decision | Meaning |
|---|---|
| Scale | Result reaches the success threshold; make the change permanent or expand it |
| Kill | Result is below the economic/learning hurdle; stop the mechanism and document why |
| Iterate | Signal lies between thresholds; run a controlled next version |
Four compulsory parts
- A single metric tied directly to the primary constraint.
- A scale threshold, derived from the hypothesised value of success.
- A kill threshold, below which the gain does not justify the cost.
- A bound: a fixed time window, sample size or both.
For Clairo, reducing onboarding from six steps to three uses the percentage of new free users who record a meeting within seven days. From a 34% baseline: scale at , kill at , iterate between those points, after 500 users per arm or 14 days. If the observed rate is 40.3%, the rule says iterate — regardless of who calls it “close enough”.
For Zoko, a day 7/14/21 photo-prompt sequence has a 22% 90-day repurchase baseline. Scale at , kill at , iterate between 24% and 30%, after 800 customers and the full 90-day observation period. A 23.1% result is a kill.
Iteration must be bounded
Iteration is a new experiment, not a permission slip to keep an idea alive.
- Change one variable only, so the cause of any change remains identifiable.
- Write a new rule and bound before the new version goes live.
- Allow at most two iterations; if version 3 still sits in the iterate band, force closure.
A successful kill captures the mechanism and lesson. Example: “The photo-prompt cadence did not move 90-day repurchase; future retention work should test another mechanism, such as ritual habit cards or referral incentives, before retesting cadence.” This raises the confidence of later ideas even though the test failed.
Four ways teams break the rules
- Moving the goalpost after seeing a result just below scale.
- Cohort cherry-picking until a noise segment appears significant; define the cohort first if it matters.
- The more-data trap: extending past the pre-set bound in hope of a different answer.
- Iterate by default despite landing below the kill threshold.
The audit question is simple: Did the rule exist in writing before activation? If not, the team is negotiating the outcome rather than experimenting.
Key takeaways
- Pre-register one metric, two thresholds and a bound.
- Scale, kill and iterate are the only legitimate result states.
- Iteration changes one variable and has a maximum of two rounds.
- Killing with a precise learning compounds knowledge; killing with “didn’t work” does not.
Validated metrics and the vanity-metric audit
A vanity metric is a true number that looks meaningful but neither changes nor predicts a business outcome. In a rigid decision rule, a wrong metric is more dangerous than no rule because it produces confident wrong decisions.
The five tests
| Test | A valid operating metric should… |
|---|---|
| Denominator | Be a rate where appropriate, not an unexplained cumulative count |
| Comparability | Be calculated consistently across periods/cohorts |
| Two-directional | Be able to rise or fall; cumulative totals usually only rise |
| Actionability | Tell the team what lever or next decision to consider when it moves |
| Business outcome | Trace to revenue, retention or margin in one or two validated hops |
Examples: replace total free sign-ups with weekly-cohort activation rate; total meetings recorded with percentage recording within seven days; a transcription-accuracy marketing claim with a customer-experience error rate; Instagram followers with Instagram-attributed first-purchase rate; and average order value with contribution margin per order.
The two-hop rule and leading indicators
A sound Clairo chain is:
The relationships must be supported by the team’s cohort data. By contrast, “Instagram followers reach visits purchases” has three assumed hops and is a story, not an operating metric.
A leading indicator is earned, not imported. It becomes valid when own cohorts show that movement in it predicts a later outcome and the relationship is revalidated after relevant product changes. Before that, it is a hypothesis and belongs in an experiment, not a decision rule.
Exam tip: Inputs — ad spend, tests run, tickets answered — measure effort. They are not automatically business outcomes.
Key takeaways
- Test every decision metric for denominator, comparability, directionality, actionability and outcome proximity.
- Prefer one or two validated hops to revenue, retention or margin.
- Engagement, cumulative investor-reporting numbers and gameable survey claims often become vanity metrics.
- A leading metric needs validation in own data before it directs decisions.
Capacity allocation and the 90-day roadmap
After ICE, capacity is the second filter. It has three currencies; plan against the one that becomes scarce first.
| Capacity currency | What constrains it |
|---|---|
| Engineering currency | Development and instrumentation hours |
| Decision currency | Founder/leadership attention needed to unblock, approve and coordinate |
| Audience currency | Number of users available without overlapping treatments destroying interpretability |
Allocate the binding resource rather than dividing all experiments equally:
| Allocation | Portfolio | Confidence guide |
|---|---|---|
| 70% | Known winners: improve/scale an already validated mechanism | |
| 20% | New bets with reasonable evidence | 5–7 |
| 10% | Exploration with high option value but low confidence |
Run experiments serially when they touch the same users, metric or mechanism; run them in parallel only when audiences, causal paths and the actual bottleneck permit clean interpretation. Review capacity biweekly and reallocate when strong evidence or a binding constraint changes; do not start everything simply because it appears on the roadmap.
A roadmap is a learning calendar
A 90-day growth roadmap converts the ranked backlog into sequenced execution. It should identify each experiment’s build/start/decision dates, owner, capacity commitment, dependency and slack/buffer. Biweekly review points are not status theatre: they are where the team applies rules, discovers constraints and reallocates responsibly.
The integrated growth engine
The AARRR funnel — acquisition, activation, retention, referral, revenue — is not a checklist of isolated tactics. It is a feedback system: activation behaviours should be chosen for their retention consequences; retained users can become referrers; referral cohorts may have different LTV and payback.
The six operating disciplines form one stack: ICE ranks ideas; the backlog holds their lifecycle; decision rules turn results into actions; validated metrics make those rules meaningful; capacity allocation limits commitment; the roadmap sequences execution. AARRR says where to work; the stack says how and when.
Quarter 1 calibrates instrumentation, cohorts and the binding currency. Quarter 2 turns Q1 learning into stronger hypotheses. Quarter 3 puts validated mechanisms into the 70% bucket. By Quarter 4, cross-stage learning can become a flywheel. Common failures are siloed workbook tabs, optimising only one funnel stage, failing to transfer closed-experiment learning to the next quarter, and maintaining a workbook that is updated but never used in decisions.
Key takeaways
- Capacity is engineering, decision attention and audience; the tightest one binds the plan.
- Use the 70/20/10 portfolio to balance scaling, evidence-building and options.
- Treat AARRR as a connected loop, not five separate workstreams.
- Growth compounds when each quarter’s closed experiments become the next quarter’s evidence.