Term 6 · Module 9 of 9

Experimentation System and Integrated Growth Engine

Business Research and Growth Systems Architecture

ICE prioritisation: choosing the experiment that deserves capacity

ICE prioritisation makes growth work deliberate when a team has far more ideas than it can run. It is not a claim that a score is perfectly objective; it is a shared discipline for deciding which idea should consume scarce time first.

ICE score=Impact+Confidence+Ease3\text{ICE score}=\frac{\text{Impact}+\text{Confidence}+\text{Ease}}{3}

Score every dimension from 1 to 10, then rank ideas in descending order. The usefulness lies less in the arithmetic than in making the evidence, constraint, effort and disagreement explicit.

DimensionQuestionHigh score meansDiscipline required
ImpactHow much will this move the metric that matters?Direct, material movement in the primary growth constraintDo not award a high score to a downstream or unrelated metric
ConfidenceHow sure are we that the predicted effect will occur?Specific, direct evidenceState the evidence in one sentence
EaseHow quickly can we build, measure and interpret the test?Fast path to a usable signalInclude measurement time, not only build time

Score impact against the constraint, not against excitement

Impact is bounded by the current primary constraint. If activation is the constraint, a homepage test that improves acquisition may be worthwhile, but it does not deserve the same impact score as an intervention that directly improves activation.

Impact scoreInterpretation
9–10Could shift the primary constraint by about 20% or more; rare
6–8Meaningful movement in the constraint, or strong movement in a secondary metric
3–5Modest movement in a supporting metric
1–2Mainly changes a vanity metric

For Clairo, only 34% of free sign-ups record a meeting. A pricing-page test may lift free-to-paid conversion, but it affects a thin downstream slice while activation remains leaky. Reducing onboarding friction deserves the higher impact score.

Earn confidence with evidence

Confidence is the most easily inflated score: liking an idea is not evidence for it. A useful ceiling rule is: if the team cannot name supporting evidence in one sentence, confidence cannot exceed 5.

Confidence scoreEvidence standard
9–10Strong direct evidence, such as an earlier equivalent test in the same product/metric
6–8Indirect but relevant evidence: comparable-company pattern plus own-user qualitative signal
3–5Sound hypothesis without direct evidence
1–2Guess prompted by brainstorming or fashion

An earlier mobile-onboarding shortcut that lifted activation by 11% is direct evidence. A comparable SaaS case study showing a 15% lift after cutting onboarding steps, combined with own-user feedback that onboarding is long, is useful indirect evidence.

Ease is time to trustworthy learning

Ease includes building, instrumentation, waiting for data and analysis. A two-day build that needs six weeks to generate a credible signal is not a 9.

Ease scoreTypical total effort
9–10Under a day; measurement already exists (copy, subject line, simple price change)
6–8One to two weeks with light instrumentation
3–5Three to six weeks; several teams, new metrics, referral mechanic or feature
1–2Two months or more; major infrastructure/new platform bet

Exam tip: An expensive feature can be valuable, but it often belongs in a product-bet roadmap rather than an experimental backlog.

ICE in use

At Clairo, shortening onboarding from four steps to two scores I=9I=9, C=8C=8, E=7E=7, hence 8.08.0. A fast homepage-hero test can have E=9E=9 but only I=5I=5 when it targets acquisition rather than the activation constraint, so it can rank below slower activation tests. ICE prevents the team from confusing speed with importance.

At Zoko, 62% of first-time starter-kit buyers do not repurchase within 90 days. Day 7, 14 and 21 WhatsApp progress-photo prompts directly address the observed habit/retention problem, whereas influencer content, TikTok partnerships and subscription toggles primarily target acquisition or revenue and should not crowd out the constraint.

Key takeaways

  • ICE is a transparent prioritisation habit, not a precision machine.
  • Score impact against the primary constraint, not a flattering adjacent metric.
  • Confidence above 5 needs named evidence; ease includes time to usable learning.
  • Rank the backlog, then let capacity decide how many high-ranking ideas can run.

The growth backlog: a system, not an idea graveyard

A list merely stores ideas. A growth backlog determines what runs, what stops and what the team learned. It needs stages, fields, owners, limits and a recurring operating ritual.

Five irreversible lifecycle stages

StageMeaning
IdeaThought recorded, but not yet evaluated
ScoredImpact, confidence and ease completed; rank is known
ActiveRunning now; exactly one owner and a written decision rule
DecidingTime/sample bound reached; apply the rule
ClosedScale/kill/iterate decision and reusable learning captured

Every row must have a stage. An experiment that succeeds still closes after it is made permanent; a failed idea closes with a learning, not by returning silently to “idea”.

The eight essential fields

FieldWhy it exists
ExperimentOne-line description of what changes
Hypothesis“If we do X, Y will change because Z” — makes the causal claim falsifiable
MetricOne decisive number
ICEInputs and live average, not an unexplained final score
StageShows operational state
OwnerOne named person accountable for status
Decision ruleScale, kill, iterate thresholds and time/sample bound, written before launch
LearningMandatory when closed; prevents repeat waste

WIP limits and weekly cadence

Work in progress (WIP) is capped to protect signal quality and attention. Growth teams commonly use three to five simultaneous active experiments. More overlapping treatments make causal attribution difficult, diffuse ownership and reduce learning per experiment.

RitualPurpose
Monday decision meetingApply completed rules; write learning; rescore non-active ideas on new evidence; fill vacant WIP slots from the ranking
Friday check-inEach owner gives a concise status; flag an experiment that has reached a decision bound

Confidence can be rescored as new evidence arrives, but not mid-test: an active experiment must run to its already-written rule.

Exam tip: “Team A and Team B” is not ownership. If every active row does not name one human, nobody is accountable.

Key takeaways

  • A backlog has defined stages, a single owner, rules and learning — not just ideas.
  • A closed experiment without a learning statement is not truly closed.
  • WIP limits protect causality and focus; active rows are not rescored midstream.
  • A lightweight Monday/Friday ritual keeps the system alive.

Decision rules: pre-commit before looking at results

A decision rule is a contract with the team’s future self. Without it, a middling result permits endless post-hoc debate. With it, every outcome becomes scale, kill or iterate.

DecisionMeaning
ScaleResult reaches the success threshold; make the change permanent or expand it
KillResult is below the economic/learning hurdle; stop the mechanism and document why
IterateSignal lies between thresholds; run a controlled next version

Four compulsory parts

  1. A single metric tied directly to the primary constraint.
  2. A scale threshold, derived from the hypothesised value of success.
  3. A kill threshold, below which the gain does not justify the cost.
  4. A bound: a fixed time window, sample size or both.

For Clairo, reducing onboarding from six steps to three uses the percentage of new free users who record a meeting within seven days. From a 34% baseline: scale at ≥42%\geq42\%, kill at ≤36%\leq36\%, iterate between those points, after 500 users per arm or 14 days. If the observed rate is 40.3%, the rule says iterate — regardless of who calls it “close enough”.

For Zoko, a day 7/14/21 photo-prompt sequence has a 22% 90-day repurchase baseline. Scale at ≥30%\geq30\%, kill at ≤24%\leq24\%, iterate between 24% and 30%, after 800 customers and the full 90-day observation period. A 23.1% result is a kill.

Iteration must be bounded

Iteration is a new experiment, not a permission slip to keep an idea alive.

  • Change one variable only, so the cause of any change remains identifiable.
  • Write a new rule and bound before the new version goes live.
  • Allow at most two iterations; if version 3 still sits in the iterate band, force closure.

A successful kill captures the mechanism and lesson. Example: “The photo-prompt cadence did not move 90-day repurchase; future retention work should test another mechanism, such as ritual habit cards or referral incentives, before retesting cadence.” This raises the confidence of later ideas even though the test failed.

Four ways teams break the rules

  • Moving the goalpost after seeing a result just below scale.
  • Cohort cherry-picking until a noise segment appears significant; define the cohort first if it matters.
  • The more-data trap: extending past the pre-set bound in hope of a different answer.
  • Iterate by default despite landing below the kill threshold.

The audit question is simple: Did the rule exist in writing before activation? If not, the team is negotiating the outcome rather than experimenting.

Key takeaways

  • Pre-register one metric, two thresholds and a bound.
  • Scale, kill and iterate are the only legitimate result states.
  • Iteration changes one variable and has a maximum of two rounds.
  • Killing with a precise learning compounds knowledge; killing with “didn’t work” does not.

Validated metrics and the vanity-metric audit

A vanity metric is a true number that looks meaningful but neither changes nor predicts a business outcome. In a rigid decision rule, a wrong metric is more dangerous than no rule because it produces confident wrong decisions.

The five tests

TestA valid operating metric should…
DenominatorBe a rate where appropriate, not an unexplained cumulative count
ComparabilityBe calculated consistently across periods/cohorts
Two-directionalBe able to rise or fall; cumulative totals usually only rise
ActionabilityTell the team what lever or next decision to consider when it moves
Business outcomeTrace to revenue, retention or margin in one or two validated hops

Examples: replace total free sign-ups with weekly-cohort activation rate; total meetings recorded with percentage recording within seven days; a transcription-accuracy marketing claim with a customer-experience error rate; Instagram followers with Instagram-attributed first-purchase rate; and average order value with contribution margin per order.

The two-hop rule and leading indicators

A sound Clairo chain is:

The relationships must be supported by the team’s cohort data. By contrast, “Instagram followers →\rightarrow reach →\rightarrow visits →\rightarrow purchases” has three assumed hops and is a story, not an operating metric.

A leading indicator is earned, not imported. It becomes valid when own cohorts show that movement in it predicts a later outcome and the relationship is revalidated after relevant product changes. Before that, it is a hypothesis and belongs in an experiment, not a decision rule.

Exam tip: Inputs — ad spend, tests run, tickets answered — measure effort. They are not automatically business outcomes.

Key takeaways

  • Test every decision metric for denominator, comparability, directionality, actionability and outcome proximity.
  • Prefer one or two validated hops to revenue, retention or margin.
  • Engagement, cumulative investor-reporting numbers and gameable survey claims often become vanity metrics.
  • A leading metric needs validation in own data before it directs decisions.

Capacity allocation and the 90-day roadmap

After ICE, capacity is the second filter. It has three currencies; plan against the one that becomes scarce first.

Capacity currencyWhat constrains it
Engineering currencyDevelopment and instrumentation hours
Decision currencyFounder/leadership attention needed to unblock, approve and coordinate
Audience currencyNumber of users available without overlapping treatments destroying interpretability

Allocate the binding resource rather than dividing all experiments equally:

AllocationPortfolioConfidence guide
70%Known winners: improve/scale an already validated mechanism≥7\geq7
20%New bets with reasonable evidence5–7
10%Exploration with high option value but low confidence<5<5

Run experiments serially when they touch the same users, metric or mechanism; run them in parallel only when audiences, causal paths and the actual bottleneck permit clean interpretation. Review capacity biweekly and reallocate when strong evidence or a binding constraint changes; do not start everything simply because it appears on the roadmap.

A roadmap is a learning calendar

A 90-day growth roadmap converts the ranked backlog into sequenced execution. It should identify each experiment’s build/start/decision dates, owner, capacity commitment, dependency and slack/buffer. Biweekly review points are not status theatre: they are where the team applies rules, discovers constraints and reallocates responsibly.

The integrated growth engine

The AARRR funnel — acquisition, activation, retention, referral, revenue — is not a checklist of isolated tactics. It is a feedback system: activation behaviours should be chosen for their retention consequences; retained users can become referrers; referral cohorts may have different LTV and payback.

The six operating disciplines form one stack: ICE ranks ideas; the backlog holds their lifecycle; decision rules turn results into actions; validated metrics make those rules meaningful; capacity allocation limits commitment; the roadmap sequences execution. AARRR says where to work; the stack says how and when.

Quarter 1 calibrates instrumentation, cohorts and the binding currency. Quarter 2 turns Q1 learning into stronger hypotheses. Quarter 3 puts validated mechanisms into the 70% bucket. By Quarter 4, cross-stage learning can become a flywheel. Common failures are siloed workbook tabs, optimising only one funnel stage, failing to transfer closed-experiment learning to the next quarter, and maintaining a workbook that is updated but never used in decisions.

Key takeaways

  • Capacity is engineering, decision attention and audience; the tightest one binds the plan.
  • Use the 70/20/10 portfolio to balance scaling, evidence-building and options.
  • Treat AARRR as a connected loop, not five separate workstreams.
  • Growth compounds when each quarter’s closed experiments become the next quarter’s evidence.