Term 6 · Module 3 of 9

AB Testing and Startup experimentation Math

Business Research and Growth Systems Architecture

Null Hypothesis

The null hypothesis (H₀) is the assumption of no effect — the default position that any observed difference between variants is due to random chance alone. It is not a belief you hold; it is a disciplined starting point that protects you from confirmation bias. Every A/B test is designed to reject this assumption, thereby providing evidence for the alternative hypothesis (H₁) — the claim you actually want to prove.

Definition: The null hypothesis states that any change you make (e.g., a new CTA) produces zero difference in the outcome metric.

Hunch vs. Hypothesis vs. Null Hypothesis

Only two of these three are testable. The table summarises the distinctions:

ConceptDefinitionExampleTestable?
HunchA gut feeling with no metric, timeframe, or measurable prediction“I think the green button will get more clicks.”❌ Not testable
Hypothesis (H₁)A specific, measurable, and falsifiable prediction“Changing CTA to ‘Record my first meeting’ will increase free-to-paid conversion by ≥10% within 30 days.”✅ Testable
Null Hypothesis (H₀)The assumption that the change produces no difference“Changing the CTA will produce no difference in sign‑up rate.”✅ Testable (you attempt to reject it)

A good hypothesis must be falsifiable — there must be a measurable way to show it is wrong. A hunch lacks this property and cannot be tested scientifically.

The Four-Step Testing Process

Key nuance: “Failing to reject H₀” is not a failure. It means your test did not find strong enough evidence to rule out random chance — it does not prove your variant is worthless. The difference might become detectable with a larger sample.

Why Every Experiment Starts from “No Difference”

Starting from the null hypothesis forces you to:

  • Set a significance threshold before looking at data.
  • Run the test to completion (never stop early).
  • Interpret results objectively, not as confirmation of your hunch.

If you instead start from optimism (assuming your variant is better), you will:

  • Check results daily and stop as soon as numbers favour your variant.
  • Fall prey to confirmation bias — seeing what you want to see.
  • Risk costly decisions (e.g., rolling out a change based on a few days’ data that may vanish later).

Exam tip: Confirmation bias is the #1 threat in A/B testing. The null hypothesis is the mechanism that guards against it.

Real Brand Examples

Clairo (B2B)

  • Baseline: free‑to‑paid conversion = 3.7%
  • H₁: Changing CTA to “Record my first meeting” will improve conversion to ≥4.1% (a 10% relative lift).
  • H₀ : Changing the CTA produces no difference in free‑to‑paid conversion.

Zoko (B2C)

  • Baseline: starter kit attach rate = 38%
  • H₁: Changing CTA to “Try the 21‑Day Starter Kit” will increase the rate to ≥46% (an 8‑percentage‑point absolute lift, ~20% relative).
  • H₀ : Changing the CTA produces no difference in attach rate.

In both cases, the test is designed to disprove H₀, not to “prove” the variant is better.

Important Truths About the Null Hypothesis

  1. You can never prove H₀ true — you can only fail to reject it.
  2. Rejecting H₀ is not the same as proving H₁ — a significant result supports H₁ but does not guarantee it with 100% certainty. Statistics always leaves room for uncertainty.
  3. H₀ protects your business from expensive mistakes — it turns a gut feeling into a rigorous test, preventing overconfident decisions based on small or biased samples.

Key takeaways

  • The null hypothesis assumes no effect and is the starting point of every A/B test.
  • A hypothesis (H₁) must be specific, measurable, and falsifiable.
  • A hunch is not testable — it lacks metrics and a timeframe.
  • The goal of a test is to reject H₀ (statistical significance); failing to reject H₀ does not mean the variant is ineffective.
  • Starting from H₀ guards against confirmation bias and forces honest, data‑driven decisions.

Sample Size Logic Without Heavy Statistics

Sample size is the minimum number of observations you need before an A/B test result can be trusted. Without it, you are not running a test — you are just making an observation. The logic: a coin flipped 3 times giving 3 heads does not prove the coin is rigged; flip it 1,000 times and get 700 heads, and you can say a real pattern exists. Sample size is the number of “flips” (users, events) your test needs.

Three inputs every sample size calculation needs

You cannot run a valid test without deciding all three before looking at any data.

InputWhat it isExample (Clairo)Example (Zoko)
Baseline conversion rateCurrent performance of the control3.7% (free→paid)38% (starter kit attach rate)
Minimum detectable effect (MDE)Smallest improvement that would change a business decision10% relative lift8 percentage points absolute (≈21% relative)
Significance level (α\alpha) and power (1−β1-\beta)Risk tolerance for errors0.05 (95% confidence) and 80% powerSame defaults

Exam tip: B2B (Clairo) typically has lower baseline rates than B2C (Zoko). This difference affects sample size directly — lower baselines need more data.

Significance level (α\alpha) and power explained

α\alpha — false positive protection

  • Plain language definition: The probability that your test will declare a winner when there is no real difference.
  • Standard default: α=0.05\alpha = 0.05 (95% confidence). Reason: proposed by Ronald Fisher in 1925 as a pragmatic convention. 1 in 20 tests will show a false positive by chance.
  • Lowering α\alpha (e.g., to 0.01) increases required sample size. For Clairo (3.7% baseline, 10% MDE): at α=0.05\alpha=0.05 → ~1,400 per variant; at α=0.01\alpha=0.01 → ~2,100 per variant.
  • Pharma/finance may use 99% confidence; for startup experimentation, leave α=0.05\alpha=0.05 as default unless a senior statistician documents a reason to change.

Power ( 1−β1-\beta ) — false negative protection

  • Plain language definition: The probability that your test will detect a real effect when one genuinely exists.
  • Standard default: 80% power (proposed by Jacob Cohen, 1969). Means: if variant B actually works at the MDE, the test will find it 4 out of 5 times. 1 in 5 real effects will be missed.
  • Higher power (e.g., 90%) increases sample size. For Clairo: 80% → ~1,400; 90% → ~1,870 per variant.
  • 80% is the industry default for marketers. Increasing power costs more data and time.

Connection between α\alpha, power, and sample size: Both control different error types. Adjusting either one changes the required sample size. More perfection → more data.

Minimum detectable effect (MDE) — the business-driven input

MDE sets the smallest lift that would change your decision. It is not what is easiest to detect; it is what matters.

MDE sizeImplicationExample sample size
Tiny (1% absolute)Requires hundreds of thousands of users per variant; months to years of trafficImpractical for startups
Practical (10-20% relative)Sample size in hundreds or low thousands; feasible in days/weeksClairo’s 10% MDE → 9,500 per variant on free-to-paid conversion (impractical)
Large (e.g., 20% relative)Much smaller sample size; blinds test to any smaller effect20% MDE from 34% baseline → ~770 per variant (but misses a 12% lift)

Rule: Set MDE based on the minimum improvement that would change your action. If you wouldn’t update pricing unless a CTA lifts conversion by at least 10%, then MDE = 10%.

Worked example: Clairo vs. Zoko (live calculator outputs)

Evan Miller’s sample size calculator (raw number per variant)

  • Clairo (activation rate proxy): Baseline 34%, MDE 10% relative, α=0.05\alpha=0.05, 80% power → ~1,400 per variant.
  • Zoko: Baseline 38%, MDE 8% absolute (to 46%), same defaults → ~560 per variant.

Exam tip: Use relative MDE when baseline is small; absolute MDE is often clearer for higher baselines. Always specify which type you use.

VWO test duration calculator (adds traffic)

  • Clairo (free→paid conversion): Baseline 3.7%, MDE 10% relative → needs ~9,500 per variant. With 620 sign-ups/month (≈21 visitors/day) → test would take 15 months — likely not runnable.
  • Clairo (activation rate): Baseline 34%, MDE 10% relative → needs ~1,400 per variant. With 620 sign-ups/month → test completes in ~2 months — runnable.
  • Zoko: Baseline 38%, MDE 8% absolute → needs ~560 per variant. With 38,000 monthly visitors → sample reached in <1 day. But never stop before 7 calendar days to capture day-of-week patterns.

Three mistakes that break sample size logic

MistakeWhat happensFix
Peeking – checking results before sample is reachedInflates false positive rate; early lucky streak looks like a winnerSet sample size in advance; do not open results until it is met
Setting MDE too lowNeed hundreds of thousands of users → test impossibleSet MDE based on minimum decision-changing lift (10–20% relative is practical)
Forgetting sample size is per variantOnly plan for total, not per arm; A/B test needs ×2, A/B/C needs ×3Always double/triple according to number of variants

Connecting sample size to experiment modes (template)

Teams record the calculation in a template with columns:

  1. Hypothesis & metric (e.g., H1: changing CTA lifts free→paid from 3.7% to ≥4.07%)
  2. Experiment mode (full valid / directional / proxy)
  3. Sample target vs. actual (e.g., target 1,400 per variation)
  4. Result & decision (e.g., “101% reached → conclusive” or “39% reached → directional, flag as inconclusive”)

The target column is the output of the calculator; the mode column forces the question: is this test runnable?

Key takeaways

  • Sample size separates signal from luck; it is the minimum observations needed before you can trust the result.
  • Three pre‑test inputs: baseline conversion rate, MDE (business‑driven), and α/power (use defaults: 0.05 and 80%).
  • MDE must be set on business logic, not statistical convenience. Practical MDEs are 10–20% relative.
  • Always run the test for at least 7 calendar days even if the sample size is reached earlier.
  • Peeking, too‑low MDE, and forgetting “per variant” are the three common pitfalls.

Type I and Type II Errors

Every A/B test can produce a wrong answer in exactly two ways. The core idea: You are comparing a null hypothesis (no difference) against a real effect. A test can say “there is a difference” when there isn’t (false positive, Type I error) or say “no difference” when there really is one (false negative, Type II error). Understanding these two failure modes lets you control them — and prevents costly misread data.


The 2×2 Matrix of Test Outcomes

The columns represent reality (H₀ true or false). The rows represent your decision (reject H₀ or fail to reject H₀). Four cells, two of which are errors.

H₀ true (no real effect)H₀ false (real effect exists)
Reject H₀ (say “effect found”)Type I error (false positive) – wrong convictionCorrect rejection – true positive
Fail to reject H₀ (say “no effect”)Correct non-rejection – true negativeType II error (false negative) – wrong acquittal

Exam tip: Type I = false positive; Type II = false negative. Mnemonic: “I” sounds like “lie” (positive lie); “II” sounds like “nein” (negative lie).


Definitions in Plain Terms

  • Type I error: You think the variant works when it actually does not. Analogy: Convicting an innocent person. The evidence pointed to guilt, but the person didn’t do it.

  • Type II error: You think the variant does not work when it actually does. Analogy: Acquitting a guilty person. Evidence was insufficient, but they really did it.


Key Quantities: Alpha and Beta

SymbolNameMeaningTypical valueControls
α\alphaSignificance levelProbability of Type I error (false positive)0.05 (5%)Set before test
β\beta–Probability of Type II error (false negative)0.20 (20%)Reduced by larger sample
1−β1-\betaStatistical powerProbability of detecting a real effect of size MDE0.80Increased by larger sample

If you set α=0.05\alpha = 0.05, then 5% of tests where H₀ is true will still produce a false positive by chance alone — 1 in 20.

If power = 80%, then 20% of the time a real effect of the chosen MDE will not be detected.


Causes & Business Costs

Type I Error

  • Most common cause: Peeking — stopping the test too early with very little traffic. A small early fluctuation looks like a signal.
  • Business cost: You scale a change that does not actually improve performance. Resources (time, money, effort) are wasted on a wrong direction. False beliefs about what drives conversion become embedded.

Example – Clairo (CTA copy test): Team saw 4.2% conversion vs 4.0% after only 100 users and declared B the winner. Required sample was 9,500 per variant. The 0.2-point lift was random variation. Switching to the wrong CTA could drop conversion from 3.7% to under 3%, costing millions.

Example – Zoko (starter kit CTA test): High traffic reached sample size in 18 hours (weekday afternoon). Observed a 9-point lift, but the test should have run a full 7-day week. The real effect might have been 2 points or the opposite.

Type II Error

  • Most common cause: Underpowered test — sample size too small. Or setting MDE too high (e.g., expecting a 20% lift when the true lift is 9%).
  • Business cost: You miss an opportunity. The better variant is discarded, and the worse baseline continues. Unlike Type I, this loss is invisible — you never see what you could have gained. Over time, the missed benefit compounds.

Example – Clairo (onboarding flow test): Required 1,400 users per variant, but budget constraints limited it to 310. Test showed no significant result; team abandoned a genuinely better onboarding flow.

Example – Zoko (unboxing insert test): MDE set at 20% relative change; true effect was 9% improvement in 90-day repurchase. Underpowered test missed it.


How to Control Each Error

ErrorControlled byMechanismSide effect
Type ILower α\alpha (significance level)Reduce the threshold for “statistical significance”Increases required sample size; may raise Type II risk if power drops
Type IIIncrease power (via larger sample size)More data makes real effects easier to detectMore time, traffic, and resources needed

Both are design decisions made before the test runs.


When to Worry More About Which Error

Prioritize controlling Type I when:

  • Scaling the change requires a major resource commitment (e.g., paid acquisition campaign, engineering sprint, pricing change) — hard to reverse.
  • The market is slow-moving; acting on noise can waste months of momentum before the mistake is visible.

Prioritize controlling Type II when:

  • Test opportunities are scarce (e.g., low traffic, few chances to rerun). Missing one variant is a big loss.
  • You are in a fast competitive market; even a small improvement can give an edge. Missing it may let a competitor capture that advantage first.

Exam tip: There is no universal “worst error” — it depends on context. Type I bites when scaling is costly; Type II bites when opportunities are rare.


Key Takeaways

  • Type I error (false positive): reject H₀ when H₀ is true. Caused by peeking, small samples. Cost: wasted resources on a false winner.
  • Type II error (false negative): fail to reject H₀ when H₀ is false. Caused by underpowered tests, MDE too high. Cost: lost opportunity, invisible.
  • The 2×2 matrix maps both errors; two cells are errors, two are correct outcomes.
  • α\alpha controls Type I; power (via sample size) controls Type II.
  • Which error is more damaging depends on context: costly scaling → Type I; scarce/fast opportunities → Type II.

Startup Experimentation Math

Real-world startups face limited traffic, time, and patience — ideal statistical conditions (large samples, no peeking, fixed duration) are often impossible. The goal is informed trade-offs, not shortcuts: make calculated compromises without abandoning statistical thinking.

The Speed–Accuracy Trade-off

Any experiment exists on a spectrum from maximum speed (no controls) to maximum accuracy (full statistical validity). The choice is a design decision based on business context, not a mistake.

  • Ship & watch (left): make a change, watch aggregate metric. No hypothesis, no control — cannot isolate cause.
  • Directional testing (middle): run with a hypothesis but a sample below the required size. Gives trend, not proof.
  • Full valid test (right): pre-calculated sample, run to completion, no peeking. Produces conclusive, defensible results.

Three Experiment Modes

ModeWhen to useExampleAcceptable conclusion
Full valid testDecision is costly, hard to reverse, high stakesPricing change, new acquisition channel, major product updateSignificant (or not) at pre-set confidence level → scale or abandon
Directional testDecision has low cost, reversible, fast iterationEmail copy, secondary CTA, onboarding messageDirection only — do not use to justify scaling
Proxy testFinal metric needs too much traffic/time; measure a leading indicator insteadFree→paid conversion too slow → test activation rateRun a full valid test on the proxy; not a permanent substitute

Choosing the mode is a business decision: If this test gives a wrong result, how big is the impact? How reversible?

Partial Sample Results: What You Actually Know

Running a test on less than 100% of the required sample degrades effective power and inflates Type I error. The table below shows the practical meaning:

% of required sampleEffective powerType I error rateWhat you know
20%~30%15–20%Almost nothing — as good as no test
40%~50%10–12%Weak direction, highly uncertain
60–80%~60–70%7–8%Signal becomes more reliable, still not conclusive
100%80%5%Valid, actionable result

Exam tip: Never use a partial-sample result to justify scaling. Only a full valid test (100% sample, 80% power, 5% Type I) produces a decision. Anything else is a directional signal — useful for planning the next full test, not for rolling out changes.

Proxy Metrics

A proxy metric is a leading indicator that reliably predicts the final (lagging) metric and can be measured faster and with less data.

Leading vs. Lagging Metric (example: Clairo)

  • North star metric (lagging): number of paid users.
  • Leading metrics (can be improved to affect the north star):
    • Visitors → free user conversion rate
    • Free user → paid user conversion rate
    • Number of visitors (via campaigns)

To increase paid users, you can improve any leading metric. A proxy test chooses one leading metric as a stand-in.

Three Requirements for a Valid Proxy

  1. Correlation — must be documented (e.g., 10% of free users become paid → free users strongly predict paid users).
  2. Sensitivity — must respond to the change you are testing (e.g., an activation-rate change will affect free→paid conversion).
  3. Speed advantage — must genuinely require less traffic/time than testing the final metric directly.

Real Example: Clairo

  • End metric: free→paid conversion rate (3.7% → 4.1%)
  • Required sample: ~19,000 users → 15 months at 620 sign-ups/month ❌
  • Proxy chosen after cross-team discussion: Activation rate (user reaches aha moment within 7 days).
  • Proxy sample: ~2,800 users (1,400 per variant) → testable in 2–3 months ✅
  • Validation: users who never activate never convert; activation directly correlates with free→paid.
  • Mode for proxy test: full valid test on the proxy (no compromise on accuracy for this high-stakes decision).

Exam tip: Proxy tests are speed tools. As soon as you have enough traffic to test the final metric directly, switch back to it.

Experiment Decision Log

A 4-column log keeps the experimentation system honest. Record before results are known.

ColumnContentPurpose
1Hypothesis + Metric“I will increase activation rate from 34% to 39%.” Prevents vague descriptions like “improve customer journey.”
2ModeFull valid / Directional / Proxy (chosen before seeing data). Forces team to agree on acceptable conclusion type.
3Sample target vs. actualRequired sample size (from calculator) and actual sample reached. The ratio instantly tells you if result is valid or just directional.
4Decision + Acknowledged riskFor mode 1: scale decision. For mode 2: state that risk is elevated and result is directional only — do not treat as proof.

This log is used throughout the course to document experiments consistently.


Key takeaways

  • The speed–accuracy trade-off is a deliberate business decision, not a statistical error.
  • Three modes: full valid (high stakes), directional (low cost/reversible), proxy (when final metric is too slow).
  • Partial samples ( < 100%) give only directional signals — never use them to justify scaling.
  • A valid proxy must be correlated, sensitive, and faster to measure; test it with full validity.
  • Always log every experiment with hypothesis, mode, sample ratio, and acknowledged risk.

Practical Landing Page A/B Cases – Clairo & Zoko

Two complete A/B tests are run end-to-end to show the full 6‑step workflow in action. One test (Clairo, B2B) produces a significant result; the other (Zoko, B2C) does not. Both are valid, designed tests – the difference lies only in the outcome and the decision that follows.

The 6‑Step Workflow (Applied Identically to Both Cases)

  1. State hypotheses – Define H0H_0 (null: no difference) and H1H_1 (alternative: expected effect) before any data is collected.
  2. Calculate required sample size – Based on baseline rate, minimum detectable effect (MDE), α=0.05\alpha = 0.05, power = 80%.
  3. Choose experiment mode & decision rule – Select mode (full valid test, directional, etc.) and pre‑commit to the rule: if p<αp < \alpha at full sample → deploy; otherwise don’t.
  4. Run test – No peeking; reach the required sample; run at least 7 calendar days.
  5. Read results – Check pp vs α\alpha and actual sample vs required. Apply error‑type framework.
  6. Log decisions – Hypothesis, mode, sample ratio, final decision, risk acknowledged.

Exam tip: Steps 1–3 must be written down before the test starts. Peeking invalidates the error guarantees.


Case A: Clairo (B2B Startup)

Goal: Increase 7‑day activation rate from 34% to at least 37.4% (10% relative lift) using a new guided first‑meeting flow.

Step 1 – Hypotheses

  • H0H_0: New onboarding flow produces no difference in 7‑day activation rate.
  • H1H_1: New flow increases 7‑day activation rate from 34% to at least 37.4%.

Step 2 – Sample Size

  • Baseline: 34%
  • MDE: 10% relative (3.4% absolute).
  • α=0.05\alpha = 0.05, power = 80%
  • Required: ~1,400 users per variant → total 2,800 users.

Clairo gets 620 sign‑ups/month → 2.3 months needed; achievable within a single quarter.

Step 3 – Mode & Decision Rule

  • Mode: Full valid test (metric is the primary growth constraint; result drives next quarter’s roadmap).
  • Decision rule:
    • If p<0.05p < 0.05 at full sample → deploy new flow.
    • If p≥0.05p \geq 0.05 → do not scale; conduct UI audit / redesign and retest.

Steps 4–6 – Results & Decision

  • Ran 70 days, reached 1,447 users per variant (103% of required). No peeking.
  • Control: 34% activation
  • Variant: 38.4% activation
  • Lift: +4.4 p.p. (12.9% relative)
  • p=0.031p = 0.031 (confidence 96.9%) → p<0.05p < 0.05 → significant.

Decision: Deploy the new flow. Log all details (hypotheses, mode, sample ratio, decision, risk acknowledged).

Key point: A properly validated proxy metric (7‑day activation predicting paid conversion) can be a legitimate test metric if the correlation is documented.

Key Takeaways – Clairo

  • Hypotheses must be specific (baseline → target) and stated pre‑test.
  • Sample size is derived from baseline, MDE, α\alpha, and power; ensure traffic can support it.
  • A significant result (p<αp < \alpha) allows deployment, but still monitor the actual downstream metric.

Case B: Zoko (B2C Brand)

Goal: Increase starter‑kit attach rate from 38% to at least 46% (8 p.p. absolute lift) by changing the CTA from “Start your skin journey” to “Try the 21‑day starter kit”.

Step 1 – Hypotheses

  • H0H_0: CTA change produces no difference in attach rate.
  • H1H_1: New CTA increases attach rate from 38% to at least 46%.

Step 2 – Sample Size

  • Baseline: 38%
  • MDE: 8% absolute (relative ~21%).
  • α=0.05\alpha = 0.05, power = 80%
  • Required: ~560 users per variant → total 1,120 users.

Zoko receives 38,000 visitors/month → can reach in <1 day, but test is still run for 7 calendar days to capture day‑of‑week variation.

Step 3 – Mode & Decision Rule

  • Mode: Full valid test.
  • Decision rule:
    • If p<0.05p < 0.05 at 7‑day end → deploy new CTA.
    • If p≥0.05p \geq 0.05 → do not deploy; iterate copy/design and retest.

Steps 4–6 – Results & Decision

  • Ran 7 days, reached 572 users per variant (102% of required).
  • Control: 38% attach rate
  • Variant: 40.3% attach rate
  • Lift: +2.3 p.p. (absolute)
  • p=0.21p = 0.21 (confidence 79%) → p>0.05p > 0.05 → not significant.

Decision: Fail to reject H0H_0. Do not deploy. But do not abandon – the observed 2.3% lift is directional; the copy may have a small real effect that a larger sample or a lower MDE could detect. The next step is to iterate (better copy, different design) and retest.

Interpreting a Non‑Significant Result – The Critical Lesson

  • ✅ The result is evidence that the specific CTA did NOT produce the 8% lift we designed for.
  • ✅ It is a valid, complete test at the designed error rates.
  • ✅ It provides directional data (variant performed 2.3% better than control).
  • ❌ It does not prove the CTA has zero effect.
  • ❌ It does not prove the variant is worse than control (it’s better).
  • ❌ It does not give permission to stop testing the hypothesis.

Exam tip: A non‑significant result does not mean “nothing happened”. It means the observed effect is smaller than the MDE the test was powered to detect. The correct response: iterate, don’t conclude.


Comparison Table: Clairo vs. Zoko

DimensionClairo (B2B)Zoko (B2C)
Metric7‑day activation rate (proxy for paid conversion)Starter‑kit attach rate (direct)
Baseline34%38%
Target lift+3.4 p.p. (10% relative)+8 p.p. (absolute)
Required sample/variant1,400560
Actual sample/variant1,447 (103%)572 (102%)
Observed lift+4.4 p.p. (12.9% relative)+2.3 p.p. (absolute)
pp‑value0.0310.21
Decision on H0H_0Reject (significant)Cannot reject (not significant)
Business actionDeploy new flowDo not deploy; iterate & retest
LessonA validated proxy metric can be trusted if correlation is documented.Observed effect < designed MDE → directional signal, not conclusion.

Key Takeaways – Practical A/B Cases

  • The 6‑step workflow is the same for every test – only the numbers and the outcome differ.
  • A significant result (p<αp < \alpha) permits deployment; a non‑significant result signals iteration, not abandonment.
  • Never peek, always reach the pre‑calculated sample, and run at least 7 days.
  • A non‑significant test does not confirm the null; it only fails to reject it. The observed effect is a directional input for the next experiment.
  • Proxy metrics are acceptable only if their correlation to the true business metric is validated and documented.
  • Growth is iterative – no single test gives the final answer.

Practical Landing Page A/B Case – II

This case walks through reading real A/B test results from VWO, interpreting significance, making deploy decisions, and logging experiments across multiple runs. The standard reading order is: sample → rates → statistics.


Clairo Activation Test – Significant Result

MetricControl (A)Variant (B)Difference
Sample per variant1,4471,447–
Target per variant1,4001,400103% reached
7‑day activation rate34.0%38.4%+4.4 pp absolute
Relative lift––+12.9% ((38.4−34)/34)((38.4-34)/34)
p-value––0.031
Confidence level––96.9%
  • Sample: 2,894 total visitors; full sample reached (103% of target).
  • Rates: Variant outperforms control by 4.4 percentage points.
  • Statistics: p=0.031<0.05p = 0.031 < 0.05, confidence 96.9%>95%96.9\% > 95\% → statistically significant.
  • Decision: Reject the null hypothesis (H₀) – the new onboarding flow produces a real improvement. Deploy the variant across the site.

Exam tip: Significance alone is not enough. Always verify that the target sample size was reached (as designed for the chosen alpha and power). Here both conditions are met.


Zoko CTA Test – Non‑Significant Result

MetricControl (A)Variant (B)Difference
Sample per variant572572–
Target per variant560560102% reached
Attach rate38.0%40.3%+2.3 pp absolute
Relative lift––+6.1% (approx.)
p-value––0.21
Confidence level––79%
  • Sample: Full sample reached (102% of target).
  • Rates: Observed absolute lift of 2.3 percentage points.
  • Statistics: p=0.21>0.05p = 0.21 > 0.05, confidence 79%<95%79\% < 95\% → not statistically significant.
  • Decision: Fail to reject H₀. Do not deploy the new CTA copy.

What a Non‑Significant Result Actually Means

  • The test provides evidence that the new copy did not produce the designed minimum detectable effect (MDE) (here, an 8% relative lift).
  • It does not prove the copy has zero effect – the observed +2.3 pp lift suggests a real effect smaller than the MDE.
  • Therefore, the hypothesis is retained, and the team should iterate (e.g., refine the copy) rather than abandon the idea.

Directional Tests vs Full Valid Tests

  • Directional test: Run when insufficient traffic prevents reaching the target sample. Results are inconclusive but can be logged to preserve the hypothesis for a future full run.
  • Full valid test: Reaches the pre‑specified sample size and satisfies the design parameters (alpha, power). Only this mode supports statistical conclusions.

Example from the Experiment Log (Zoko CTA)

RowTestModeSample Reached% of TargetResultDecision
2Zoko starter kit CTADirectional220 per variant39%InconclusiveHypothesis retained
5Zoko starter kit CTAFull valid572 per variant102%Not significant (p=0.21)Do not deploy; hypothesis retained

The log captures both runs – the same hypothesis tested twice at increasing efficiency. This is the experimentation life cycle in action.


Experiment Logging – The Three‑Act Story

The experiment log tracks the history of each hypothesis across runs:

  1. Row 2 (directional): Zoko CTA hypothesis tested with insufficient sample → inconclusive → hypothesis retained.
  2. Row 5 (full valid): Same hypothesis tested with full sample → still non‑significant → hypothesis retained again, but with clear evidence that the effect is below the MDE.
  3. Row 1 (Clairo): Full valid test, significant → deployed.

The log provides continuity and informs the next iteration. It is not a contradiction to log two entries for the same hypothesis – that is the system working as designed.

Exam tip: A non‑significant result from a full valid test is a valid and complete outcome. It means the test was properly powered but the effect was not large enough to detect. Use it as input for the next iteration – do not treat it as failure.

Key Takeaways

  • Read results in order: sample → rates → statistics.
  • Clairo: p=0.031, full sample → significant → deploy.
  • Zoko: p=0.21, full sample → not significant → do not deploy, but retain hypothesis.
  • Directional tests are valid for learning; only full valid tests support statistical conclusions.
  • Experiment logs capture the lifecycle of a hypothesis across multiple runs.
  • Non‑significance does not imply zero effect – only that the observed effect is smaller than the MDE.
  • Systematic iteration, not one‑shot wins, drives growth.