Null Hypothesis
The null hypothesis (H₀) is the assumption of no effect — the default position that any observed difference between variants is due to random chance alone. It is not a belief you hold; it is a disciplined starting point that protects you from confirmation bias. Every A/B test is designed to reject this assumption, thereby providing evidence for the alternative hypothesis (H₁) — the claim you actually want to prove.
Definition: The null hypothesis states that any change you make (e.g., a new CTA) produces zero difference in the outcome metric.
Hunch vs. Hypothesis vs. Null Hypothesis
Only two of these three are testable. The table summarises the distinctions:
| Concept | Definition | Example | Testable? |
|---|---|---|---|
| Hunch | A gut feeling with no metric, timeframe, or measurable prediction | “I think the green button will get more clicks.” | ❌ Not testable |
| Hypothesis (H₁) | A specific, measurable, and falsifiable prediction | “Changing CTA to ‘Record my first meeting’ will increase free-to-paid conversion by ≥10% within 30 days.” | ✅ Testable |
| Null Hypothesis (H₀) | The assumption that the change produces no difference | “Changing the CTA will produce no difference in sign‑up rate.” | ✅ Testable (you attempt to reject it) |
A good hypothesis must be falsifiable — there must be a measurable way to show it is wrong. A hunch lacks this property and cannot be tested scientifically.
The Four-Step Testing Process
Key nuance: “Failing to reject H₀” is not a failure. It means your test did not find strong enough evidence to rule out random chance — it does not prove your variant is worthless. The difference might become detectable with a larger sample.
Why Every Experiment Starts from “No Difference”
Starting from the null hypothesis forces you to:
- Set a significance threshold before looking at data.
- Run the test to completion (never stop early).
- Interpret results objectively, not as confirmation of your hunch.
If you instead start from optimism (assuming your variant is better), you will:
- Check results daily and stop as soon as numbers favour your variant.
- Fall prey to confirmation bias — seeing what you want to see.
- Risk costly decisions (e.g., rolling out a change based on a few days’ data that may vanish later).
Exam tip: Confirmation bias is the #1 threat in A/B testing. The null hypothesis is the mechanism that guards against it.
Real Brand Examples
Clairo (B2B)
- Baseline: free‑to‑paid conversion = 3.7%
- H₁: Changing CTA to “Record my first meeting” will improve conversion to ≥4.1% (a 10% relative lift).
- H₀ : Changing the CTA produces no difference in free‑to‑paid conversion.
Zoko (B2C)
- Baseline: starter kit attach rate = 38%
- H₁: Changing CTA to “Try the 21‑Day Starter Kit” will increase the rate to ≥46% (an 8‑percentage‑point absolute lift, ~20% relative).
- H₀ : Changing the CTA produces no difference in attach rate.
In both cases, the test is designed to disprove H₀, not to “prove” the variant is better.
Important Truths About the Null Hypothesis
- You can never prove H₀ true — you can only fail to reject it.
- Rejecting H₀ is not the same as proving H₁ — a significant result supports H₁ but does not guarantee it with 100% certainty. Statistics always leaves room for uncertainty.
- H₀ protects your business from expensive mistakes — it turns a gut feeling into a rigorous test, preventing overconfident decisions based on small or biased samples.
Key takeaways
- The null hypothesis assumes no effect and is the starting point of every A/B test.
- A hypothesis (H₁) must be specific, measurable, and falsifiable.
- A hunch is not testable — it lacks metrics and a timeframe.
- The goal of a test is to reject H₀ (statistical significance); failing to reject H₀ does not mean the variant is ineffective.
- Starting from H₀ guards against confirmation bias and forces honest, data‑driven decisions.
Sample Size Logic Without Heavy Statistics
Sample size is the minimum number of observations you need before an A/B test result can be trusted. Without it, you are not running a test — you are just making an observation. The logic: a coin flipped 3 times giving 3 heads does not prove the coin is rigged; flip it 1,000 times and get 700 heads, and you can say a real pattern exists. Sample size is the number of “flips” (users, events) your test needs.
Three inputs every sample size calculation needs
You cannot run a valid test without deciding all three before looking at any data.
| Input | What it is | Example (Clairo) | Example (Zoko) |
|---|---|---|---|
| Baseline conversion rate | Current performance of the control | 3.7% (free→paid) | 38% (starter kit attach rate) |
| Minimum detectable effect (MDE) | Smallest improvement that would change a business decision | 10% relative lift | 8 percentage points absolute (≈21% relative) |
| Significance level () and power () | Risk tolerance for errors | 0.05 (95% confidence) and 80% power | Same defaults |
Exam tip: B2B (Clairo) typically has lower baseline rates than B2C (Zoko). This difference affects sample size directly — lower baselines need more data.
Significance level () and power explained
— false positive protection
- Plain language definition: The probability that your test will declare a winner when there is no real difference.
- Standard default: (95% confidence). Reason: proposed by Ronald Fisher in 1925 as a pragmatic convention. 1 in 20 tests will show a false positive by chance.
- Lowering (e.g., to 0.01) increases required sample size. For Clairo (3.7% baseline, 10% MDE): at → ~1,400 per variant; at → ~2,100 per variant.
- Pharma/finance may use 99% confidence; for startup experimentation, leave as default unless a senior statistician documents a reason to change.
Power ( ) — false negative protection
- Plain language definition: The probability that your test will detect a real effect when one genuinely exists.
- Standard default: 80% power (proposed by Jacob Cohen, 1969). Means: if variant B actually works at the MDE, the test will find it 4 out of 5 times. 1 in 5 real effects will be missed.
- Higher power (e.g., 90%) increases sample size. For Clairo: 80% → ~1,400; 90% → ~1,870 per variant.
- 80% is the industry default for marketers. Increasing power costs more data and time.
Connection between , power, and sample size: Both control different error types. Adjusting either one changes the required sample size. More perfection → more data.
Minimum detectable effect (MDE) — the business-driven input
MDE sets the smallest lift that would change your decision. It is not what is easiest to detect; it is what matters.
| MDE size | Implication | Example sample size |
|---|---|---|
| Tiny (1% absolute) | Requires hundreds of thousands of users per variant; months to years of traffic | Impractical for startups |
| Practical (10-20% relative) | Sample size in hundreds or low thousands; feasible in days/weeks | Clairo’s 10% MDE → 9,500 per variant on free-to-paid conversion (impractical) |
| Large (e.g., 20% relative) | Much smaller sample size; blinds test to any smaller effect | 20% MDE from 34% baseline → ~770 per variant (but misses a 12% lift) |
Rule: Set MDE based on the minimum improvement that would change your action. If you wouldn’t update pricing unless a CTA lifts conversion by at least 10%, then MDE = 10%.
Worked example: Clairo vs. Zoko (live calculator outputs)
Evan Miller’s sample size calculator (raw number per variant)
- Clairo (activation rate proxy): Baseline 34%, MDE 10% relative, , 80% power → ~1,400 per variant.
- Zoko: Baseline 38%, MDE 8% absolute (to 46%), same defaults → ~560 per variant.
Exam tip: Use relative MDE when baseline is small; absolute MDE is often clearer for higher baselines. Always specify which type you use.
VWO test duration calculator (adds traffic)
- Clairo (free→paid conversion): Baseline 3.7%, MDE 10% relative → needs ~9,500 per variant. With 620 sign-ups/month (≈21 visitors/day) → test would take 15 months — likely not runnable.
- Clairo (activation rate): Baseline 34%, MDE 10% relative → needs ~1,400 per variant. With 620 sign-ups/month → test completes in ~2 months — runnable.
- Zoko: Baseline 38%, MDE 8% absolute → needs ~560 per variant. With 38,000 monthly visitors → sample reached in <1 day. But never stop before 7 calendar days to capture day-of-week patterns.
Three mistakes that break sample size logic
| Mistake | What happens | Fix |
|---|---|---|
| Peeking – checking results before sample is reached | Inflates false positive rate; early lucky streak looks like a winner | Set sample size in advance; do not open results until it is met |
| Setting MDE too low | Need hundreds of thousands of users → test impossible | Set MDE based on minimum decision-changing lift (10–20% relative is practical) |
| Forgetting sample size is per variant | Only plan for total, not per arm; A/B test needs ×2, A/B/C needs ×3 | Always double/triple according to number of variants |
Connecting sample size to experiment modes (template)
Teams record the calculation in a template with columns:
- Hypothesis & metric (e.g., H1: changing CTA lifts free→paid from 3.7% to ≥4.07%)
- Experiment mode (full valid / directional / proxy)
- Sample target vs. actual (e.g., target 1,400 per variation)
- Result & decision (e.g., “101% reached → conclusive” or “39% reached → directional, flag as inconclusive”)
The target column is the output of the calculator; the mode column forces the question: is this test runnable?
Key takeaways
- Sample size separates signal from luck; it is the minimum observations needed before you can trust the result.
- Three pre‑test inputs: baseline conversion rate, MDE (business‑driven), and α/power (use defaults: 0.05 and 80%).
- MDE must be set on business logic, not statistical convenience. Practical MDEs are 10–20% relative.
- Always run the test for at least 7 calendar days even if the sample size is reached earlier.
- Peeking, too‑low MDE, and forgetting “per variant” are the three common pitfalls.
Type I and Type II Errors
Every A/B test can produce a wrong answer in exactly two ways. The core idea: You are comparing a null hypothesis (no difference) against a real effect. A test can say “there is a difference” when there isn’t (false positive, Type I error) or say “no difference” when there really is one (false negative, Type II error). Understanding these two failure modes lets you control them — and prevents costly misread data.
The 2×2 Matrix of Test Outcomes
The columns represent reality (H₀ true or false). The rows represent your decision (reject H₀ or fail to reject H₀). Four cells, two of which are errors.
| H₀ true (no real effect) | H₀ false (real effect exists) | |
|---|---|---|
| Reject H₀ (say “effect found”) | Type I error (false positive) – wrong conviction | Correct rejection – true positive |
| Fail to reject H₀ (say “no effect”) | Correct non-rejection – true negative | Type II error (false negative) – wrong acquittal |
Exam tip: Type I = false positive; Type II = false negative. Mnemonic: “I” sounds like “lie” (positive lie); “II” sounds like “nein” (negative lie).
Definitions in Plain Terms
-
Type I error: You think the variant works when it actually does not. Analogy: Convicting an innocent person. The evidence pointed to guilt, but the person didn’t do it.
-
Type II error: You think the variant does not work when it actually does. Analogy: Acquitting a guilty person. Evidence was insufficient, but they really did it.
Key Quantities: Alpha and Beta
| Symbol | Name | Meaning | Typical value | Controls |
|---|---|---|---|---|
| Significance level | Probability of Type I error (false positive) | 0.05 (5%) | Set before test | |
| – | Probability of Type II error (false negative) | 0.20 (20%) | Reduced by larger sample | |
| Statistical power | Probability of detecting a real effect of size MDE | 0.80 | Increased by larger sample |
If you set , then 5% of tests where H₀ is true will still produce a false positive by chance alone — 1 in 20.
If power = 80%, then 20% of the time a real effect of the chosen MDE will not be detected.
Causes & Business Costs
Type I Error
- Most common cause: Peeking — stopping the test too early with very little traffic. A small early fluctuation looks like a signal.
- Business cost: You scale a change that does not actually improve performance. Resources (time, money, effort) are wasted on a wrong direction. False beliefs about what drives conversion become embedded.
Example – Clairo (CTA copy test): Team saw 4.2% conversion vs 4.0% after only 100 users and declared B the winner. Required sample was 9,500 per variant. The 0.2-point lift was random variation. Switching to the wrong CTA could drop conversion from 3.7% to under 3%, costing millions.
Example – Zoko (starter kit CTA test): High traffic reached sample size in 18 hours (weekday afternoon). Observed a 9-point lift, but the test should have run a full 7-day week. The real effect might have been 2 points or the opposite.
Type II Error
- Most common cause: Underpowered test — sample size too small. Or setting MDE too high (e.g., expecting a 20% lift when the true lift is 9%).
- Business cost: You miss an opportunity. The better variant is discarded, and the worse baseline continues. Unlike Type I, this loss is invisible — you never see what you could have gained. Over time, the missed benefit compounds.
Example – Clairo (onboarding flow test): Required 1,400 users per variant, but budget constraints limited it to 310. Test showed no significant result; team abandoned a genuinely better onboarding flow.
Example – Zoko (unboxing insert test): MDE set at 20% relative change; true effect was 9% improvement in 90-day repurchase. Underpowered test missed it.
How to Control Each Error
| Error | Controlled by | Mechanism | Side effect |
|---|---|---|---|
| Type I | Lower (significance level) | Reduce the threshold for “statistical significance” | Increases required sample size; may raise Type II risk if power drops |
| Type II | Increase power (via larger sample size) | More data makes real effects easier to detect | More time, traffic, and resources needed |
Both are design decisions made before the test runs.
When to Worry More About Which Error
Prioritize controlling Type I when:
- Scaling the change requires a major resource commitment (e.g., paid acquisition campaign, engineering sprint, pricing change) — hard to reverse.
- The market is slow-moving; acting on noise can waste months of momentum before the mistake is visible.
Prioritize controlling Type II when:
- Test opportunities are scarce (e.g., low traffic, few chances to rerun). Missing one variant is a big loss.
- You are in a fast competitive market; even a small improvement can give an edge. Missing it may let a competitor capture that advantage first.
Exam tip: There is no universal “worst error” — it depends on context. Type I bites when scaling is costly; Type II bites when opportunities are rare.
Key Takeaways
- Type I error (false positive): reject H₀ when H₀ is true. Caused by peeking, small samples. Cost: wasted resources on a false winner.
- Type II error (false negative): fail to reject H₀ when H₀ is false. Caused by underpowered tests, MDE too high. Cost: lost opportunity, invisible.
- The 2×2 matrix maps both errors; two cells are errors, two are correct outcomes.
- controls Type I; power (via sample size) controls Type II.
- Which error is more damaging depends on context: costly scaling → Type I; scarce/fast opportunities → Type II.
Startup Experimentation Math
Real-world startups face limited traffic, time, and patience — ideal statistical conditions (large samples, no peeking, fixed duration) are often impossible. The goal is informed trade-offs, not shortcuts: make calculated compromises without abandoning statistical thinking.
The Speed–Accuracy Trade-off
Any experiment exists on a spectrum from maximum speed (no controls) to maximum accuracy (full statistical validity). The choice is a design decision based on business context, not a mistake.
- Ship & watch (left): make a change, watch aggregate metric. No hypothesis, no control — cannot isolate cause.
- Directional testing (middle): run with a hypothesis but a sample below the required size. Gives trend, not proof.
- Full valid test (right): pre-calculated sample, run to completion, no peeking. Produces conclusive, defensible results.
Three Experiment Modes
| Mode | When to use | Example | Acceptable conclusion |
|---|---|---|---|
| Full valid test | Decision is costly, hard to reverse, high stakes | Pricing change, new acquisition channel, major product update | Significant (or not) at pre-set confidence level → scale or abandon |
| Directional test | Decision has low cost, reversible, fast iteration | Email copy, secondary CTA, onboarding message | Direction only — do not use to justify scaling |
| Proxy test | Final metric needs too much traffic/time; measure a leading indicator instead | Free→paid conversion too slow → test activation rate | Run a full valid test on the proxy; not a permanent substitute |
Choosing the mode is a business decision: If this test gives a wrong result, how big is the impact? How reversible?
Partial Sample Results: What You Actually Know
Running a test on less than 100% of the required sample degrades effective power and inflates Type I error. The table below shows the practical meaning:
| % of required sample | Effective power | Type I error rate | What you know |
|---|---|---|---|
| 20% | ~30% | 15–20% | Almost nothing — as good as no test |
| 40% | ~50% | 10–12% | Weak direction, highly uncertain |
| 60–80% | ~60–70% | 7–8% | Signal becomes more reliable, still not conclusive |
| 100% | 80% | 5% | Valid, actionable result |
Exam tip: Never use a partial-sample result to justify scaling. Only a full valid test (100% sample, 80% power, 5% Type I) produces a decision. Anything else is a directional signal — useful for planning the next full test, not for rolling out changes.
Proxy Metrics
A proxy metric is a leading indicator that reliably predicts the final (lagging) metric and can be measured faster and with less data.
Leading vs. Lagging Metric (example: Clairo)
- North star metric (lagging): number of paid users.
- Leading metrics (can be improved to affect the north star):
- Visitors → free user conversion rate
- Free user → paid user conversion rate
- Number of visitors (via campaigns)
To increase paid users, you can improve any leading metric. A proxy test chooses one leading metric as a stand-in.
Three Requirements for a Valid Proxy
- Correlation — must be documented (e.g., 10% of free users become paid → free users strongly predict paid users).
- Sensitivity — must respond to the change you are testing (e.g., an activation-rate change will affect free→paid conversion).
- Speed advantage — must genuinely require less traffic/time than testing the final metric directly.
Real Example: Clairo
- End metric: free→paid conversion rate (3.7% → 4.1%)
- Required sample: ~19,000 users → 15 months at 620 sign-ups/month ❌
- Proxy chosen after cross-team discussion: Activation rate (user reaches aha moment within 7 days).
- Proxy sample: ~2,800 users (1,400 per variant) → testable in 2–3 months ✅
- Validation: users who never activate never convert; activation directly correlates with free→paid.
- Mode for proxy test: full valid test on the proxy (no compromise on accuracy for this high-stakes decision).
Exam tip: Proxy tests are speed tools. As soon as you have enough traffic to test the final metric directly, switch back to it.
Experiment Decision Log
A 4-column log keeps the experimentation system honest. Record before results are known.
| Column | Content | Purpose |
|---|---|---|
| 1 | Hypothesis + Metric | “I will increase activation rate from 34% to 39%.” Prevents vague descriptions like “improve customer journey.” |
| 2 | Mode | Full valid / Directional / Proxy (chosen before seeing data). Forces team to agree on acceptable conclusion type. |
| 3 | Sample target vs. actual | Required sample size (from calculator) and actual sample reached. The ratio instantly tells you if result is valid or just directional. |
| 4 | Decision + Acknowledged risk | For mode 1: scale decision. For mode 2: state that risk is elevated and result is directional only — do not treat as proof. |
This log is used throughout the course to document experiments consistently.
Key takeaways
- The speed–accuracy trade-off is a deliberate business decision, not a statistical error.
- Three modes: full valid (high stakes), directional (low cost/reversible), proxy (when final metric is too slow).
- Partial samples ( < 100%) give only directional signals — never use them to justify scaling.
- A valid proxy must be correlated, sensitive, and faster to measure; test it with full validity.
- Always log every experiment with hypothesis, mode, sample ratio, and acknowledged risk.
Practical Landing Page A/B Cases – Clairo & Zoko
Two complete A/B tests are run end-to-end to show the full 6‑step workflow in action. One test (Clairo, B2B) produces a significant result; the other (Zoko, B2C) does not. Both are valid, designed tests – the difference lies only in the outcome and the decision that follows.
The 6‑Step Workflow (Applied Identically to Both Cases)
- State hypotheses – Define (null: no difference) and (alternative: expected effect) before any data is collected.
- Calculate required sample size – Based on baseline rate, minimum detectable effect (MDE), , power = 80%.
- Choose experiment mode & decision rule – Select mode (full valid test, directional, etc.) and pre‑commit to the rule: if at full sample → deploy; otherwise don’t.
- Run test – No peeking; reach the required sample; run at least 7 calendar days.
- Read results – Check vs and actual sample vs required. Apply error‑type framework.
- Log decisions – Hypothesis, mode, sample ratio, final decision, risk acknowledged.
Exam tip: Steps 1–3 must be written down before the test starts. Peeking invalidates the error guarantees.
Case A: Clairo (B2B Startup)
Goal: Increase 7‑day activation rate from 34% to at least 37.4% (10% relative lift) using a new guided first‑meeting flow.
Step 1 – Hypotheses
- : New onboarding flow produces no difference in 7‑day activation rate.
- : New flow increases 7‑day activation rate from 34% to at least 37.4%.
Step 2 – Sample Size
- Baseline: 34%
- MDE: 10% relative (3.4% absolute).
- , power = 80%
- Required: ~1,400 users per variant → total 2,800 users.
Clairo gets 620 sign‑ups/month → 2.3 months needed; achievable within a single quarter.
Step 3 – Mode & Decision Rule
- Mode: Full valid test (metric is the primary growth constraint; result drives next quarter’s roadmap).
- Decision rule:
- If at full sample → deploy new flow.
- If → do not scale; conduct UI audit / redesign and retest.
Steps 4–6 – Results & Decision
- Ran 70 days, reached 1,447 users per variant (103% of required). No peeking.
- Control: 34% activation
- Variant: 38.4% activation
- Lift: +4.4 p.p. (12.9% relative)
- (confidence 96.9%) → → significant.
Decision: Deploy the new flow. Log all details (hypotheses, mode, sample ratio, decision, risk acknowledged).
Key point: A properly validated proxy metric (7‑day activation predicting paid conversion) can be a legitimate test metric if the correlation is documented.
Key Takeaways – Clairo
- Hypotheses must be specific (baseline → target) and stated pre‑test.
- Sample size is derived from baseline, MDE, , and power; ensure traffic can support it.
- A significant result () allows deployment, but still monitor the actual downstream metric.
Case B: Zoko (B2C Brand)
Goal: Increase starter‑kit attach rate from 38% to at least 46% (8 p.p. absolute lift) by changing the CTA from “Start your skin journey” to “Try the 21‑day starter kit”.
Step 1 – Hypotheses
- : CTA change produces no difference in attach rate.
- : New CTA increases attach rate from 38% to at least 46%.
Step 2 – Sample Size
- Baseline: 38%
- MDE: 8% absolute (relative ~21%).
- , power = 80%
- Required: ~560 users per variant → total 1,120 users.
Zoko receives 38,000 visitors/month → can reach in <1 day, but test is still run for 7 calendar days to capture day‑of‑week variation.
Step 3 – Mode & Decision Rule
- Mode: Full valid test.
- Decision rule:
- If at 7‑day end → deploy new CTA.
- If → do not deploy; iterate copy/design and retest.
Steps 4–6 – Results & Decision
- Ran 7 days, reached 572 users per variant (102% of required).
- Control: 38% attach rate
- Variant: 40.3% attach rate
- Lift: +2.3 p.p. (absolute)
- (confidence 79%) → → not significant.
Decision: Fail to reject . Do not deploy. But do not abandon – the observed 2.3% lift is directional; the copy may have a small real effect that a larger sample or a lower MDE could detect. The next step is to iterate (better copy, different design) and retest.
Interpreting a Non‑Significant Result – The Critical Lesson
- ✅ The result is evidence that the specific CTA did NOT produce the 8% lift we designed for.
- ✅ It is a valid, complete test at the designed error rates.
- ✅ It provides directional data (variant performed 2.3% better than control).
- ❌ It does not prove the CTA has zero effect.
- ❌ It does not prove the variant is worse than control (it’s better).
- ❌ It does not give permission to stop testing the hypothesis.
Exam tip: A non‑significant result does not mean “nothing happened”. It means the observed effect is smaller than the MDE the test was powered to detect. The correct response: iterate, don’t conclude.
Comparison Table: Clairo vs. Zoko
| Dimension | Clairo (B2B) | Zoko (B2C) |
|---|---|---|
| Metric | 7‑day activation rate (proxy for paid conversion) | Starter‑kit attach rate (direct) |
| Baseline | 34% | 38% |
| Target lift | +3.4 p.p. (10% relative) | +8 p.p. (absolute) |
| Required sample/variant | 1,400 | 560 |
| Actual sample/variant | 1,447 (103%) | 572 (102%) |
| Observed lift | +4.4 p.p. (12.9% relative) | +2.3 p.p. (absolute) |
| ‑value | 0.031 | 0.21 |
| Decision on | Reject (significant) | Cannot reject (not significant) |
| Business action | Deploy new flow | Do not deploy; iterate & retest |
| Lesson | A validated proxy metric can be trusted if correlation is documented. | Observed effect < designed MDE → directional signal, not conclusion. |
Key Takeaways – Practical A/B Cases
- The 6‑step workflow is the same for every test – only the numbers and the outcome differ.
- A significant result () permits deployment; a non‑significant result signals iteration, not abandonment.
- Never peek, always reach the pre‑calculated sample, and run at least 7 days.
- A non‑significant test does not confirm the null; it only fails to reject it. The observed effect is a directional input for the next experiment.
- Proxy metrics are acceptable only if their correlation to the true business metric is validated and documented.
- Growth is iterative – no single test gives the final answer.
Practical Landing Page A/B Case – II
This case walks through reading real A/B test results from VWO, interpreting significance, making deploy decisions, and logging experiments across multiple runs. The standard reading order is: sample → rates → statistics.
Clairo Activation Test – Significant Result
| Metric | Control (A) | Variant (B) | Difference |
|---|---|---|---|
| Sample per variant | 1,447 | 1,447 | – |
| Target per variant | 1,400 | 1,400 | 103% reached |
| 7‑day activation rate | 34.0% | 38.4% | +4.4 pp absolute |
| Relative lift | – | – | +12.9% |
| p-value | – | – | 0.031 |
| Confidence level | – | – | 96.9% |
- Sample: 2,894 total visitors; full sample reached (103% of target).
- Rates: Variant outperforms control by 4.4 percentage points.
- Statistics: , confidence → statistically significant.
- Decision: Reject the null hypothesis (H₀) – the new onboarding flow produces a real improvement. Deploy the variant across the site.
Exam tip: Significance alone is not enough. Always verify that the target sample size was reached (as designed for the chosen alpha and power). Here both conditions are met.
Zoko CTA Test – Non‑Significant Result
| Metric | Control (A) | Variant (B) | Difference |
|---|---|---|---|
| Sample per variant | 572 | 572 | – |
| Target per variant | 560 | 560 | 102% reached |
| Attach rate | 38.0% | 40.3% | +2.3 pp absolute |
| Relative lift | – | – | +6.1% (approx.) |
| p-value | – | – | 0.21 |
| Confidence level | – | – | 79% |
- Sample: Full sample reached (102% of target).
- Rates: Observed absolute lift of 2.3 percentage points.
- Statistics: , confidence → not statistically significant.
- Decision: Fail to reject H₀. Do not deploy the new CTA copy.
What a Non‑Significant Result Actually Means
- The test provides evidence that the new copy did not produce the designed minimum detectable effect (MDE) (here, an 8% relative lift).
- It does not prove the copy has zero effect – the observed +2.3 pp lift suggests a real effect smaller than the MDE.
- Therefore, the hypothesis is retained, and the team should iterate (e.g., refine the copy) rather than abandon the idea.
Directional Tests vs Full Valid Tests
- Directional test: Run when insufficient traffic prevents reaching the target sample. Results are inconclusive but can be logged to preserve the hypothesis for a future full run.
- Full valid test: Reaches the pre‑specified sample size and satisfies the design parameters (alpha, power). Only this mode supports statistical conclusions.
Example from the Experiment Log (Zoko CTA)
| Row | Test | Mode | Sample Reached | % of Target | Result | Decision |
|---|---|---|---|---|---|---|
| 2 | Zoko starter kit CTA | Directional | 220 per variant | 39% | Inconclusive | Hypothesis retained |
| 5 | Zoko starter kit CTA | Full valid | 572 per variant | 102% | Not significant (p=0.21) | Do not deploy; hypothesis retained |
The log captures both runs – the same hypothesis tested twice at increasing efficiency. This is the experimentation life cycle in action.
Experiment Logging – The Three‑Act Story
The experiment log tracks the history of each hypothesis across runs:
- Row 2 (directional): Zoko CTA hypothesis tested with insufficient sample → inconclusive → hypothesis retained.
- Row 5 (full valid): Same hypothesis tested with full sample → still non‑significant → hypothesis retained again, but with clear evidence that the effect is below the MDE.
- Row 1 (Clairo): Full valid test, significant → deployed.
The log provides continuity and informs the next iteration. It is not a contradiction to log two entries for the same hypothesis – that is the system working as designed.
Exam tip: A non‑significant result from a full valid test is a valid and complete outcome. It means the test was properly powered but the effect was not large enough to detect. Use it as input for the next iteration – do not treat it as failure.
Key Takeaways
- Read results in order: sample → rates → statistics.
- Clairo: p=0.031, full sample → significant → deploy.
- Zoko: p=0.21, full sample → not significant → do not deploy, but retain hypothesis.
- Directional tests are valid for learning; only full valid tests support statistical conclusions.
- Experiment logs capture the lifecycle of a hypothesis across multiple runs.
- Non‑significance does not imply zero effect – only that the observed effect is smaller than the MDE.
- Systematic iteration, not one‑shot wins, drives growth.