Intuition and Construction
The sampling distribution of the sample mean is approximately normal (Central Limit Theorem) with mean and standard error . For a normal distribution, 95% of all values lie within of .
Now consider forming an interval around any observed :
If that particular is within of , then the interval will contain ; if is farther away (in the 5% tails), the interval will not contain . Because 95% of all possible are close enough, 95% of all such intervals will capture the true population mean. This is the core idea of a confidence interval.
Formal definition (σ known): A confidence interval for is where is the confidence coefficient and is the critical value from the standard normal distribution (area in the upper tail).
The term is the margin of error.
Worked Example: Basavaraju's Customer Satisfaction Scores
- Population standard deviation (historical value)
- Sample size →
- Sample mean
For any confidence level, the interval is .
| Confidence Level | Margin of Error | Confidence Interval | |
|---|---|---|---|
| 80% | 1.28 | 1.07 | |
| 90% | 1.64 | 1.37 | |
| 95% | 1.96 | 1.63 | |
| 99% | 2.58 | 2.15 |
How to obtain : Use the standard normal table (or inverse‑normal function).
- For 95% confidence, , →
- For 99% confidence, , →
- For 90% confidence, , →
- For 80% confidence, , →
Effect of the Confidence Coefficient
- Higher confidence → wider interval (larger ). A 99% interval is wider than a 90% interval because we need to cover more of the sampling distribution.
- 100% confidence would require , producing an interval → useless.
- Choosing a confidence level is a trade‑off: higher confidence gives more assurance that the interval contains , but at the cost of precision (wider interval). The decision depends on the cost of being wrong:
- Critical applications (e.g., FDA drug approval) often require 95% or 99% intervals.
- For exploratory or low‑stakes decisions, 80% or 90% may be acceptable.
- Consistency: When comparing different intervals, always use the same confidence coefficient.
Exam tip: A 95% confidence interval does not mean “there is a 95% probability that lies in this particular interval.” It means that if we repeated the sampling process many times, 95% of the resulting intervals would contain .
Effect of Sample Size
The margin of error is .
- Larger → smaller margin of error (more precise estimate).
- Because appears in the denominator, doubling reduces the margin by a factor of (not by half).
- Conversely, to achieve a desired margin of error, we can solve for the required sample size (covered later).
Key Takeaways
- The confidence interval for (σ known) is .
- 95% of all intervals constructed this way contain ; the confidence level describes the long‑run success rate.
- Increasing the confidence level widens the interval; increasing sample size narrows it.
- The choice of confidence level (80%, 90%, 95%, 99%) depends on the context and the cost of being wrong.
- Always compute intervals with a fixed confidence coefficient when comparing results.
Confidence Intervals for the Mean (σ Known)
A confidence interval (CI) for the population mean quantifies the uncertainty around the sample mean . When the population standard deviation is known, the interval is built using the standard normal () distribution.
The formula:
- = confidence coefficient (e.g., 0.95).
- = -value that cuts off area in the upper tail of the standard normal.
- = standard error of , denoted .
Why it works
Because is normally distributed (exactly if population is normal, approximately by the CLT for ) with mean and SD , the standardized variable
follows a standard normal distribution. Hence, for a given , we can find such that , which rearranges to the CI above.
Assumptions
- Independent random sample.
- known (e.g., from historical data).
- Either or the population itself is normally distributed.
Worked Example 1 – Large Sample ()
Amit’s laddu shop: , , (known).
| Confidence Level | Margin of Error | CI | |
|---|---|---|---|
| 80% | 1.28 | ||
| 95% | 1.96 |
Exam tip: Increasing the confidence level (e.g., 80% → 95%) widens the interval because the margin of error grows with .
Worked Example 2 – Small Sample, Known ()
Puneet’s spa: , , (known). Since is small, we must assume the population is normally distributed for to be normal.
| Confidence Level | Margin of Error | CI | |
|---|---|---|---|
| 95% | 1.96 | ||
| 99% | 2.58 |
Interpretation: “99% of such intervals will contain the true population mean .” The margin of error increases with higher confidence.
Verifying normality: Plot sample data or use prior knowledge.
Key Takeaways (σ Known)
- CI for when is known: .
- Rely on normal distribution of (CLT for large , population normality for small ).
- Larger confidence level → larger → wider interval.
- Standard error .
Confidence Intervals for the Mean (σ Unknown)
In most real problems is unknown and must be estimated from the sample. Using the sample standard deviation in place of introduces extra uncertainty. The correct sampling distribution is the Student’s -distribution, not the standard normal.
Why can’t we just replace with and keep using ? The quantity is not standard normal because is a random variable. Its distribution is a -distribution with degrees of freedom.
Formula
- = critical value from a -distribution with degrees of freedom (df).
- = sample standard deviation.
- Margin of error: .
Properties of the -distribution
- Bell-shaped, symmetric about 0 (like ).
- Flatter (more spread) than the standard normal, reflecting the extra uncertainty.
- As df increases ( grows), the -distribution converges to the standard normal.
- Developed by William Gosset (working at Guinness).
Assumptions for Valid Use of
- Random sample.
- Population is approximately normally distributed (the procedure is robust to mild deviations, especially for ).
- No requirement to know .
Comparison: Known vs. Unknown
| Aspect | Known | Unknown |
|---|---|---|
| Critical value | ||
| Std. error | ||
| Distribution of | Standard normal | with df |
| Sample size requirement | or normal pop. | any size (but normality important for small ) |
Exam tip: Always check whether is known. If “population standard deviation is given” → . If “sample standard deviation ” → with df. The exam will often test this distinction.
Key Takeaways (σ Unknown)
- Use -distribution with df when is unknown and estimated by .
- Formula: .
- -distribution is wider than , giving larger margins of error for the same confidence level (especially for small ).
- Assumption of population normality is important for small samples; the procedure is robust for larger .
- As increases, and the two methods converge.
Confidence Intervals for Population Mean: σ Unknown
When the population standard deviation σ is unknown — the realistic case — we replace σ with the sample standard deviation and use the t-distribution with degrees of freedom instead of the standard normal. The structure remains identical: point estimate ± margin of error.
Margin of error (σ unknown):
Confidence interval:
The t‑distribution is slightly wider than the normal for small (heavy tails), reflecting the extra uncertainty from estimating σ.
Worked Example 1: Joti Hegde’s Salary Survey
Data summary (n = 103)
- Sample mean lakhs
- Sample standard deviation lakhs
- Standard error:
90% confidence interval (, , df = 102)
- Interval: lakhs
95% confidence interval (, , df = 102)
- Interval: lakhs
Exam tip: The only change from σ known to σ unknown is swapping for and using in place of σ. The mechanics are identical.
Worked Example 2: Hanumantha Pai’s Credit Card Expenditure
Data summary (n = 55)
- Sample mean ₹
- Sample standard deviation ₹
- Standard error:
80% confidence interval (, , df = 54)
- Interval: ₹
95% confidence interval (, , df = 54)
- Interval: ₹
Notice: a higher confidence level (95% vs. 80%) produces a wider margin of error, as expected.
When Is the t‑Interval Exact? (Reliability & Sample Size)
| Population shape | Sample size | Validity of t‑interval |
|---|---|---|
| Normal | Any | Exact – the formula is exact. |
| Approximately symmetric, bell‑shaped | Small () | Good approximation – often adequate. |
| Moderately skewed, no outliers | Adequate approximation – safe rule of thumb. | |
| Highly skewed or contains outliers | Needed approximation – increase to improve. |
In practice, for the t‑interval works well for most populations unless the distribution is severely non‑normal.
Determining the Sample Size for a Desired Margin of Error (σ Known)
Sample size planning happens before data collection. We want the smallest that yields a pre‑specified margin of error at a chosen confidence level.
For σ known:
Since σ is usually unknown before sampling, use a planning value for σ:
- Historical data – σ from a previous study.
- Pilot study – collect a small preliminary sample and use its .
- Range/4 rule – (crude but simple).
Example: Basavaraju’s Customer Satisfaction
- Historical σ = 5
- Desired 99% confidence interval ()
- Desired margin of error
Thus Basavaraju should have collected a sample of 166 (instead of 36) to achieve a 99% CI with margin ±1.
Exam tip: Always round the sample size up to the next integer (e.g., 165.8 → 166). A fraction gives a margin larger than desired.
Key Takeaways
- When σ is unknown, use the t‑distribution with degrees of freedom and in the margin of error.
- The t‑interval is exact if the population is normal; for non‑normal populations, (or for strong skew/outliers) produces good approximations.
- To plan sample size for a desired margin of error (σ known case), use , with a planning value for σ.
- A higher confidence level or a smaller desired requires a larger sample size.
Confidence Interval for Population Proportion
A confidence interval for a population proportion estimates the unknown true proportion of "successes" in the population based on a sample. Intuitively: if 47% of sampled customers are satisfied, the true proportion is likely within a range around that value, with a stated level of confidence.
The general form is the same as for the mean:
Here the sample estimate is the sample proportion (where = number of successes in the sample), and the margin of error depends on the sampling distribution of .
Sampling Distribution of
- is a binomial random variable ( successes out of trials).
- For large , the binomial is approximated by a normal distribution.
- Condition for approximation: and , where .
- The sampling distribution is centered at the population proportion with standard error:
Exam tip: Always check and to validate the normal approximation. If the condition fails, the confidence interval may be unreliable.
Margin of Error and Confidence Interval
If is approximately normal, the margin of error for a confidence level is . However, depends on the unknown . The solution is to use as an estimate for in the standard error:
Thus, a confidence interval for is:
where is the -score cutting off area in the upper tail of the standard normal distribution.
Worked Example 1: Basavaraju's Customer Satisfaction
- Sample size , number of successes (score ).
- .
- For a 95% confidence interval (): .
- Compute .
- Margin of error .
- 95% CI: .
Worked Example 2: Hanumantha's Credit Card Spend
- customers. Number spending less than ₹3000: .
- .
- .
80% confidence interval (, ):
- Margin of error .
- CI: .
99% confidence interval (, ):
- Margin of error .
- CI: .
The 99% interval is much wider – a trade-off between higher confidence and precision.
Sample Size Determination for a Desired Margin of Error
To achieve a desired margin of error (e.g., ) at a given confidence level, solve for in the margin-of-error formula:
But is unknown before sampling. A planning value must be used. Options:
- Use the sample proportion from a previous study.
- Conduct a pilot study and use its .
- Use a best guess / judgment.
- If no information, use (the most conservative – gives the largest ).
Exam tip: Using yields the maximum possible sample size for a given and , because is maximized at . This ensures the actual margin of error will not exceed even if the true proportion is far from 0.5.
Example: Hanumantha's desired margin at 99% confidence
- , .
- Use planning value (conservative).
- .
With only , the margin of error was 0.16; a sample of 666 would achieve the desired 0.05.
Key takeaways
- Confidence interval for : , valid when and .
- The margin of error uses as a substitute for unknown in the standard error.
- Higher confidence levels widen the interval; larger sample sizes narrow it.
- To determine sample size, use a planning value (previous data, pilot, guess, or ).
- is the safest (most conservative) choice, maximizing the required .
Confidence Interval for Population Variance
We now construct a confidence interval (CI) for the population variance (or standard deviation ). The key insight: while the sample variance is our best point estimate, its sampling distribution is not symmetric, so the CI cannot be written as “ margin of error”. Instead we use the chi-square distribution.
Intuition and the Sampling Distribution
If the population is normally distributed, the random variable
follows a chi-square distribution with degrees of freedom:
This is the pivot we invert to obtain a CI for .
Formula for the Confidence Interval
Choose a confidence coefficient . Let and be the critical values from the distribution such that
Rearranging (taking reciprocals, multiplying by ) gives the CI:
Definition where and are from .
To obtain a CI for , take square roots of the endpoints.
Worked Example: Basavaraju’s Customer Satisfaction Survey (95% CI)
Given: .
Step 1 – Degrees of freedom: .
Step 2 – Critical values (from chi-square table or template):
| Tail probability () | Value |
|---|---|
Thus of lies between and .
Step 3 – Plug into formula:
Step 4 – CI for : Take square roots:
Exam tip: Because the chi‑square distribution is not symmetric, we cannot write the interval as something. This is a common trick – the CI for variance is always of the form .
Varying Confidence Levels: 80% and 90% CI
For the same Basavaraju data, we can compute intervals at other confidence levels. The process and critical values change.
80% CI ()
, .
| Tail probability | Value |
|---|---|
Computation:
90% CI ()
, .
| Tail probability | Value |
|---|---|
Computation:
Observation: As confidence increases, the interval widens, just as with means and proportions.
Example: Jyothi Hegde’s Salary Data
Given: , .
. Critical values:
| Tail probability | Value |
|---|---|
CI for :
CI for :
Process Overview
Key Takeaways
- Assumption: Population must be normally distributed for the chi‑square pivot to be exact.
- Formula: .
- Asymmetry: Because is not symmetric, the CI is not of the form margin of error.
- Critical values: Always use the upper‑tail critical value in the denominator of the left endpoint and the lower‑tail value in the denominator of the right endpoint.
- Standard deviation CI: Simply take square roots of the variance CI endpoints.
- Interpretation: Over repeated sampling, of such intervals will contain the true .
Hypothesis Testing – Introduction
Statistical inference goes beyond confidence intervals. Hypothesis tests use sample data to decide whether a claim about a population parameter (mean , proportion , or variance ) is plausible. The logic mirrors confidence intervals: we make an inference about a population from limited sample evidence, because the decisions that follow will affect the entire population, not just the sample.
Purpose and logic
The core question: Does the sample provide enough evidence to support a specific statement about a population parameter? Typical claims:
- Is greater than a certain value?
- Is equal to a certain value?
- Is less than a certain value?
If the evidence is strong, we can confidently act on that claim for the population. If not, we must be cautious.
Null and alternative hypotheses
Every test involves two competing statements:
- Null hypothesis () – the tentative assumption about the parameter. It represents the status quo or the default position.
- Alternative hypothesis () – the opposite of ; it is the claim the test is designed to support.
The test then evaluates whether sample data can reject in favour of .
Important: The outcome is never “accept ”. It is reject or do not reject . “Do not reject” means the evidence is insufficient to prove false – it does not prove true.
Two possible outcomes
| Sample evidence | Decision | Meaning |
|---|---|---|
| Definitive evidence that is false | Reject | The claim () is supported beyond reasonable doubt. |
| Not definitive; too risky to reject | Do not reject | Sample is inconclusive; status quo remains. |
Legal analogy (Innocent until proven guilty)
The logic of hypothesis testing is exactly the logic of a criminal trial.
- Null hypothesis (): The accused is innocent.
- Alternative hypothesis (): The accused is guilty (the claim brought by the prosecution).
- Evidence: Sample data (video footage, fingerprints, alibi).
- Test outcome:
- If the evidence is convincing beyond reasonable doubt → reject → guilty.
- If the evidence is not convincing (reasonable doubt remains) → do not reject → innocent until proven guilty.
Note: A “not guilty” verdict does not prove innocence; it only means the prosecution failed to prove guilt. Similarly, failing to reject does not prove is true.
Setting up the hypotheses correctly
Two questions guide the setup:
- Who is conducting the test and trying to make a claim about the parameter?
- What specific claim are they making?
The claim being argued for is always placed in the alternative hypothesis . The opposite statement becomes .
Example: A manufacturer claims the mean battery life exceeds 100 hours. – Claim: (what they want to prove) → . – Default (null): .
Exam tip: The way hypotheses are stated depends entirely on context. Always identify who wants to prove what before writing and .
Errors in hypothesis testing
Because sample evidence is probabilistic, two types of errors can occur:
- Type I error: Rejecting when it is actually true (false positive).
- Type II error: Not rejecting when it is actually false (false negative).
The goal is to minimise both, but a trade-off always exists. Quantifying and controlling these errors (similar to confidence level in interval estimation) will be covered in later sections.
Key takeaways
- Hypothesis tests decide whether a claim about , , or is supported by sample data.
- is the tentative assumption (status quo); is the claim being tested.
- The only decisions are reject (strong evidence) or do not reject (insufficient evidence). Never “accept ”.
- Legal analogy: = innocent (presumed true); = guilty (must be proven); the verdict mirrors the test outcome.
- To set hypotheses: ask who is making a claim and what the claim is – the claim goes in .
- Two unavoidable errors exist (Type I and Type II); they will be formalised later.
Types of Hypothesis Tests
Hypothesis tests about a population parameter (here the mean ) take one of three forms depending on the direction of the claim being tested:
- Lower‑tail test – claim is
- Upper‑tail test – claim is
- Two‑tail test – claim is
The null hypothesis () always contains the equality (, , or ). The alternate hypothesis () is the claim the researcher wants to support using sample data.
| Test Type | When to use | ||
|---|---|---|---|
| Lower‑tail | Want to prove is less than | ||
| Upper‑tail | Want to prove is greater than | ||
| Two‑tail | Want to prove is different from (not specifically smaller or larger) |
Exam tip: The equality sign always belongs in . Identify the claim first; then write as that claim, and as its opposite (including the equality).
Setting up the hypotheses – two guiding questions
- Who is making the claim? (the party using the sample data)
- What is the specific claim about ? (less than, greater than, not equal to)
The answer to question 2 determines the form of and, consequently, the test type.
Worked examples
Example 1: Under‑filling of lentil packets (Lower‑tail test)
- Context: Consumer Affairs Department suspects packets labelled “500 g” contain less.
- Claim: (the average weight of all packets is deliberately below the label).
- Sample: , g, g, significance level .
- Hypotheses:
- Decision rule: If the p‑value (probability of observing a sample mean when is true) is less than , reject and conclude under‑filling is occurring. Otherwise, the test is inconclusive – the sample does not prove under‑filling.
- Key point: “Not rejecting ” does not prove the packets contain 500 g; it only means the evidence is insufficient.
Example 2: Service response time (Upper‑tail test)
- Context: Geeta Kumari wants to check if mean response time has increased beyond the advertised 48 hours.
- Claim: (the average time exceeds the policy).
- Sample: , h, h, .
- Hypotheses:
- Decision rule: Reject if p‑value ; then corrective action is needed. Otherwise, the data do not prove that response time has increased.
Example 3: Petrol dispensing accuracy (Two‑tail test)
- Context: Ibrahim Khan worries the pumps dispense an amount different from the requested 30 L (over‑filling hurts profit, under‑filling cheats customers).
- Claim: (the average amount dispensed is not 30 L).
- Sample: , L, L, .
- Hypotheses:
- Decision rule: Reject if p‑value ; then recalibration is required. If not, the sample does not provide enough evidence of a deviation.
What does the p‑value mean?
The p‑value is the probability of observing a sample result as extreme as the one obtained (or more extreme) assuming the null hypothesis is true. A small p‑value (smaller than the chosen ) indicates that such an extreme result is unlikely under , leading to rejection of .
Key takeaways
- Three test forms: lower‑tail (), upper‑tail (), two‑tail ().
- The claim being tested always goes into ; equality always stays in .
- The test type determines the direction of “extremeness” for the p‑value calculation.
- “Not rejecting ” does not prove is true – only that the evidence is insufficient to support .
- Use the two‑guiding‑question method to set up hypotheses correctly: who is making the claim, and what is the claim?
Hypotheses, Errors, and the Testing Procedure
Every hypothesis test pits two competing claims about a population parameter against each other: the null hypothesis and the alternative hypothesis (or ). Only one of them is true. The test uses sample evidence to decide whether to reject in favour of . Because decisions rest on a sample, errors are possible.
Type I and Type II Errors
| Decision → Truth ↓ | Reject | Do not reject |
|---|---|---|
| true | Type I error (false positive) | Correct |
| true | Correct | Type II error (false negative) |
-
Type I error: rejecting when it is actually true. Probability = level of significance , chosen by the analyst (common: 0.05, 0.10, 0.01). Intuition: crying “wolf” when there is none.
-
Type II error: failing to reject when is true. Probability = (not directly controlled, depends on sample size, effect size, and ). Intuition: missing a real effect.
Trade-off: lowering reduces Type I error but increases (and vice‑versa). The only way to reduce both is to increase the sample size.
Exam tip: If the cost of a false claim (Type I) is high – e.g. convicting an innocent person – choose a small . If missing a real effect is more costly (Type II), use a larger or a bigger sample.
Example – Geeta’s response time hr, hr.
- Type I error: concluding that mean time exceeds 48 hr when it actually is 48 hr.
- Type II error: concluding that mean time is 48 hr when it actually exceeds 48 hr.
Key takeaways
- Type I = false rejection of ; probability = .
- Type II = false acceptance of (failure to reject ); probability = .
- is set by the analyst; depends on , sample size, and true effect.
- Trade‑off: small → large ; increase sample size to control both.
The Five‑Step Hypothesis Test
Step 1 – Hypotheses Identify (the status quo, often an equality or “≤”/“≥”) and (the claim the test seeks to support). Three common forms:
- Lower‑tail test: ,
- Upper‑tail test: ,
- Two‑tail test: ,
Step 2 – Significance level Set (maximum acceptable Type I error probability).
Step 3 – Test statistic Assume true at the boundary (). Use the sample to compute which follows a ‑distribution with (valid when or the population is normal).
Step 4 – Decision
- P‑value approach: Compute = probability, under , of observing a test statistic as extreme as (or more extreme than) the one obtained.
- If → reject .
- If → do not reject .
- Critical value approach: Determine the rejection region from ; reject if falls in that region. Both approaches yield the same decision.
Step 5 – Conclusion Translate the statistical decision into a real‑world interpretation for the decision maker.
Exam tip: “Do not reject ” is not the same as “accept ”. It means the sample lacked sufficient evidence to support . The test result is inconclusive about being true.
Worked Examples
Example 1 – Jalan Supermarket (lower‑tail) , g, g, (4%).
- (labelling claim), (underfilling).
- Standard error: .
- , lower‑tail .
- → reject .
- Conclusion: Sufficient evidence that the mean packet weight is below 500 g; the chain is likely underfilling.
Example 2 – Geeta Kumari (upper‑tail, two scenarios)
-
Strong evidence (, , ): , upper‑tail → reject → mean response time exceeds 48 hr.
-
Weak evidence (, , ): (unchanged) , → do not reject → insufficient evidence that mean time exceeds 48 hr.
Example 3 – Ibrahim’s petrol pumps (two‑tail) , L, L, .
- , (pumps may over‑ or under‑fill).
- , two‑tail .
- → reject .
- Conclusion: Evidence that the mean dispensed amount is not 30 L; recalibration is needed.
Key takeaways
- Hypothesis tests always involve two errors; controls Type I, is managed through design.
- The five‑step framework works for any population mean test: state hypotheses, choose , compute ‑statistic, obtain ‑value, draw conclusion.
- A ‑value smaller than means the observed sample is unlikely under → reject.
- Three test directions: lower‑tail (negative extreme), upper‑tail (positive extreme), two‑tail (both extremes).
- Always finish with a real‑world statement of what the decision means for the decision maker.
Hypothesis Testing for Population Mean
Hypothesis testing is a formal procedure to decide whether sample data provide enough evidence to reject a claim about a population parameter. For the mean , the process always follows the same skeleton: state hypotheses, compute a test statistic, find the p-value, and compare it to the significance level . The test statistic measures how far the sample mean is from the claimed value under the null hypothesis, in units of standard error.
Three types of tests
The alternative hypothesis ( or ) determines the type of test. The claim being tested is usually placed in .
| Test | form | form | When used |
|---|---|---|---|
| Lower-tail test | Claim that mean is less than a threshold (e.g., average lentil weight below labelled weight) | ||
| Upper-tail test | Claim that mean is greater than a threshold (e.g., average service work order completion time exceeds target) | ||
| Two-tail test | Claim that mean differs from a specific value (e.g., average petrol pump fill is not exactly 5 litres) |
Test statistic (unknown population )
When the population standard deviation is unknown (the usual case), the test statistic follows a -distribution with degrees of freedom:
where is the sample standard deviation and is the value under at equality.
Computing the p-value
The p-value is the probability, under , of obtaining a test statistic as extreme as (or more extreme than) the observed value, in the direction(s) specified by .
- Lower-tail test:
- Upper-tail test:
- Two-tail test: If : If : (By symmetry of the -distribution, the two-tail p-value is double the one-tail probability in the tail of the observed statistic.)
Decision rule: Reject if (including exact equality). Otherwise, do not reject .
Special case: known
If the population standard deviation is known, replace with . The test statistic becomes a -statistic (standard normal):
The p-value is computed from the standard normal distribution () instead of the -distribution. All other steps remain identical.
Key takeaways
- Hypothesis tests for are lower-tail, upper-tail, or two-tail depending on the claim in .
- Test statistic is with df (unknown ) or (known ).
- p-value = tail probability of the test statistic; two-tail p-value doubles the one-tail probability.
- Reject if ; otherwise fail to reject.
Hypothesis Testing for Population Proportion
Testing a claim about the population proportion follows the same logic as for the mean, but uses a different sampling distribution and test statistic. The sample proportion is approximately normal for large samples.
Hypotheses and test types
Again, determines the tail. Let be the hypothesized value under at equality.
| Test | form | form |
|---|---|---|
| Lower-tail | ||
| Upper-tail | ||
| Two-tail |
Test statistic
The sampling distribution of is approximately normal with mean and standard deviation , provided is large. Under at equality (), the test statistic is a -statistic:
p-value computation
The p-value is computed exactly as for the mean, but using the standard normal distribution :
- Lower-tail:
- Upper-tail:
- Two-tail: the one-tail probability in the direction of .
Exam tip: For proportions, the test statistic uses in the denominator (under ), not . This differs from the standard error used in confidence intervals for , where is used.
The decision rule remains the same: reject if .
Key takeaways
- Hypothesis tests for proportion mirror the structure for : null vs. alternative, tail choice, p-value comparison.
- Test statistic is , using in the denominator.
- The -test is valid when the normal approximation holds (large ).
- Reject if ; otherwise fail to reject.
Hypothesis Testing for Population Proportion
Hypothesis testing for a population proportion answers whether the true proportion of a categorical outcome differs from a hypothesized value . The sample proportion serves as the point estimate, and the test checks if the observed deviation from is statistically significant.
The test statistic and its distribution
Under the null hypothesis (assumed true at equality), the sampling distribution of is approximately normal if
The standard error of under is
The test statistic follows a standard normal distribution .
Three forms of the test
| Alternative hypothesis | Test type | -value |
|---|---|---|
| Upper‑tail | ||
| Lower‑tail | ||
| Two‑tailed |
Reject if the -value (the significance level).
Worked examples
Example 1 – Upper‑tail test (Swiggy coupons)
- Claim (to test): more than 10 % of coupon recipients will use them →
- Sample: used coupons →
- Conditions: , ✓
- Standard error:
- Test statistic:
- -value:
- Decision: → do not reject .
Exam tip: Even though the sample proportion (15 %) is higher than the hypothesized value (10 %), the p‑value exceeds 5 %. The difference is not large enough given the standard error – a reminder that sample evidence must be weighed against sampling variability.
Example 2 – Lower‑tail test (Jhansi sleeper adequacy)
- Claim (to test): less than 50 % of travellers find sleeper arrangements inadequate →
- Sample: , 280 said adequate, so number finding inadequate = . (watch the definition of “success” – here it is inadequate).
- Conditions: ✓
- Standard error:
- Test statistic:
- -value:
- Decision: → reject .
Exam tip: Carefully define what “success” means for the proportion under test. In this problem the claim is about inadequate arrangements, but the raw data gave the count of adequate – always translate to the outcome of interest.
Example 3 – Two‑tailed test (Noida hotel bookings)
- Claim (to test): the proportion of fully booked hotels is not 90 % →
- Sample: , 12 have vacancies → booked →
- Conditions: , ✓
- Standard error:
- Test statistic:
- -value:
- Decision: → reject .
Key takeaways
- The test statistic for a proportion is .
- Valid only if and .
- Three alternatives: , , ; compute the -value accordingly.
- Reject when -value .
- Watch the definition of success – it must align with the claim.
Hypothesis Testing for Population Variance
Hypothesis testing for a population variance uses the sample variance to judge whether differs from a hypothesized value . The procedure requires the population to be normally distributed.
The test statistic and its distribution
If the population is normal, the quantity follows a chi‑square distribution with degrees of freedom.
Under (assuming ), the test statistic has a distribution.
Three forms of the test
| Alternative hypothesis | Test type | -value |
|---|---|---|
| Lower‑tail | ||
| Upper‑tail | ||
| Two‑tailed |
Reject if -value .
Worked examples
Example 1 (Upper‑tail test – Meerut bus arrival variance)
- Claim (to test): variance of arrival times is below 4 minutes.
- Sample:
- Conditions: population assumed normal ✓
- Test statistic:
- -value:
- Decision: → do not reject .
Interpretation: The sample does not provide evidence that the variance is below 4 minutes.
Example 2 (Two‑tailed test – Gorakhpur driving test scores)
- Claim (to test): variance of new exam scores differs from historical 100 →
- Sample:
- Test statistic:
- -value:
- Decision: → reject .
Key takeaways
- Test on variance uses with degrees of freedom.
- Requires normal population.
- Three forms: lower‑tail, upper‑tail, two‑tailed.
- For two‑tailed tests, double the smaller tail probability because the chi-square distribution is asymmetric.
- Reject when -value .
Hypothesis Testing for Population Variance
Hypothesis tests for a population variance are used to judge claims about the dispersion of a normally distributed population – e.g., whether a production process has become too variable or whether a variance equals a specified target . The test relies on the chi-square distribution because the sample variance scaled by degrees of freedom follows a distribution under normality.
Test Statistic
Under the null hypothesis , assuming a normal population:
where = sample size, = sample variance, = hypothesized value.
Three Types of Tests
The alternative hypothesis determines the type of test and how the p-value is computed.
| Type | P‑value | ||
|---|---|---|---|
| Lower‑tail | |||
| Upper‑tail | |||
| Two‑tail | … (area in the smaller tail) |
Decision Rule
Compute the test statistic from sample data. Obtain the p‑value based on the distribution. Reject if (the chosen significance level). Otherwise, do not reject .
Worked Process (Flowchart)
Exam tip: The chi‑square test for variance is valid only when sampling from a normal population. Violating this assumption can severely distort the p‑value. Always check normality (e.g., histogram, normal probability plot) before applying this test.
Key takeaways
- Test statistic: with degrees of freedom.
- Three test types: lower‑tail, upper‑tail, two‑tail – choose based on the claim.
- Decision: reject if p‑value .
- Normality of the population is a critical assumption.
Module 5 Recap: Confidence Intervals and Hypothesis Tests
This module used sampling distributions (of , , ) to perform statistical inference via two complementary approaches: confidence intervals (estimation) and hypothesis tests (decision‑making).
1. Confidence Intervals
For each parameter, a confidence interval provides a range of plausible values at a chosen confidence level.
| Parameter | Point estimate | Form of interval | Distribution used |
|---|---|---|---|
| (or if known) | |||
| Standard normal () | |||
| Direct lower and upper bounds: |
- Sample size determination for and : set to achieve a desired margin of error.
2. Hypothesis Tests
Hypothesis testing answers “Is the population parameter equal to, greater than, or less than a specific value?” Steps:
- Identify the claim and who is making it – this guides the structure of and (not always symmetric).
- Write null and alternative hypotheses based on the claim.
- Choose the test statistic according to the parameter and sampling conditions:
- : (or if known)
- :
- : (only under normality)
- Compute the p‑value – the probability of obtaining a test statistic as extreme as observed, assuming is true. This is the probability of a Type I error.
- Decision: Reject if p‑value ; otherwise do not reject.
Exam tip: When setting up and , remember that the claim is often placed in the alternative hypothesis unless it’s a statement of “no difference”. Also, always contains the equality ().
Both confidence intervals and hypothesis tests draw conclusions about a population from a single random sample. The next module extends inference to regression – modelling relationships between variables.
Key takeaways
- Confidence intervals give a range; hypothesis tests give a yes/no verdict.
- Each parameter uses a specific sampling distribution (, , ).
- The p‑value quantifies the risk of a Type I error; reject when that risk is small ().
- Correct formulation of hypotheses is the most common pitfall – always ask “What claim is being tested?”