Term 1 · Module 4 of 8

Sampling and Sampling Distributions

Business Statistics for Entrepreneurs

Sampling Process

Sampling is the process of selecting a subset of observations (a sample) from a larger population to make inferences about population parameters (e.g., mean, proportion, standard deviation). The goal is to gather data efficiently and use it to answer research questions about the entire population.

  • Element: an entity for which data is collected.
  • Population: the collection of all elements of interest.
  • Sample: a subset of the population.

Two main uses:

  1. Viren’s election poll: sample of 400 voters → 160 prefer a candidate → estimated population proportion p^=160400=0.40\hat{p} = \frac{160}{400} = 0.40 (40%).
  2. Manoj’s migrant worker study: sample of 120 workers → average yearly earnings ₹36,500 → estimated population mean xˉ=36500\bar{x} = 36500.

These are sample statistics used as estimates of the true population parameters. Because the sample is only a portion, error exists; proper sampling methods provide guarantees on reliability.

Steps in the Sampling Process

  1. Identify the target population – clearly define who (or what) is being studied. Example: Manoj’s “migrant temporary workers” needs boundaries (age, gender, type of work) to avoid vagueness.

  2. Decide the sampling frame – the source or procedure to list all elements of the target population. Example: Viren can use voter registration data; Manoj has no pre-built database of migrant workers, making this step harder.

  3. Determine the sample size – the minimum number of records required.

    • Larger samples improve precision but increase cost and time.
    • Sample size affects confidence level and margin of error.
  4. Choose the sampling method – the technique to select individual cases from the target population using the sampling frame.

    • Broad classification: probabilistic vs. non-probabilistic methods.

Exam tip: The sampling frame determines the accuracy of a study. If the frame is incomplete or biased, even a large sample can produce misleading inferences.

Key takeaways – Sampling Process

  • Sampling is a subset used to estimate population parameters (mean, proportion, etc.).
  • Sample statistics (e.g., xˉ\bar{x}, p^\hat{p}) are not exact; error is inevitable.
  • The four steps are: define target population → choose sampling frame → decide sample size → select sampling method.
  • A poor sampling frame or insufficient sample size undermines reliability.

Sampling Methods

Sampling methods fall into two categories: probabilistic (selection follows a probability distribution) and non-probabilistic (selection does not follow a probability distribution).

Probabilistic Sampling

Simple Random Sampling

  • Every subject in the population has an equal probability of being selected.
  • Best when the population is homogeneous.
  • Execution: label all elements sequentially, then generate uniform random numbers (e.g., in Excel using RAND() or RANDBETWEEN()).
  • Viren’s case: label all 4 crore voters, generate 400 random numbers, and select those entries.

Stratified Sampling

  • Population is divided into mutually exclusive and collectively exhaustive groups called strata (e.g., by age, gender, income).
  • Strata should be as homogeneous as possible; simple random sampling is applied within each stratum.
  • Sample size per stratum should be proportional to the stratum’s size in the population (e.g., if 40% of workers are women, then 40% of the sample should be women).
  • Final sample = combination of all stratum samples.
  • Useful when the population is heterogeneous; stratified sampling can achieve precision comparable to a larger simple random sample.
  • Manoj’s study: strata could be male/female or age groups (e.g., <20, 20–30, 30–40).

Cluster Sampling

  • Population is divided into separate clusters (each element belongs to exactly one cluster).
  • A simple random sample of clusters is taken; all elements within each selected cluster form the sample.
  • Clusters should be heterogeneous (each cluster a mini-representation of the population) – this is the key difference from strata (which are homogeneous).
  • Also called area sampling when clusters are geographic units (e.g., districts, municipalities).
  • Viren: could select a few municipalities and survey all voters there, assuming each municipality is a good replica of Rajasthan.
  • Typically requires a larger sample size than simple random or stratified, but can reduce cost and time because data collection is concentrated in a few clusters.

Exam tip: Strata → homogeneous groups; cluster → heterogeneous groups. The distinction is commonly tested.

Non-Probabilistic Sampling

Systematic Sampling

  • Useful when simple random sampling is time‑consuming (generating random numbers and searching the list).
  • Steps:
    1. Randomly select one element from the first kk elements of the population list (where k=population size/desired sample sizek = \text{population size} / \text{desired sample size}).
    2. Then select every kk‑th element thereafter.
  • Viren: population 4 crore, sample 400 → k=100,000k = 100{,}000; pick a random start among the first 100,000, then every 100,000th element.
  • Because the first selection is random, systematic sampling is often assumed to have properties similar to simple random sampling (though it is non‑probabilistic by strict definition).

Convenience Sampling

  • Elements are selected based on ease of access; no known probability of selection.
  • Example: Manoj surveys migrant workers he has recently interacted with at the plant.
  • Example: Traffic inspectors pulling over the next motorist at an intersection.
  • Advantages: easy and cheap.
  • Disadvantages: impossible to evaluate representativeness; no robust statistical inference possible.

Judgment Sampling

  • A knowledgeable person selects elements they believe are most representative of the population.
  • Example: TV panelists expressing opinions on a political issue; Viren asking a handful of people from a municipality.
  • Quality depends entirely on the selector’s expertise; results must be interpreted with caution.

Voluntary Sampling

  • Data comes from volunteers who choose to participate.
  • Example: Customer survey requests (mostly unhappy customers respond); clinical trial volunteers.
  • Strong self‑selection bias – results are unlikely to represent the general population.

Key takeaways – Sampling Methods

  • Probabilistic: each element has a known probability of selection → allows statistical inference.
    • Simple random: equal chance; best for homogeneous populations.
    • Stratified: homogeneous strata → proportional allocation.
    • Cluster: heterogeneous clusters → sample all elements in selected clusters.
  • Non‑probabilistic: selection is not based on probability → inference is limited.
    • Systematic: easier than simple random; first element random then fixed interval.
    • Convenience: easy but potentially very biased.
    • Judgment: depends on expert’s judgment.
    • Voluntary: severe self‑selection bias.
  • Always prefer probabilistic methods when possible; non‑probabilistic methods may be used for exploratory work or when a sampling frame is unavailable.

Analysis of Sample Data

The goal of collecting a sample is rarely to describe only that sample. Instead, we use sample data to estimate unknown population quantities — the parameters. The most common parameters are the population mean μ\mu, population standard deviation σ\sigma, and population proportion PP. The numbers computed from the sample are called sample statistics; when we use them to guess the parameter, they become point estimates.

Population Parameters vs. Sample Statistics

Population parameter (unknown)Sample statistic (computed from data)Point estimate (numerical value)
Mean μ\muSample mean xˉ=1n∑xi\bar{x} = \frac{1}{n}\sum x_ixˉ\bar{x}
Standard deviation σ\sigmaSample standard deviation s=1n−1∑(xi−xˉ)2s = \sqrt{\frac{1}{n-1}\sum (x_i-\bar{x})^2}ss
Proportion PPSample proportion p^=count in samplesample size\hat{p} = \frac{\text{count in sample}}{\text{sample size}}p^\hat{p}

Note: p^\hat{p} is often written as pp with a caret. Some materials use pˉ\bar{p} for the sample proportion.

Why Point Estimates Are Not Unique

A single survey yields one set of numbers. But if we repeated the sampling process on the same population, we would almost certainly obtain different values of xˉ\bar{x}, ss, and p^\hat{p}. Therefore:

  • xˉ\bar{x} is one of many possible estimates of μ\mu.
  • ss is one of many possible estimates of σ\sigma.
  • p^\hat{p} is one of many possible estimates of PP.

This variability is natural — the sample is only a subset. The difference between a point estimate and the true parameter is called sampling error, and understanding it is the foundation for constructing confidence intervals (covered later).

Worked Example: Chroma Store Satisfaction

Basavaraju collected data from 36 customers (sample size n=36n=36). The key variable is satisfaction score (0–100).

Descriptive statistics from the sample:

StatisticValue
Sample mean xˉ\bar{x}57.31
Sample median59
Sample mode61
Sample range40 – 66 (approx.)
Sample standard deviation ss6.38
Sample proportion scoring ≥60\ge 60p^=17/36≈0.47\hat{p} = 17/36 \approx 0.47

Interpretation as point estimates:

  • xˉ=57.31\bar{x} = 57.31 → point estimate of the population mean μ\mu.
  • s=6.38s = 6.38 → point estimate of the population standard deviation σ\sigma.
  • p^=0.47\hat{p} = 0.47 → point estimate of the population proportion PP of customers who are "reasonably satisfied" (score ≥60\ge 60).

These are our best guesses from this one sample. Other samples would yield different numbers.

Exam tip: Distinguish clearly between the symbol for the statistic (xˉ\bar{x}, ss, p^\hat{p}) and the symbol for the parameter (μ\mu, σ\sigma, PP). The usual notation for a sample proportion is p^\hat{p}; pˉ\bar{p} is also used in some materials.

Key Takeaways

  • Point estimation uses sample statistics to produce a single best guess of an unknown population parameter.
  • Common parameters: mean μ\mu, standard deviation σ\sigma, proportion PP.
  • Common sample statistics: xˉ\bar{x}, ss, p^\hat{p}.
  • Point estimates vary from sample to sample; they are not unique.
  • The computed values from Basavaraju’s data: xˉ=57.31\bar{x}=57.31, s=6.38s=6.38, p^≈0.47\hat{p}\approx0.47 — each is one possible estimate.

Sampling Distribution

A point estimator (like the sample mean xˉ\bar{x}, sample proportion pˉ\bar{p}, or sample standard deviation ss) varies from sample to sample. If we could take many samples from the same population, we would get many different estimates. The probability distribution of a point estimator over all possible samples is called its sampling distribution.

Point Estimators Are Random Variables

  • Selecting a simple random sample is a random experiment.
  • The sample mean xˉ\bar{x} is a numerical outcome of that experiment → xˉ\bar{x} is a random variable.
  • Therefore, xˉ\bar{x} has its own probability distribution (PDF or PMF) with an expected value, variance, and shape.
  • The same holds for pˉ\bar{p} and ss.

A sampling distribution is the distribution of all possible values of a sample statistic.

Properties of the Sampling Distribution of xˉ\bar{x}

1. Expected Value (Mean)

E[xˉ]=μE[\bar{x}] = \mu

  • μ\mu = population mean.
  • When E[estimator]=population parameterE[\text{estimator}] = \text{population parameter}, the estimator is unbiased.
  • Hence xˉ\bar{x} is an unbiased estimator of μ\mu.

2. Standard Deviation (Standard Error)

For an infinite population (or when population size NN is large relative to sample size nn):

σxˉ=σn\sigma_{\bar{x}} = \frac{\sigma}{\sqrt{n}}

  • σ\sigma = population standard deviation.
  • σxˉ\sigma_{\bar{x}} is called the standard error of xˉ\bar{x} — to distinguish it from the population standard deviation.
  • Larger nn → smaller standard error → xˉ\bar{x} tends to be closer to μ\mu.

3. Shape (Form)

  • If the population itself is normally distributed with mean μ\mu and standard deviation σ\sigma, then xˉ\bar{x} is normally distributed with mean μ\mu and standard error σ/n\sigma/\sqrt{n}.
  • (For non‑normal populations, the shape is determined by the Central Limit Theorem, discussed later.)

Worked Example: Basavaraju’s Survey

  • n=36n = 36 customers.
  • Assume (from a previous survey) that the population standard deviation σ=6\sigma = 6.
  • Then the standard error of xˉ\bar{x} is:

σxˉ=636=1\sigma_{\bar{x}} = \frac{6}{\sqrt{36}} = 1

  • Basavaraju’s observed xˉ=57.31\bar{x} = 57.31 is a single draw from a distribution:
    • Centered at μ\mu (unknown)
    • With spread = 1

Exam tip: The standard error σ/n\sigma / \sqrt{n} measures the precision of xˉ\bar{x} as an estimate of μ\mu. It decreases as nn increases — one reason large samples are preferred.

How Repeated Sampling Builds the Concept

In practice we take only one sample, but we know its xˉ\bar{x} comes from a distribution with known expected value and standard error.

Key takeaways

  • xˉ\bar{x} is a random variable; its probability distribution is the sampling distribution of xˉ\bar{x}.
  • E[xˉ]=μE[\bar{x}] = \mu → xˉ\bar{x} is an unbiased estimator of the population mean.
  • The standard error of xˉ\bar{x} is σxˉ=σ/n\sigma_{\bar{x}} = \sigma / \sqrt{n} (population infinite or large).
  • Larger sample size → smaller standard error → more precise estimate.
  • If the population is normal, xˉ\bar{x} is exactly normal; otherwise its shape depends on the Central Limit Theorem (later).

Examples of Central Limit Theorem

The Central Limit Theorem (CLT) states that for a random sample of size nn drawn from any population with mean μ\mu and finite standard deviation σ\sigma, the sampling distribution of the sample mean Xˉ\bar{X} is approximately normal when nn is large (typically n≥30n \ge 30):

Xˉ∼N ⁣(μ,  σXˉ2),σXˉ=σn\bar{X} \sim N\!\left(\mu,\; \sigma_{\bar{X}}^2\right),\qquad \sigma_{\bar{X}} = \frac{\sigma}{\sqrt{n}}

σXˉ\sigma_{\bar{X}} is called the standard error of the mean. The following examples demonstrate how to use the CLT to compute probabilities about Xˉ\bar{X}.


Example 1: Unemployment duration (Kailash Sharma)

  • Population: all unemployed individuals in Rajasthan
  • Known parameters: μ=17.5\mu = 17.5 weeks, σ=4\sigma = 4 weeks
  • Sample size: n=50n = 50
  • Question: probability that the sample mean Xˉ\bar{X} is within 11 week (and 0.50.5 week) of μ\mu

Step 1 – Standard error

σXˉ=450=0.57 weeks\sigma_{\bar{X}} = \frac{4}{\sqrt{50}} = 0.57 \text{ weeks}

Step 2 – Within 1 week

P(∣Xˉ−μ∣<1)=P ⁣(∣Xˉ−μσXˉ∣<10.57)=P(∣Z∣<1.77)P(|\bar{X} - \mu| < 1) = P\!\left( \left|\frac{\bar{X} - \mu}{\sigma_{\bar{X}}}\right| < \frac{1}{0.57} \right) = P(|Z| < 1.77)

From standard normal table: P(∣Z∣<1.77)=0.923P(|Z| < 1.77) = 0.923 (92.3% chance).

Step 3 – Within 0.5 week

P(∣Xˉ−μ∣<0.5)=P ⁣(∣Z∣<0.50.57)=P(∣Z∣<0.88)=0.62P(|\bar{X} - \mu| < 0.5) = P\!\left(|Z| < \frac{0.5}{0.57}\right) = P(|Z| < 0.88) = 0.62 (62% chance).

Exam tip: Dividing the margin of error by the standard error yields the zz-score bound. Always convert the Xˉ\bar{X}-probability problem into a standard normal probability.


Why sample when population parameters are known?

Kailash already knows μ\mu and σ\sigma, but sampling is still valuable: a detailed study on a representative subset allows deeper qualitative analysis (e.g., understanding challenges unemployed individuals face). Ensuring that the sample mean lies close to μ\mu (e.g., within 1 week) confirms that the sample is representative of the population, so conclusions from the detailed study can be generalized.


Example 2: Rainfall in two districts (Shantibai)

DistrictPopulation mean μ\mu (cm)Population σ\sigma (cm)Sample size nnStandard error σXˉ=σ/n\sigma_{\bar{X}} = \sigma/\sqrt{n}
Jaisalmer294304/30≈0.734/\sqrt{30} \approx 0.73
Jhalawar664444/44≈0.604/\sqrt{44} \approx 0.60

Goal: probability that Xˉ\bar{X} is within 1 cm of the respective μ\mu.

  • Jaisalmer: z=1/0.73=1.37z = 1 / 0.73 = 1.37, P(∣Z∣<1.37)=0.83P(|Z| < 1.37) = 0.83 (83%)
  • Jhalawar: z=1/0.60=1.67z = 1 / 0.60 = 1.67, P(∣Z∣<1.67)=0.91P(|Z| < 1.67) = 0.91 (91%)

The larger sample size for Jhalawar gives a smaller standard error and a higher probability that the sample mean will be close to μ\mu.


Example 3: Tax preparation fees (Motilal Oswal)

  • Population: all certified tax consultants in Rajasthan; μ=1800\mu = 1800 rupees, σ=50\sigma = 50 rupees
  • Part A: n=40n = 40 clients
  • Part B: n=75n = 75 clients
  • Question: P(∣Xˉ−μ∣<10)P(|\bar{X} - \mu| < 10)

Part A (n=40)

σXˉ=5040≈7.91,z=107.91≈1.26,P(∣Z∣<1.26)=0.792\sigma_{\bar{X}} = \frac{50}{\sqrt{40}} \approx 7.91,\quad z = \frac{10}{7.91} \approx 1.26,\quad P(|Z| < 1.26) = 0.792

Part B (n=75)

σXˉ=5075≈5.77,z=105.77≈1.73,P(∣Z∣<1.73)=0.916\sigma_{\bar{X}} = \frac{50}{\sqrt{75}} \approx 5.77,\quad z = \frac{10}{5.77} \approx 1.73,\quad P(|Z| < 1.73) = 0.916

Increasing nn reduces the standard error, making the sample mean more precise — the probability of being within 10 rupees jumps from 79.2% to 91.6%.

Remember: The CLT applies only to the sampling distribution of Xˉ\bar{X}, not to the population distribution. The population of fees may not be normal; the normal approximation is for Xˉ\bar{X} because n≥30n \ge 30.


Example 4: Alzheimer’s disease duration (Dr. Sasdev)

  • Population: Alzheimer patients; duration ranges 3–20 years, but distribution is skewed (mean = 8 years, median ≠ midpoint)
  • Known: μ=8\mu = 8 years, σ=4\sigma = 4 years
  • Sample size: n=30n = 30
  • Standard error: σXˉ=4/30≈0.73\sigma_{\bar{X}} = 4/\sqrt{30} \approx 0.73 years

Probabilities

  1. Less than 7 years: z=7−80.73≈−1.37,P(Xˉ<7)=0.08z = \frac{7-8}{0.73} \approx -1.37,\quad P(\bar{X} < 7) = 0.08

  2. Exceeds 7 years: P(Xˉ>7)=1−0.08=0.92P(\bar{X} > 7) = 1 - 0.08 = 0.92

  3. Within 1 year of μ\mu (i.e., 7 to 9 years): P(7<Xˉ<9)=P(−1.37<Z<1.37)=0.83P(7 < \bar{X} < 9) = P(-1.37 < Z < 1.37) = 0.83

Even though the population distribution is right-skewed (mean not at the midpoint of range), the CLT ensures the sampling distribution of Xˉ\bar{X} is approximately normal because n=30n=30.


Key takeaways

  • The CLT allows us to treat Xˉ\bar{X} as normal when n≥30n \ge 30, regardless of the population shape.
  • The standard error σ/n\sigma/\sqrt{n} measures the spread of Xˉ\bar{X}; larger nn yields a smaller standard error and higher confidence that Xˉ\bar{X} is near μ\mu.
  • To find probabilities: transform Xˉ\bar{X} to Z=(Xˉ−μ)/(σ/n)Z = (\bar{X} - \mu) / (\sigma/\sqrt{n}) and use the standard normal table.
  • Sampling can be employed even when population parameters are known — to verify representativeness for a deeper study.
  • Increasing sample size narrows the sampling distribution, increasing the probability of falling within a fixed margin of error.

The Central Limit Theorem (CLT)

The Problem: The population distribution of XX may not be normal (skewed, uniform, arbitrary). Yet we need to know the sampling distribution of the sample mean Xˉ\bar{X} to make probability statements about how close Xˉ\bar{X} is to μ\mu.

The Core Insight (Intuition): Even when the population is wildly non‑normal, the distribution of sample averages becomes approximately normal once the sample size is large enough. The population itself never changes shape — it is the averaging process that smooths things out.

Formal Statement (for Xˉ\bar{X}): When random samples of size nn are drawn from any population with mean μ\mu and standard deviation σ\sigma, the sampling distribution of Xˉ\bar{X} can be approximated by a normal distribution if nn is sufficiently large.

The approximation has:

  • Mean: μXˉ=μ\mu_{\bar{X}} = \mu
  • Standard error: σXˉ=σn\sigma_{\bar{X}} = \dfrac{\sigma}{\sqrt{n}}

Thus, for large nn: Xˉ  ≈  N ⁣(μ,  σn)\bar{X} \;\approx\; N\!\left(\mu,\; \frac{\sigma}{\sqrt{n}}\right)

Visual intuition

The table below summarizes four populations (normal, uniform, skewed, arbitrary) and the sampling distribution of Xˉ\bar{X} for increasing nn:

Populationn=2n=2n=30n=30+
NormalAlready normalNormal
UniformTriangular‑likeNearly normal
SkewedStill skewedNearly normal
ArbitraryIrregularNearly normal

Key: At n=2n=2 the sampling distribution is not normal and not the population shape. At n≥30n\geq 30 it becomes close to normal for most populations.

The magic number n=30n=30

  • The theorem mathematically says: as n→∞n \to \infty, the sampling distribution converges to normal.
  • Empirically, for most applications, n≥30n \ge 30 is “large enough” to use the normal approximation.
  • Exception: Highly skewed populations may need n>150n>150 before the approximation is good. But n=30n=30 works for typical cases.

Exam trap: The CLT is about the distribution of Xˉ\bar{X}, never about the population distribution. Saying “the population becomes normal with large samples” is a common mistake.

What CLT does not say

  • It does not change the population distribution.
  • It does not guarantee Xˉ=μ\bar{X} = \mu — only that the probability of Xˉ\bar{X} being close to μ\mu increases with nn.

Effect of Sample Size on Precision

The standard error of the mean is: σXˉ=σn\sigma_{\bar{X}} = \frac{\sigma}{\sqrt{n}}

  • Doubling nn reduces σXˉ\sigma_{\bar{X}} by a factor of 2\sqrt{2}, not by half.
  • Larger nn → smaller spread → Xˉ\bar{X} is more likely to be near μ\mu.

Intuition: More data yields a more precise estimate of the population mean. The center (μ\mu) stays the same, but the tails shrink.


Worked Example: Basavaraju’s Survey

Basavaraju wants the probability that his sample mean xˉ=57.31\bar{x} = 57.31 (from n=36n=36) is within ±2\pm 2 of the unknown μ\mu. Given: last year’s σ=6\sigma = 6.

Step 1: Verify CLT applies

n=36≥30n=36 \ge 30 → sampling distribution of Xˉ\bar{X} is approximately normal.

Step 2: Compute standard error

σXˉ=636=66=1\sigma_{\bar{X}} = \frac{6}{\sqrt{36}} = \frac{6}{6} = 1

Step 3: Probability statement

We want P(∣Xˉ−μ∣≤2)P(|\bar{X} - \mu| \le 2). Standardize:

Z=Xˉ−μσXˉZ = \frac{\bar{X} - \mu}{\sigma_{\bar{X}}}

P(−2≤Z≤2)=0.9544 (approx 0.95)P(-2 \le Z \le 2) = 0.9544 \ (\text{approx } 0.95)

Result: There is about a 95% chance that the sample mean is within ±2\pm 2 of the population mean. Only 5% chance it falls outside.

What if n=100n=100?

σXˉ=6100=0.6\sigma_{\bar{X}} = \frac{6}{\sqrt{100}} = 0.6

Now P(−2≤Xˉ−μ≤2)P(-2 \le \bar{X}-\mu \le 2) becomes:

P ⁣(−20.6≤Z≤20.6)=P(−3.33≤Z≤3.33)>0.999P\!\left(\frac{-2}{0.6} \le Z \le \frac{2}{0.6}\right) = P(-3.33 \le Z \le 3.33) > 0.999

Result: With n=100n=100, the probability is over 99.9% that Xˉ\bar{X} is within ±2\pm 2 of μ\mu.

Table: Effect of nn on precision (fixed σ=6\sigma=6, range ±2\pm 2)

Sample size nnσXˉ\sigma_{\bar{X}}zz-limitsProbability
361.0±2\pm 20.95
1000.6±3.33\pm 3.330.999

The larger the sample, the higher the chance that Xˉ\bar{X} is close to μ\mu.


Sampling Distribution of the Sample Proportion pˉ\bar{p}

When working with proportions (categorical data), the same CLT logic applies to the sample proportion pˉ\bar{p}.

Intuition: pˉ\bar{p} is just the average of 0/1 indicator variables, so the CLT says its sampling distribution is approximately normal for large nn.

Formal result: If nn is large (typically np≥5np \ge 5 and n(1−p)≥5n(1-p) \ge 5), then:

pˉ  ≈  N ⁣(p,  σpˉ)\bar{p} \;\approx\; N\!\left(p,\; \sigma_{\bar{p}}\right)

where:

  • Mean: μpˉ=p\mu_{\bar{p}} = p (population proportion)
  • Standard error: σpˉ=p(1−p)n\sigma_{\bar{p}} = \sqrt{\dfrac{p(1-p)}{n}}

Example 1: Manohar’s snack survey

  • Population proportion p=0.17p = 0.17 (spend > ₹2500)
  • n=80n=80 → σpˉ=0.17×0.8380=0.042\sigma_{\bar{p}} = \sqrt{\frac{0.17 \times 0.83}{80}} = 0.042
  • Want P(0.12≤pˉ≤0.22)P(0.12 \le \bar{p} \le 0.22) = P(−1.19≤Z≤1.19)≈0.766P(-1.19 \le Z \le 1.19) \approx 0.766 (Using a standard normal table gives approximately 0.766.)

If n=160n=160:

  • σpˉ=0.17×0.83160=0.028\sigma_{\bar{p}} = \sqrt{\frac{0.17\times0.83}{160}} = 0.028
  • P(0.12≤pˉ≤0.22)=P(−1.78≤Z≤1.78)≈0.925P(0.12 \le \bar{p} \le 0.22) = P(-1.78 \le Z \le 1.78) \approx 0.925 (Using a standard normal table gives approximately 0.925.)

Example 2: Sunita’s book adoption rate

  • Population p=0.25p=0.25, sample standard error given as 0.06250.0625.
  • Find nn:

0.0625=0.25×0.75n  ⇒  n=0.1875(0.0625)2=480.0625 = \sqrt{\frac{0.25\times0.75}{n}} \;\Rightarrow\; n = \frac{0.1875}{(0.0625)^2} = 48

  • Want P(pˉ≥0.3)P(\bar{p} \ge 0.3):

Z=0.3−0.250.0625=0.8  ⇒  P(Z≥0.8)=0.2119Z = \frac{0.3-0.25}{0.0625} = 0.8 \;\Rightarrow\; P(Z \ge 0.8) = 0.2119

Result: 21% chance of reaching ≥30% adoption rate.

Example 3: Maliwal’s ghevar preference

  • p=0.75p = 0.75, n=200n=200 → σpˉ=0.75×0.25200=0.03\sigma_{\bar{p}} = \sqrt{\frac{0.75\times0.25}{200}} = 0.03
  1. P(pˉ>0.8)P(\bar{p} > 0.8):

Z=0.8−0.750.03=1.67  ⇒  P(Z>1.67)=0.0475Z = \frac{0.8-0.75}{0.03} = 1.67 \;\Rightarrow\; P(Z>1.67) = 0.0475

  1. P(pˉ<0.6)P(\bar{p} < 0.6):

Z=0.6−0.750.03=−5  ⇒  P(Z<−5)≈0Z = \frac{0.6-0.75}{0.03} = -5 \;\Rightarrow\; P(Z<-5) \approx 0

  1. 95% interval for pˉ\bar{p}:

Using inverse normal: z0.025=−1.96z_{0.025} = -1.96, so:

0.75±1.96×0.03=[0.6912,  0.8088]≈[0.69,0.81]0.75 \pm 1.96 \times 0.03 = [0.6912,\; 0.8088] \approx [0.69, 0.81]

Exam tip: For proportions, always check np≥5np \ge 5 and n(1−p)≥5n(1-p) \ge 5 to justify using the normal approximation.


Key Takeaways

  • CLT: For large nn (≥30\ge 30 in practice), the sampling distribution of Xˉ\bar{X} is approximately normal, regardless of the population shape.
  • Standard error: σXˉ=σ/n\sigma_{\bar{X}} = \sigma / \sqrt{n}; increasing nn reduces spread and improves precision.
  • Proportions: pˉ\bar{p} also follows CLT; σpˉ=p(1−p)/n\sigma_{\bar{p}} = \sqrt{p(1-p)/n}.
  • Population vs. sampling distribution: CLT does not change the population; it describes the averages.
  • Probability interpretation: With large nn, we can compute exact probabilities of Xˉ\bar{X} (or pˉ\bar{p}) being within a given distance of the population parameter.

Sampling Distribution of Proportion

The sample proportion pˉ\bar{p} is a point estimator of the population proportion pp. It is computed as:

pˉ=xn\bar{p} = \frac{x}{n}

where xx = number of elements in the sample that possess the characteristic of interest, and nn = sample size.

Example: In Basavaraj’s customer survey, x=17x = 17 customers gave a satisfaction score ≥ 60 out of n=36n = 36 respondents → pˉ=17/36=0.47\bar{p} = 17/36 = 0.47.

Because pˉ\bar{p} is a random variable, its probability distribution is called the sampling distribution of pˉ\bar{p}.

Expected Value and Unbiasedness

The expected value (mean) of the sampling distribution of pˉ\bar{p} is the population proportion:

E(pˉ)=pE(\bar{p}) = p

Thus pp is the center of the distribution. Since E(pˉ)=pE(\bar{p}) = p, pˉ\bar{p} is an unbiased estimator of pp (analogous to xˉ\bar{x} for μ\mu).

Example: If Basavaraj believes p=0.5p = 0.5, then pˉ=0.47\bar{p} = 0.47 comes from a distribution centered at 0.50.5.

Standard Error of the Proportion

The standard deviation of pˉ\bar{p} is called the standard error of the proportion, denoted σpˉ\sigma_{\bar{p}}:

σpˉ=p(1−p)n\sigma_{\bar{p}} = \sqrt{\frac{p(1-p)}{n}}

This is derived from the binomial distribution of xx (see below). For a finite population, use the finite population correction if NN is not large relative to nn; here we assume a large population.

Example: For p=0.5p = 0.5, n=36n = 36:

σpˉ=0.5×0.536=0.2536=0.0833\sigma_{\bar{p}} = \sqrt{\frac{0.5 \times 0.5}{36}} = \sqrt{\frac{0.25}{36}} = 0.0833

Shape: Normal Approximation

The number of successes xx in a simple random sample from a large population is a binomial random variable: x∼Bin(n,p)x \sim \text{Bin}(n, p). Its mean is npnp and variance np(1−p)np(1-p).

When the sample size is large — specifically, when both np≥5np \ge 5 and n(1−p)≥5n(1-p) \ge 5 — the binomial can be approximated by a normal distribution:

x≈N(np,  np(1−p))x \approx N\bigl(np,\; np(1-p)\bigr)

Dividing by the constant nn leaves pp normal:

pˉ=xn≈N ⁣(p,  p(1−p)n)\bar{p} = \frac{x}{n} \approx N\!\left(p,\; \frac{p(1-p)}{n}\right)

Hence, for large nn, the sampling distribution of pˉ\bar{p} is approximately normal with mean pp and standard error σpˉ=p(1−p)/n\sigma_{\bar{p}} = \sqrt{p(1-p)/n}.

Exam tip: Always verify np≥5np \ge 5 and n(1−p)≥5n(1-p) \ge 5 before using the normal approximation for pˉ\bar{p}. Otherwise the distribution may be skewed.

Applying the Sampling Distribution: Probability Calculations

The normal approximation allows us to compute probabilities about how close pˉ\bar{p} is to pp.

Basavaraj’s Customer Survey

Scenario: Assume p=0.5p = 0.5 (50% of customers score ≥ 60). Question: What is the probability that pˉ\bar{p} from a sample of n=36n = 36 lies within ±0.05\pm 0.05 of pp (i.e., between 0.45 and 0.55)?

  • σpˉ=0.0833\sigma_{\bar{p}} = 0.0833
  • pˉ∼N(0.5,0.0833)\bar{p} \sim N(0.5, 0.0833)
  • P(0.45≤pˉ≤0.55)=0.452P(0.45 \le \bar{p} \le 0.55) = 0.452

If nn increases to 100:

  • σpˉ=0.25/100=0.05\sigma_{\bar{p}} = \sqrt{0.25/100} = 0.05
  • P(0.45≤pˉ≤0.55)=0.683P(0.45 \le \bar{p} \le 0.55) = 0.683

Increasing sample size narrows the standard error → higher probability of being close to pp.

Bikaner Doctors (Dr. Pradak Chandak)

Scenario: National report claims p=0.42p = 0.42 of primary care doctors feel patients receive unnecessary care. A random sample of n=150n = 150 doctors is surveyed in the district.

Compute:

  1. σpˉ=0.42×0.58150=0.04\sigma_{\bar{p}} = \sqrt{\frac{0.42 \times 0.58}{150}} = 0.04
  2. pˉ∼N(0.42,0.04)\bar{p} \sim N(0.42, 0.04)

Probability that pˉ\bar{p} is within ±0.03\pm 0.03 of pp (i.e., between 0.39 and 0.45): P(0.39≤pˉ≤0.45)=0.547P(0.39 \le \bar{p} \le 0.45) = 0.547

Probability that pˉ>0.5\bar{p} > 0.5: P(pˉ>0.5)=0.023(tail area)P(\bar{p} > 0.5) = 0.023 \quad (\text{tail area})

Summary: Properties of the Sampling Distribution of pˉ\bar{p}

PropertyExpression
Mean (expected value)E(pˉ)=pE(\bar{p}) = p
Standard errorσpˉ=p(1−p)n\sigma_{\bar{p}} = \sqrt{\dfrac{p(1-p)}{n}}
Shape (large nn)Approximately normal if np≥5np \ge 5 and n(1−p)≥5n(1-p) \ge 5
Sourcex∼Binomial(n,p)x \sim \text{Binomial}(n,p) → divide by nn

Key takeaways

  • pˉ=x/n\bar{p} = x/n is the point estimator of pp; it is unbiased because E(pˉ)=pE(\bar{p}) = p.
  • The standard error σpˉ=p(1−p)/n\sigma_{\bar{p}} = \sqrt{p(1-p)/n} measures the spread of pˉ\bar{p}.
  • For large samples (np≥5np \ge 5, n(1−p)≥5n(1-p) \ge 5), pˉ\bar{p} is approximately normal.
  • Probability calculations about the difference between pˉ\bar{p} and pp use the normal distribution.
  • Larger sample sizes → smaller standard error → higher probability that pˉ\bar{p} is close to pp.

Sample Variance and Degrees of Freedom

The sample variance s2s^2 measures spread in a sample. For a random sample x1,x2,…,xnx_1, x_2, \dots, x_n from a population with sample mean xˉ\bar{x}:

s2=1n−1∑i=1n(xi−xˉ)2s^2 = \frac{1}{n-1} \sum_{i=1}^{n} (x_i - \bar{x})^2

The denominator is n−1n-1, not nn, because one degree of freedom (DOF) is lost when estimating the mean. If the sample mean is fixed, only n−1n-1 of the observations can vary freely; the last is determined.

Intuition: Given three numbers with mean 10 and two numbers are 1 and 2, the third must be 30−1−2=2730-1-2=27. Only two numbers were free → n−1=2n-1=2 DOF.

The Chi‑Square Connection for Sample Variance

The sample variance s2s^2 is a random variable. Unlike the sample mean xˉ\bar{x}, there is no Central Limit Theorem for s2s^2 — its distribution is unknown without assuming the population is normal.

Key result: If the population is normally distributed with variance σ2\sigma^2, then:

(n−1)s2σ2∼χ(n−1)2\frac{(n-1)s^2}{\sigma^2} \sim \chi^2_{(n-1)}

i.e., it follows a chi‑square distribution with n−1n-1 degrees of freedom. The chi‑square distribution is:

  • Non‑negative and right‑skewed (for small df)
  • Becomes more symmetric and bell‑shaped as df increases

This does not directly give the distribution of s2s^2 alone, but it allows us to make probability statements about σ2\sigma^2 and σ\sigma using the sample value s2s^2.

Inference on Population Variance and Standard Deviation

Given one observed s2s^2 from a normal population, the chi‑square relation lets us bound σ2\sigma^2 with a given probability. For example, if ()(\sqrt{ } ) for a chi‑square with n−1n-1 df the lower percentile is χα2\chi^2_{\alpha} and the upper percentile is χ1−α2\chi^2_{1-\alpha}, then:

P((n−1)s2χ1−α/22<σ2<(n−1)s2χα/22)=1−αP\left( \frac{(n-1)s^2}{\chi^2_{1-\alpha/2}} < \sigma^2 < \frac{(n-1)s^2}{\chi^2_{\alpha/2}} \right) = 1-\alpha


Basaraju Example (n=36, s² = 40.73, s = 6.38)

Assume normality. Then 35⋅s2σ2∼χ(35)2\frac{35 \cdot s^2}{\sigma^2} \sim \chi^2_{(35)}.

  • 90% of the time χ352<46.06\chi^2_{35} < 46.06 → σ2>35⋅40.7346.06≈30.92\sigma^2 > \frac{35 \cdot 40.73}{46.06} \approx 30.92 → σ>30.92≈5.56\sigma > \sqrt{30.92} \approx 5.56.
  • 90% of the time χ352>24.8\chi^2_{35} > 24.8 → σ2<35⋅40.7324.8≈57.44\sigma^2 < \frac{35 \cdot 40.73}{24.8} \approx 57.44 → σ<57.44≈7.58\sigma < \sqrt{57.44} \approx 7.58.

Thus, based on this single sample, the population standard deviation is between approximately 5.56 and 7.58 with roughly 90% confidence (using the two one‑sided 90% statements together).


Worked Examples

Kailash Sharma (unemployment data)

  • Population: μ=17.5\mu = 17.5 weeks, σ=4\sigma = 4 weeks.
  • Sample size n=50n = 50. Assume normality.
  • Random variable: 49⋅s216∼χ(49)2\frac{49 \cdot s^2}{16} \sim \chi^2_{(49)}.

1. Probability s2>20s^2 > 20

P(s2>20)=P(χ492>49⋅2016=61.25)=0.113P(s^2 > 20) = P\left( \chi^2_{49} > \frac{49 \cdot 20}{16} = 61.25 \right) = 0.113 (from chi‑square template).

2. Probability s<3s < 3

s<3  ⟺  s2<9s < 3 \iff s^2 < 9.

P(s<3)=P(χ492<49⋅916=27.5625)=0.006P(s < 3) = P\left( \chi^2_{49} < \frac{49 \cdot 9}{16} = 27.5625 \right) = 0.006.


Motilal Oswal (tax preparation fees)

  • Population: μ=1800\mu = 1800 rupees, σ=50\sigma = 50 rupees.
  • Sample size n=40n = 40. Assume normality.
  • Distribution: 39⋅s22500∼χ(39)2\frac{39 \cdot s^2}{2500} \sim \chi^2_{(39)}.

1. Probability s>60s > 60

s>60  ⟺  s2>3600s > 60 \iff s^2 > 3600.

P(s>60)=P(χ392>39⋅36002500=56.16)=0.037P(s > 60) = P\left( \chi^2_{39} > \frac{39 \cdot 3600}{2500} = 56.16 \right) = 0.037.

2. Claim: 90% chance s<55s < 55

s<55  ⟺  s2<3025s < 55 \iff s^2 < 3025.

P(s<55)=P(χ392<39⋅30252500=47.19)=0.827P(s < 55) = P\left( \chi^2_{39} < \frac{39 \cdot 3025}{2500} = 47.19 \right) = 0.827.

The claim is false — the actual probability is about 82.7%, not 90%.


Dr. Naveen Sachdev (Alzheimer's survival time)

  • Population: μ=8\mu = 8 years, σ=4\sigma = 4 years. Population is skewed (range 3–20 years, not normal).
  • Sample size n=30n = 30. Normality assumption is violated — use with caution.
  • Distribution under normality: 29⋅s216∼χ(29)2\frac{29 \cdot s^2}{16} \sim \chi^2_{(29)}.

1. Probability s<3s < 3

s<3  ⟺  s2<9s < 3 \iff s^2 < 9.

P(s<3)=P(χ292<29⋅916=16.3125)=0.028P(s < 3) = P\left( \chi^2_{29} < \frac{29 \cdot 9}{16} = 16.3125 \right) = 0.028.

2. Probability 3.5<s<4.53.5 < s < 4.5

s∈[3.5,4.5]  ⟺  s2∈[12.25,20.25]s \in [3.5, 4.5] \iff s^2 \in [12.25, 20.25].

P(3.5<s<4.5)=P(22.20<χ292<36.70)=0.658P(3.5 < s < 4.5) = P\left( 22.20 < \chi^2_{29} < 36.70 \right) = 0.658.

Exam tip: The chi‑square result for sample variance requires the population to be normally distributed. For skewed populations (like Dr. Sachdev’s), the calculated probabilities are approximate at best. A larger sample may help, but no CLT equivalent exists for variance.


Caution: Normality Assumption

All of the above inferences rely on the assumption that the population is normally distributed. If the population is not normal (e.g., skewed, like the Alzheimer's data), the distribution of (n−1)s2σ2\frac{(n-1)s^2}{\sigma^2} is not exactly chi‑square. In practice:

  • For moderate to large samples, the chi‑square approximation may still be reasonable if the population is not too non‑normal.
  • Dr. Sachdev’s example illustrates a case where the assumption is clearly violated, so the computed probabilities (2.8%, 65.8%) should be interpreted as approximate, not exact.

Key takeaways

  • Sample variance uses denominator n−1n-1 to account for the loss of one degree of freedom.
  • If the population is normal, (n−1)s2σ2∼χ(n−1)2\frac{(n-1)s^2}{\sigma^2} \sim \chi^2_{(n-1)}.
  • This relation allows probability statements about the population variance and standard deviation without having the direct distribution of s2s^2.
  • Use chi‑square tables or templates to find probabilities and thresholds.
  • The normality assumption is critical; without it, results are unreliable, especially for small samples.

Properties of Point Estimators

A point estimator is a sample statistic (e.g., xˉ\bar{x}, pˉ\bar{p}, ss) used to estimate a population parameter (e.g., μ\mu, pp, σ\sigma). Not all sample statistics make good estimators; four desirable properties distinguish excellent estimators.

Let θ\theta denote the population parameter of interest and θˉ\bar{\theta} the point estimator (sample statistic) for θ\theta.

Unbiasedness

An estimator θˉ\bar{\theta} is unbiased if its expected value equals the population parameter:

E(θˉ)=θE(\bar{\theta}) = \theta

  • For unbiased estimators, the mean of the sampling distribution is centred exactly on θ\theta; over many samples the over- and under-estimates balance out.
  • Bias = E(θˉ)−θE(\bar{\theta}) - \theta. A biased estimator systematically over- or under-estimates θ\theta.
Unbiased estimatorPopulation parameterReason
xˉ\bar{x}μ\muE(xˉ)=μE(\bar{x}) = \mu
pˉ\bar{p}ppE(pˉ)=pE(\bar{p}) = p
s2s^2σ2\sigma^2E(s2)=σ2E(s^2) = \sigma^2 (uses n−1n-1 in denominator to correct bias)

Exam tip: The sample variance s2s^2 uses n−1n-1 exactly to make it an unbiased estimator of σ2\sigma^2. The sample standard deviation ss is not unbiased, but bias is small for moderate nn.

Efficiency

When two unbiased estimators exist for the same θ\theta, the one with smaller standard error is more efficient. A smaller standard error means values cluster more tightly around θ\theta.

  • Example: For a normal population, SE(xˉ)<SE(sample median)\text{SE}(\bar{x}) < \text{SE}(\text{sample median}), so sample mean is more efficient than sample median as an estimator of μ\mu.

Consistency

A consistent estimator improves as sample size grows: larger nn → estimates cluster closer to θ\theta.

  • Formally: lim⁡n→∞P(∣θˉ−θ∣<ϵ)=1\lim_{n\to\infty} P(|\bar{\theta} - \theta| < \epsilon) = 1 for any ϵ>0\epsilon > 0.
  • This follows from standard error shrinking with nn:
    • SE(xˉ)=σ/n\text{SE}(\bar{x}) = \sigma / \sqrt{n} → n↑n\uparrow → SE ↓\downarrow
    • SE(pˉ)=p(1−p)/n\text{SE}(\bar{p}) = \sqrt{p(1-p)/n} → n↑n\uparrow → SE ↓\downarrow

Sufficiency

A sufficient estimator uses all available data points in the sample.

  • Both xˉ\bar{x} and pˉ\bar{p} are sufficient: they sum or count every observation.
  • Sample median is not sufficient because it ignores the magnitude of values away from the centre.

Key takeaways

  • Unbiasedness: E(θˉ)=θE(\bar{\theta}) = \theta; bias can sometimes be corrected (e.g., s2s^2 with n−1n-1).
  • Efficiency: prefer unbiased estimators with smaller standard error.
  • Consistency: larger samples give more precise estimates.
  • Sufficiency: use every data point in the estimate.
  • Sample mean xˉ\bar{x} and sample proportion pˉ\bar{p} satisfy all four properties.

Sampling Error

The inevitable deviation of a single sample from the population due to random chance. It is unavoidable when taking a probability sample; we quantify it with the standard error of the estimator. Increasing sample size reduces sampling error.

Non‑sampling Errors

Errors that arise from causes unrelated to random sampling. They introduce bias and can mislead decisions regardless of sample size.

TypeDescriptionExample (Basaraju's customer survey)
Coverage errorThe sampled population does not match the target population.Surveying only current customers ignores competitors’ customers and non‑customers — the real growth opportunity.
Non‑response errorSome segments are over‑ or under‑represented due to differing response rates.In‑store surveys miss online customers; online surveys miss in‑store customers. Response rates may not reflect actual purchasing proportions.
Measurement errorResponses are inaccurate due to question design, interviewer influence, or respondent dishonesty.Non‑anonymous forms; staff watching customers fill forms; confusing online questions; rushed or dishonest answers.

Reducing Non‑sampling Errors

  1. Define the target population precisely before drawing the sample.
  2. Design the data collection process carefully and train collectors.
  3. Pre‑test the data collection procedure (pilot study) to identify and fix issues.
  4. Choose an appropriate sampling method (stratified, cluster, systematic) to ensure key variables are reflected in the sample.

Key takeaways

  • Sampling error is random and manageable via sample size; non‑sampling errors are systematic.
  • Three major non‑sampling errors: coverage, non‑response, measurement.
  • Non‑sampling errors cannot be fixed by increasing sample size — they require careful study design.
  • Good point estimators quantify sampling error (standard error); reducing non‑sampling errors is a separate, critical step for valid inference.