Sampling Process
Sampling is the process of selecting a subset of observations (a sample) from a larger population to make inferences about population parameters (e.g., mean, proportion, standard deviation). The goal is to gather data efficiently and use it to answer research questions about the entire population.
- Element: an entity for which data is collected.
- Population: the collection of all elements of interest.
- Sample: a subset of the population.
Two main uses:
- Viren’s election poll: sample of 400 voters → 160 prefer a candidate → estimated population proportion (40%).
- Manoj’s migrant worker study: sample of 120 workers → average yearly earnings ₹36,500 → estimated population mean .
These are sample statistics used as estimates of the true population parameters. Because the sample is only a portion, error exists; proper sampling methods provide guarantees on reliability.
Steps in the Sampling Process
-
Identify the target population – clearly define who (or what) is being studied. Example: Manoj’s “migrant temporary workers” needs boundaries (age, gender, type of work) to avoid vagueness.
-
Decide the sampling frame – the source or procedure to list all elements of the target population. Example: Viren can use voter registration data; Manoj has no pre-built database of migrant workers, making this step harder.
-
Determine the sample size – the minimum number of records required.
- Larger samples improve precision but increase cost and time.
- Sample size affects confidence level and margin of error.
-
Choose the sampling method – the technique to select individual cases from the target population using the sampling frame.
- Broad classification: probabilistic vs. non-probabilistic methods.
Exam tip: The sampling frame determines the accuracy of a study. If the frame is incomplete or biased, even a large sample can produce misleading inferences.
Key takeaways – Sampling Process
- Sampling is a subset used to estimate population parameters (mean, proportion, etc.).
- Sample statistics (e.g., , ) are not exact; error is inevitable.
- The four steps are: define target population → choose sampling frame → decide sample size → select sampling method.
- A poor sampling frame or insufficient sample size undermines reliability.
Sampling Methods
Sampling methods fall into two categories: probabilistic (selection follows a probability distribution) and non-probabilistic (selection does not follow a probability distribution).
Probabilistic Sampling
Simple Random Sampling
- Every subject in the population has an equal probability of being selected.
- Best when the population is homogeneous.
- Execution: label all elements sequentially, then generate uniform random numbers (e.g., in Excel using
RAND()orRANDBETWEEN()). - Viren’s case: label all 4 crore voters, generate 400 random numbers, and select those entries.
Stratified Sampling
- Population is divided into mutually exclusive and collectively exhaustive groups called strata (e.g., by age, gender, income).
- Strata should be as homogeneous as possible; simple random sampling is applied within each stratum.
- Sample size per stratum should be proportional to the stratum’s size in the population (e.g., if 40% of workers are women, then 40% of the sample should be women).
- Final sample = combination of all stratum samples.
- Useful when the population is heterogeneous; stratified sampling can achieve precision comparable to a larger simple random sample.
- Manoj’s study: strata could be male/female or age groups (e.g., <20, 20–30, 30–40).
Cluster Sampling
- Population is divided into separate clusters (each element belongs to exactly one cluster).
- A simple random sample of clusters is taken; all elements within each selected cluster form the sample.
- Clusters should be heterogeneous (each cluster a mini-representation of the population) – this is the key difference from strata (which are homogeneous).
- Also called area sampling when clusters are geographic units (e.g., districts, municipalities).
- Viren: could select a few municipalities and survey all voters there, assuming each municipality is a good replica of Rajasthan.
- Typically requires a larger sample size than simple random or stratified, but can reduce cost and time because data collection is concentrated in a few clusters.
Exam tip: Strata → homogeneous groups; cluster → heterogeneous groups. The distinction is commonly tested.
Non-Probabilistic Sampling
Systematic Sampling
- Useful when simple random sampling is time‑consuming (generating random numbers and searching the list).
- Steps:
- Randomly select one element from the first elements of the population list (where ).
- Then select every ‑th element thereafter.
- Viren: population 4 crore, sample 400 → ; pick a random start among the first 100,000, then every 100,000th element.
- Because the first selection is random, systematic sampling is often assumed to have properties similar to simple random sampling (though it is non‑probabilistic by strict definition).
Convenience Sampling
- Elements are selected based on ease of access; no known probability of selection.
- Example: Manoj surveys migrant workers he has recently interacted with at the plant.
- Example: Traffic inspectors pulling over the next motorist at an intersection.
- Advantages: easy and cheap.
- Disadvantages: impossible to evaluate representativeness; no robust statistical inference possible.
Judgment Sampling
- A knowledgeable person selects elements they believe are most representative of the population.
- Example: TV panelists expressing opinions on a political issue; Viren asking a handful of people from a municipality.
- Quality depends entirely on the selector’s expertise; results must be interpreted with caution.
Voluntary Sampling
- Data comes from volunteers who choose to participate.
- Example: Customer survey requests (mostly unhappy customers respond); clinical trial volunteers.
- Strong self‑selection bias – results are unlikely to represent the general population.
Key takeaways – Sampling Methods
- Probabilistic: each element has a known probability of selection → allows statistical inference.
- Simple random: equal chance; best for homogeneous populations.
- Stratified: homogeneous strata → proportional allocation.
- Cluster: heterogeneous clusters → sample all elements in selected clusters.
- Non‑probabilistic: selection is not based on probability → inference is limited.
- Systematic: easier than simple random; first element random then fixed interval.
- Convenience: easy but potentially very biased.
- Judgment: depends on expert’s judgment.
- Voluntary: severe self‑selection bias.
- Always prefer probabilistic methods when possible; non‑probabilistic methods may be used for exploratory work or when a sampling frame is unavailable.
Analysis of Sample Data
The goal of collecting a sample is rarely to describe only that sample. Instead, we use sample data to estimate unknown population quantities — the parameters. The most common parameters are the population mean , population standard deviation , and population proportion . The numbers computed from the sample are called sample statistics; when we use them to guess the parameter, they become point estimates.
Population Parameters vs. Sample Statistics
| Population parameter (unknown) | Sample statistic (computed from data) | Point estimate (numerical value) |
|---|---|---|
| Mean | Sample mean | |
| Standard deviation | Sample standard deviation | |
| Proportion | Sample proportion |
Note: is often written as with a caret. Some materials use for the sample proportion.
Why Point Estimates Are Not Unique
A single survey yields one set of numbers. But if we repeated the sampling process on the same population, we would almost certainly obtain different values of , , and . Therefore:
- is one of many possible estimates of .
- is one of many possible estimates of .
- is one of many possible estimates of .
This variability is natural — the sample is only a subset. The difference between a point estimate and the true parameter is called sampling error, and understanding it is the foundation for constructing confidence intervals (covered later).
Worked Example: Chroma Store Satisfaction
Basavaraju collected data from 36 customers (sample size ). The key variable is satisfaction score (0–100).
Descriptive statistics from the sample:
| Statistic | Value |
|---|---|
| Sample mean | 57.31 |
| Sample median | 59 |
| Sample mode | 61 |
| Sample range | 40 – 66 (approx.) |
| Sample standard deviation | 6.38 |
| Sample proportion scoring |
Interpretation as point estimates:
- → point estimate of the population mean .
- → point estimate of the population standard deviation .
- → point estimate of the population proportion of customers who are "reasonably satisfied" (score ).
These are our best guesses from this one sample. Other samples would yield different numbers.
Exam tip: Distinguish clearly between the symbol for the statistic (, , ) and the symbol for the parameter (, , ). The usual notation for a sample proportion is ; is also used in some materials.
Key Takeaways
- Point estimation uses sample statistics to produce a single best guess of an unknown population parameter.
- Common parameters: mean , standard deviation , proportion .
- Common sample statistics: , , .
- Point estimates vary from sample to sample; they are not unique.
- The computed values from Basavaraju’s data: , , — each is one possible estimate.
Sampling Distribution
A point estimator (like the sample mean , sample proportion , or sample standard deviation ) varies from sample to sample. If we could take many samples from the same population, we would get many different estimates. The probability distribution of a point estimator over all possible samples is called its sampling distribution.
Point Estimators Are Random Variables
- Selecting a simple random sample is a random experiment.
- The sample mean is a numerical outcome of that experiment → is a random variable.
- Therefore, has its own probability distribution (PDF or PMF) with an expected value, variance, and shape.
- The same holds for and .
A sampling distribution is the distribution of all possible values of a sample statistic.
Properties of the Sampling Distribution of
1. Expected Value (Mean)
- = population mean.
- When , the estimator is unbiased.
- Hence is an unbiased estimator of .
2. Standard Deviation (Standard Error)
For an infinite population (or when population size is large relative to sample size ):
- = population standard deviation.
- is called the standard error of — to distinguish it from the population standard deviation.
- Larger → smaller standard error → tends to be closer to .
3. Shape (Form)
- If the population itself is normally distributed with mean and standard deviation , then is normally distributed with mean and standard error .
- (For non‑normal populations, the shape is determined by the Central Limit Theorem, discussed later.)
Worked Example: Basavaraju’s Survey
- customers.
- Assume (from a previous survey) that the population standard deviation .
- Then the standard error of is:
- Basavaraju’s observed is a single draw from a distribution:
- Centered at (unknown)
- With spread = 1
Exam tip: The standard error measures the precision of as an estimate of . It decreases as increases — one reason large samples are preferred.
How Repeated Sampling Builds the Concept
In practice we take only one sample, but we know its comes from a distribution with known expected value and standard error.
Key takeaways
- is a random variable; its probability distribution is the sampling distribution of .
- → is an unbiased estimator of the population mean.
- The standard error of is (population infinite or large).
- Larger sample size → smaller standard error → more precise estimate.
- If the population is normal, is exactly normal; otherwise its shape depends on the Central Limit Theorem (later).
Examples of Central Limit Theorem
The Central Limit Theorem (CLT) states that for a random sample of size drawn from any population with mean and finite standard deviation , the sampling distribution of the sample mean is approximately normal when is large (typically ):
is called the standard error of the mean. The following examples demonstrate how to use the CLT to compute probabilities about .
Example 1: Unemployment duration (Kailash Sharma)
- Population: all unemployed individuals in Rajasthan
- Known parameters: weeks, weeks
- Sample size:
- Question: probability that the sample mean is within week (and week) of
Step 1 – Standard error
Step 2 – Within 1 week
From standard normal table: (92.3% chance).
Step 3 – Within 0.5 week
(62% chance).
Exam tip: Dividing the margin of error by the standard error yields the -score bound. Always convert the -probability problem into a standard normal probability.
Why sample when population parameters are known?
Kailash already knows and , but sampling is still valuable: a detailed study on a representative subset allows deeper qualitative analysis (e.g., understanding challenges unemployed individuals face). Ensuring that the sample mean lies close to (e.g., within 1 week) confirms that the sample is representative of the population, so conclusions from the detailed study can be generalized.
Example 2: Rainfall in two districts (Shantibai)
| District | Population mean (cm) | Population (cm) | Sample size | Standard error |
|---|---|---|---|---|
| Jaisalmer | 29 | 4 | 30 | |
| Jhalawar | 66 | 4 | 44 |
Goal: probability that is within 1 cm of the respective .
- Jaisalmer: , (83%)
- Jhalawar: , (91%)
The larger sample size for Jhalawar gives a smaller standard error and a higher probability that the sample mean will be close to .
Example 3: Tax preparation fees (Motilal Oswal)
- Population: all certified tax consultants in Rajasthan; rupees, rupees
- Part A: clients
- Part B: clients
- Question:
Part A (n=40)
Part B (n=75)
Increasing reduces the standard error, making the sample mean more precise — the probability of being within 10 rupees jumps from 79.2% to 91.6%.
Remember: The CLT applies only to the sampling distribution of , not to the population distribution. The population of fees may not be normal; the normal approximation is for because .
Example 4: Alzheimer’s disease duration (Dr. Sasdev)
- Population: Alzheimer patients; duration ranges 3–20 years, but distribution is skewed (mean = 8 years, median ≠ midpoint)
- Known: years, years
- Sample size:
- Standard error: years
Probabilities
-
Less than 7 years:
-
Exceeds 7 years:
-
Within 1 year of (i.e., 7 to 9 years):
Even though the population distribution is right-skewed (mean not at the midpoint of range), the CLT ensures the sampling distribution of is approximately normal because .
Key takeaways
- The CLT allows us to treat as normal when , regardless of the population shape.
- The standard error measures the spread of ; larger yields a smaller standard error and higher confidence that is near .
- To find probabilities: transform to and use the standard normal table.
- Sampling can be employed even when population parameters are known — to verify representativeness for a deeper study.
- Increasing sample size narrows the sampling distribution, increasing the probability of falling within a fixed margin of error.
The Central Limit Theorem (CLT)
The Problem: The population distribution of may not be normal (skewed, uniform, arbitrary). Yet we need to know the sampling distribution of the sample mean to make probability statements about how close is to .
The Core Insight (Intuition): Even when the population is wildly non‑normal, the distribution of sample averages becomes approximately normal once the sample size is large enough. The population itself never changes shape — it is the averaging process that smooths things out.
Formal Statement (for ): When random samples of size are drawn from any population with mean and standard deviation , the sampling distribution of can be approximated by a normal distribution if is sufficiently large.
The approximation has:
- Mean:
- Standard error:
Thus, for large :
Visual intuition
The table below summarizes four populations (normal, uniform, skewed, arbitrary) and the sampling distribution of for increasing :
| Population | + | |
|---|---|---|
| Normal | Already normal | Normal |
| Uniform | Triangular‑like | Nearly normal |
| Skewed | Still skewed | Nearly normal |
| Arbitrary | Irregular | Nearly normal |
Key: At the sampling distribution is not normal and not the population shape. At it becomes close to normal for most populations.
The magic number
- The theorem mathematically says: as , the sampling distribution converges to normal.
- Empirically, for most applications, is “large enough” to use the normal approximation.
- Exception: Highly skewed populations may need before the approximation is good. But works for typical cases.
Exam trap: The CLT is about the distribution of , never about the population distribution. Saying “the population becomes normal with large samples” is a common mistake.
What CLT does not say
- It does not change the population distribution.
- It does not guarantee — only that the probability of being close to increases with .
Effect of Sample Size on Precision
The standard error of the mean is:
- Doubling reduces by a factor of , not by half.
- Larger → smaller spread → is more likely to be near .
Intuition: More data yields a more precise estimate of the population mean. The center () stays the same, but the tails shrink.
Worked Example: Basavaraju’s Survey
Basavaraju wants the probability that his sample mean (from ) is within of the unknown . Given: last year’s .
Step 1: Verify CLT applies
→ sampling distribution of is approximately normal.
Step 2: Compute standard error
Step 3: Probability statement
We want . Standardize:
Result: There is about a 95% chance that the sample mean is within of the population mean. Only 5% chance it falls outside.
What if ?
Now becomes:
Result: With , the probability is over 99.9% that is within of .
Table: Effect of on precision (fixed , range )
| Sample size | -limits | Probability | |
|---|---|---|---|
| 36 | 1.0 | 0.95 | |
| 100 | 0.6 | 0.999 |
The larger the sample, the higher the chance that is close to .
Sampling Distribution of the Sample Proportion
When working with proportions (categorical data), the same CLT logic applies to the sample proportion .
Intuition: is just the average of 0/1 indicator variables, so the CLT says its sampling distribution is approximately normal for large .
Formal result: If is large (typically and ), then:
where:
- Mean: (population proportion)
- Standard error:
Example 1: Manohar’s snack survey
- Population proportion (spend > ₹2500)
- →
- Want = (Using a standard normal table gives approximately 0.766.)
If :
- (Using a standard normal table gives approximately 0.925.)
Example 2: Sunita’s book adoption rate
- Population , sample standard error given as .
- Find :
- Want :
Result: 21% chance of reaching ≥30% adoption rate.
Example 3: Maliwal’s ghevar preference
- , →
- :
- :
- 95% interval for :
Using inverse normal: , so:
Exam tip: For proportions, always check and to justify using the normal approximation.
Key Takeaways
- CLT: For large ( in practice), the sampling distribution of is approximately normal, regardless of the population shape.
- Standard error: ; increasing reduces spread and improves precision.
- Proportions: also follows CLT; .
- Population vs. sampling distribution: CLT does not change the population; it describes the averages.
- Probability interpretation: With large , we can compute exact probabilities of (or ) being within a given distance of the population parameter.
Sampling Distribution of Proportion
The sample proportion is a point estimator of the population proportion . It is computed as:
where = number of elements in the sample that possess the characteristic of interest, and = sample size.
Example: In Basavaraj’s customer survey, customers gave a satisfaction score ≥ 60 out of respondents → .
Because is a random variable, its probability distribution is called the sampling distribution of .
Expected Value and Unbiasedness
The expected value (mean) of the sampling distribution of is the population proportion:
Thus is the center of the distribution. Since , is an unbiased estimator of (analogous to for ).
Example: If Basavaraj believes , then comes from a distribution centered at .
Standard Error of the Proportion
The standard deviation of is called the standard error of the proportion, denoted :
This is derived from the binomial distribution of (see below). For a finite population, use the finite population correction if is not large relative to ; here we assume a large population.
Example: For , :
Shape: Normal Approximation
The number of successes in a simple random sample from a large population is a binomial random variable: . Its mean is and variance .
When the sample size is large — specifically, when both and — the binomial can be approximated by a normal distribution:
Dividing by the constant leaves normal:
Hence, for large , the sampling distribution of is approximately normal with mean and standard error .
Exam tip: Always verify and before using the normal approximation for . Otherwise the distribution may be skewed.
Applying the Sampling Distribution: Probability Calculations
The normal approximation allows us to compute probabilities about how close is to .
Basavaraj’s Customer Survey
Scenario: Assume (50% of customers score ≥ 60). Question: What is the probability that from a sample of lies within of (i.e., between 0.45 and 0.55)?
If increases to 100:
Increasing sample size narrows the standard error → higher probability of being close to .
Bikaner Doctors (Dr. Pradak Chandak)
Scenario: National report claims of primary care doctors feel patients receive unnecessary care. A random sample of doctors is surveyed in the district.
Compute:
Probability that is within of (i.e., between 0.39 and 0.45):
Probability that :
Summary: Properties of the Sampling Distribution of
| Property | Expression |
|---|---|
| Mean (expected value) | |
| Standard error | |
| Shape (large ) | Approximately normal if and |
| Source | → divide by |
Key takeaways
- is the point estimator of ; it is unbiased because .
- The standard error measures the spread of .
- For large samples (, ), is approximately normal.
- Probability calculations about the difference between and use the normal distribution.
- Larger sample sizes → smaller standard error → higher probability that is close to .
Sample Variance and Degrees of Freedom
The sample variance measures spread in a sample. For a random sample from a population with sample mean :
The denominator is , not , because one degree of freedom (DOF) is lost when estimating the mean. If the sample mean is fixed, only of the observations can vary freely; the last is determined.
Intuition: Given three numbers with mean 10 and two numbers are 1 and 2, the third must be . Only two numbers were free → DOF.
The Chi‑Square Connection for Sample Variance
The sample variance is a random variable. Unlike the sample mean , there is no Central Limit Theorem for — its distribution is unknown without assuming the population is normal.
Key result: If the population is normally distributed with variance , then:
i.e., it follows a chi‑square distribution with degrees of freedom. The chi‑square distribution is:
- Non‑negative and right‑skewed (for small df)
- Becomes more symmetric and bell‑shaped as df increases
This does not directly give the distribution of alone, but it allows us to make probability statements about and using the sample value .
Inference on Population Variance and Standard Deviation
Given one observed from a normal population, the chi‑square relation lets us bound with a given probability. For example, if for a chi‑square with df the lower percentile is and the upper percentile is , then:
Basaraju Example (n=36, s² = 40.73, s = 6.38)
Assume normality. Then .
- 90% of the time → → .
- 90% of the time → → .
Thus, based on this single sample, the population standard deviation is between approximately 5.56 and 7.58 with roughly 90% confidence (using the two one‑sided 90% statements together).
Worked Examples
Kailash Sharma (unemployment data)
- Population: weeks, weeks.
- Sample size . Assume normality.
- Random variable: .
1. Probability
(from chi‑square template).
2. Probability
.
.
Motilal Oswal (tax preparation fees)
- Population: rupees, rupees.
- Sample size . Assume normality.
- Distribution: .
1. Probability
.
.
2. Claim: 90% chance
.
.
The claim is false — the actual probability is about 82.7%, not 90%.
Dr. Naveen Sachdev (Alzheimer's survival time)
- Population: years, years. Population is skewed (range 3–20 years, not normal).
- Sample size . Normality assumption is violated — use with caution.
- Distribution under normality: .
1. Probability
.
.
2. Probability
.
.
Exam tip: The chi‑square result for sample variance requires the population to be normally distributed. For skewed populations (like Dr. Sachdev’s), the calculated probabilities are approximate at best. A larger sample may help, but no CLT equivalent exists for variance.
Caution: Normality Assumption
All of the above inferences rely on the assumption that the population is normally distributed. If the population is not normal (e.g., skewed, like the Alzheimer's data), the distribution of is not exactly chi‑square. In practice:
- For moderate to large samples, the chi‑square approximation may still be reasonable if the population is not too non‑normal.
- Dr. Sachdev’s example illustrates a case where the assumption is clearly violated, so the computed probabilities (2.8%, 65.8%) should be interpreted as approximate, not exact.
Key takeaways
- Sample variance uses denominator to account for the loss of one degree of freedom.
- If the population is normal, .
- This relation allows probability statements about the population variance and standard deviation without having the direct distribution of .
- Use chi‑square tables or templates to find probabilities and thresholds.
- The normality assumption is critical; without it, results are unreliable, especially for small samples.
Properties of Point Estimators
A point estimator is a sample statistic (e.g., , , ) used to estimate a population parameter (e.g., , , ). Not all sample statistics make good estimators; four desirable properties distinguish excellent estimators.
Let denote the population parameter of interest and the point estimator (sample statistic) for .
Unbiasedness
An estimator is unbiased if its expected value equals the population parameter:
- For unbiased estimators, the mean of the sampling distribution is centred exactly on ; over many samples the over- and under-estimates balance out.
- Bias = . A biased estimator systematically over- or under-estimates .
| Unbiased estimator | Population parameter | Reason |
|---|---|---|
| (uses in denominator to correct bias) |
Exam tip: The sample variance uses exactly to make it an unbiased estimator of . The sample standard deviation is not unbiased, but bias is small for moderate .
Efficiency
When two unbiased estimators exist for the same , the one with smaller standard error is more efficient. A smaller standard error means values cluster more tightly around .
- Example: For a normal population, , so sample mean is more efficient than sample median as an estimator of .
Consistency
A consistent estimator improves as sample size grows: larger → estimates cluster closer to .
- Formally: for any .
- This follows from standard error shrinking with :
- → → SE
- → → SE
Sufficiency
A sufficient estimator uses all available data points in the sample.
- Both and are sufficient: they sum or count every observation.
- Sample median is not sufficient because it ignores the magnitude of values away from the centre.
Key takeaways
- Unbiasedness: ; bias can sometimes be corrected (e.g., with ).
- Efficiency: prefer unbiased estimators with smaller standard error.
- Consistency: larger samples give more precise estimates.
- Sufficiency: use every data point in the estimate.
- Sample mean and sample proportion satisfy all four properties.
Sampling Error
The inevitable deviation of a single sample from the population due to random chance. It is unavoidable when taking a probability sample; we quantify it with the standard error of the estimator. Increasing sample size reduces sampling error.
Non‑sampling Errors
Errors that arise from causes unrelated to random sampling. They introduce bias and can mislead decisions regardless of sample size.
| Type | Description | Example (Basaraju's customer survey) |
|---|---|---|
| Coverage error | The sampled population does not match the target population. | Surveying only current customers ignores competitors’ customers and non‑customers — the real growth opportunity. |
| Non‑response error | Some segments are over‑ or under‑represented due to differing response rates. | In‑store surveys miss online customers; online surveys miss in‑store customers. Response rates may not reflect actual purchasing proportions. |
| Measurement error | Responses are inaccurate due to question design, interviewer influence, or respondent dishonesty. | Non‑anonymous forms; staff watching customers fill forms; confusing online questions; rushed or dishonest answers. |
Reducing Non‑sampling Errors
- Define the target population precisely before drawing the sample.
- Design the data collection process carefully and train collectors.
- Pre‑test the data collection procedure (pilot study) to identify and fix issues.
- Choose an appropriate sampling method (stratified, cluster, systematic) to ensure key variables are reflected in the sample.
Key takeaways
- Sampling error is random and manageable via sample size; non‑sampling errors are systematic.
- Three major non‑sampling errors: coverage, non‑response, measurement.
- Non‑sampling errors cannot be fixed by increasing sample size — they require careful study design.
- Good point estimators quantify sampling error (standard error); reducing non‑sampling errors is a separate, critical step for valid inference.