Parametric and Nonparametric Methods
Statistical methods answer questions about populations using sample data: summarizing, testing relationships, making predictions. The choice between parametric and nonparametric methods depends entirely on what we assume about the underlying data distribution.
What is a parametric method?
The term comes from parameters — numbers that summarize a population (e.g., population mean , population proportion , population variance ). Parametric methods assume the data follow a specific probability distribution completely defined by these parameters.
- Example: Binary data → Bernoulli distribution, defined by .
- Example: Continuous data → Normal distribution, defined by and .
Key idea: Parametric methods require a “blueprint” — the data must come from a known distribution (e.g., normality for a -test). When the assumption holds, these methods are efficient and powerful.
What is a nonparametric method?
Nonparametric methods make no assumption about the underlying distribution. They are also called distribution-free methods. They work with ranks, relative positions, or counts rather than raw values and parameters.
- Data can be skewed, multimodal, have outliers — nonparametric methods remain valid.
- They are more flexible and robust when assumptions fail.
Core problems addressed by nonparametric methods in this module
- Independence between two variables (e.g., product type vs. satisfaction).
- Goodness of fit — does the data follow a specified distribution?
- Two-sample distribution comparison — do two samples come from the same distribution?
Worked example: Customer satisfaction scores (Product A vs. Product B)
A company collects satisfaction scores (1–10) for two products. The goal: determine whether a significant difference exists.
Parametric approach (two-sample -test)
- Null hypothesis: (population means equal).
- Alternative: .
- Assumptions: scores are normally distributed in both groups; often equal variances assumed.
- Output: estimate of mean difference, -value, confidence interval.
Nonparametric approach (Mann‑Whitney U test / Wilcoxon rank sum test)
- Null hypothesis: No difference in the distribution of scores between groups.
- Alternative: distributions differ.
- Method: Ranks all observations together; uses relative ranks to compute test statistic.
- No normality assumption; robust to skewness, outliers, bimodality.
- Output: a -value indicating whether one group tends to have higher/lower scores, but no direct estimate of effect size (e.g., mean difference).
Comparison: parametric vs. nonparametric
| Aspect | Parametric | Nonparametric |
|---|---|---|
| Assumptions | Data follows a specific distribution (e.g., normal) | None (distribution‑free) |
| Parameters estimated | , etc. | Not directly; uses ranks |
| Efficiency (given assumptions hold) | High – fewer data needed, more powerful | Lower – larger samples needed for same precision |
| Robustness to violations | Low – results can be misleading | High – valid even with skewness, outliers |
| Interpretation | Direct: mean difference, regression coefficients | Indirect: only direction or existence of difference |
| Example test | Two‑sample -test | Mann‑Whitney U test |
Exam tip: Parametric tests are more powerful only when their assumptions are met. If data is skewed or has outliers, a nonparametric test is the safer choice — it may require a larger sample but avoids invalid conclusions.
Advantages and disadvantages
Parametric
- ✅ Simpler, fewer parameters to estimate.
- ✅ Powerful and precise when assumptions hold.
- ❌ Results can be completely wrong if assumptions violated.
Nonparametric
- ✅ Flexible, works with any distribution.
- ✅ Robust to outliers and skewed data.
- ❌ Requires larger sample sizes for equivalent power.
- ❌ Provides less detailed inference (e.g., no mean difference).
Practical workflow
- Start with parametric methods if assumptions seem reasonable.
- Always check assumptions (e.g., using goodness‑of‑fit tests, Q‑Q plots).
- If assumptions are violated, switch to a nonparametric alternative.
Key takeaways
- Parametric methods assume a known distribution (e.g., normal) and estimate population parameters.
- Nonparametric methods (distribution‑free) use ranks or counts and require no distributional assumption.
- Same goal — hypothesis testing or inference — but different trade‑offs: efficiency vs. robustness.
- The Mann‑Whitney U test is the nonparametric counterpart to the two‑sample -test.
- Nonparametric tests often cannot quantify effect size beyond “different or not”.
Intuition and Definition
Independence between two events means the occurrence or non‑occurrence of one does not influence the occurrence of the other. Intuitively: knowing whether one event happened gives zero information about whether the other happened.
Example: Tossing two coins — the outcome of the first toss tells you nothing about the outcome of the second.
Formal definition – events and are independent iff
This is the multiplication rule for independent events. If the equality holds, the chance that both events occur is simply the product of their individual probabilities; no “adjustment” for influence.
Business Example: Shoe and Phone Case
Retail manager analyzing shopping behaviour:
- Event A: customer buys a particular brand of shoes
- Event B: customer buys a mobile phone case
Suppose and .
If the two purchases are independent, then
i.e., about 6% of customers would buy both.
Dependence can be positive or negative:
- Positive: customers who buy shoes are more likely to buy socks →
- Negative: customers who buy shoes are less likely to buy slippers →
For any two dependent events, the multiplication rule does not hold.
Dependent Events: Age and Gaming Console
- Event A: customer aged 18‑25
- Event B: customer purchases a gaming console
Intuition says younger people are more likely to buy a console → the events are dependent.
In this case:
A larger joint probability than the independence assumption would predict.
Independent Random Variables
Independence extends naturally to random variables and :
and are independent iff for every pair of values ,
Knowing the value of gives no information about (and vice versa).
Business context – an online platform tracks two variables per customer per month:
- = number of digital products purchased
- = number of physical goods purchased
If and are independent, marketing strategies for digital and physical products can be developed separately.
Testing Independence with Data: Toy Example
The company samples 100 customers and observes:
Marginal distribution of (digital products):
| Frequency | ||
|---|---|---|
| 0 | 10 | 0.10 |
| 1 | 50 | 0.50 |
| 2 | 30 | 0.30 |
| 3 | 10 | 0.10 |
| Sum | 100 | 1.00 |
Marginal distribution of (physical goods):
| Frequency | ||
|---|---|---|
| 0 | 5 | 0.05 |
| 1 | 40 | 0.40 |
| 2 | 40 | 0.40 |
| 3 | 15 | 0.15 |
| Sum | 100 | 1.00 |
If and are independent, the expected joint probability for any combination is the product of the marginals:
- → expected 0.5 customers
- → expected 20 customers
Similarly, the full independence‑implied joint distribution would be computed for all pairs.
Key insight: In real data, even when independence holds, the observed counts will never exactly match the products due to sampling variability. A non‑parametric test (e.g., chi‑square test of independence) determines whether the deviation is statistically significant or just random noise.
In the toy example, we expect 0.5 customers with , but we could observe 1 or 0. If observed = 1, the difference is small relative to sampling error; if observed = 10, it would strongly suggest dependence.
The goal: detect whether, for all ,
“Approximately” because sampling variation is accounted for.
Key takeaways
- Independence means ; knowing one event gives no info about the other.
- For random variables: equality of joint and product of marginals must hold for all values.
- Dependence can be positive (increases joint probability) or negative (decreases it).
- In practice, we compare observed joint frequencies to the product of marginal frequencies; sampling variability prevents exact match.
- A non‑parametric method (covered next) tests whether the gap is statistically significant.
Chi-square Test of Independence
A chi-square test of independence determines whether two categorical variables are associated. Intuitively, if the variables are independent, the observed pattern in a contingency table should match what we expect by chance; large mismatches signal dependence.
Hypotheses
- Null : The two variables are independent.
- Alternative : The two variables are not independent (dependent).
Contingency Table (Observed Frequencies)
Data are arranged in an table, where = number of rows (categories of variable 1) and = number of columns (categories of variable 2). Each cell contains the observed frequency .
Expected Frequencies
Under , the expected frequency for cell is derived from the multiplication rule of probability:
Test Statistic
The chi-square test statistic measures total squared deviation between observed and expected, standardized by expected:
Degrees of Freedom
Under , follows a chi-square distribution with:
Decision via p-value
Compute the p-value = . If p-value (commonly 0.05), reject and conclude the variables are dependent.
Worked Example: Bloomberg Businessweek
A survey asked business travelers about the type of ticket (first class, business class, economy) and trip type (domestic, international).
Observed frequencies (partial):
| Ticket type | Domestic | International | Row total |
|---|---|---|---|
| First class | 29 | 22 | 51 |
| Others (business, economy) | not specified | not specified | (remainder) |
| Column total | 642 | ? | n = 920 |
(Only the first‑class row was explicitly given; the domestic column total is 642.)
Expected frequency for first class / domestic under independence:
Contribution of this cell to :
Summing over all six cells gives:
Degrees of freedom: .
p-value: Using Excel’s CHISQ.DIST.RT(100.43, 2), the p-value is approximately (essentially 0).
Conclusion: Since p-value , reject . The type of ticket purchased is not independent of the type of trip; a significant association exists.
Exam tip: The test requires that expected frequencies are at least 5 for most cells (commonly 80% of cells). If violated, combine categories or use Fisher’s exact test. The test only detects existence of association, not strength or direction.
Key takeaways
- Chi-square test of independence assesses association between two categorical variables.
- : independence; : dependence.
- Expected frequency per cell = (row total × column total) / grand total.
- Test statistic: .
- Degrees of freedom = .
- A small p-value (e.g., < 0.05) leads to rejection of independence.
Chi-Square Test of Independence (Categorical Variables)
The chi-square test of independence determines whether two categorical variables are related or independent. Intuitively, if the variables are independent, the distribution of one variable should look roughly the same across all levels of the other variable. If they are dependent, some categories will have systematically different proportions.
Example: Department vs. Salary in an HR Dataset
- Dataset: 14,999 employees, 10 variables. Focus on two categorical variables:
- Department (10 categories): Sales, Accounting, HR, Technical, Support, Management, IT, Product Management, Marketing, R&D.
- Salary (3 categories): Low, Medium, High.
- Research question: Is salary independent of department? (Equivalently: Does salary distribution vary across departments?)
Hypotheses
- : Department and salary are independent.
- : Department and salary are not independent (i.e., there is an association).
Steps of the Test
Step 1 – Contingency Table of Observed Frequencies
Create a two-way table using COUNTIFS in Excel. Rows = departments, columns = salary levels.
| Department | Low | Medium | High | Row total |
|---|---|---|---|---|
| Sales | 2099 | ... | ... | 7316 |
| Accounting | 358 | ... | ... | ... |
| HR | ... | ... | ... | ... |
| Technical | ... | ... | ... | ... |
| Support | ... | ... | ... | ... |
| Management | ... | ... | ... | ... |
| IT | ... | ... | ... | ... |
| Product Mgmt | ... | ... | ... | ... |
| Marketing | ... | ... | ... | ... |
| R&D | ... | ... | ... | ... |
| Column total | 4140 | ... | ... | 14999 |
(Exact counts are not fully listed; the computed χ² uses the full 30-cell table.)
Step 2 – Expected Frequencies under
Under independence, the expected count for cell is:
Example for Sales × Low:
Compute all expected frequencies; verify that row and column totals match the observed marginals.
Step 3 – Chi-Square Test Statistic
For Sales × Low:
Summing over all 30 cells yields:
Step 4 – Degrees of Freedom
Step 5 – p-value
The p-value is the right-tail probability of obtaining a value at least as extreme as 700.92 under a distribution.
In Excel: =CHISQ.DIST.RT(700.92, 18) → (essentially zero).
Step 6 – Conclusion
Since the p-value is far below any common significance level (e.g., 0.05), we reject . There is strong evidence that salary and department are not independent – salary distributions differ across departments.
Understanding which categories drive the association
The largest contributions to come from the Management department:
- Deviations: (3 values) sum to ≈ 636.77 out of 700.92.
- This suggests that Management’s salary pattern is markedly different from the average.
Further analysis could test independence after removing Management (or Management, Sales, HR, Support) to see if the remaining departments show independence.
Exam tip: The chi-square test only tells you whether there is an association, not how strong it is. Large cell contributions (like Management’s) can hint at which categories differ, but follow-up tests are needed to confirm.
Key Takeaways
- The chi-square test of independence compares observed frequencies to expected frequencies under independence.
- Expected frequencies are calculated using the product of marginal totals divided by the total sample size.
- Test statistic: .
- Degrees of freedom = .
- A very small p-value leads to rejection of independence.
- The test is non-parametric – it makes no distributional assumptions about the underlying population, only that the data are counts.
Chi-Square Test of Independence for Continuous Variables
Pearson’s correlation coefficient measures linear association between two numerical variables, but it has two critical limitations:
- Normality assumption – conventional inference requires the data to be normally distributed, a condition often violated in practice.
- Linear focus only – it cannot capture non‑linear relationships. Two variables may be strongly dependent yet have near‑zero correlation.
The chi‑square test of independence overcomes both problems when applied to continuous data. It requires no distributional assumptions, detects any form of association (linear or non‑linear), and is robust to data quality issues (e.g., outliers, entry errors).
How it works: discretize then test
The core idea is to convert continuous variables into categorical bins (discretize), then apply the standard chi‑square test of independence on the resulting contingency table.
Example: Satisfaction vs. Average Monthly Hours (HR dataset)
- Satisfaction level (continuous, 0–1) → discretised into 5 bins:
[0,0.2),[0.2,0.4),[0.4,0.6),[0.6,0.8),[0.8,1]. - Average monthly hours (continuous, 96–310) → discretised into 4 bins:
≤150,(150,200],(200,250],>250.
These choices are a judgment call – guided by descriptive statistics (range, mean, median).
Rules of thumb:
- Not too few or too many bins.
- Each cell of the table should have sufficient (≥5) observations; avoid highly disproportionate counts.
The discretised data produce a joint frequency table (rows = satisfaction bins, columns = hours bins):
| Satisfaction \ Hours | ≤150 | 150–200 | 200–250 | >250 | Row total |
|---|---|---|---|---|---|
| [0.0 – 0.2) | ... | ... | ... | ... | ... |
| [0.2 – 0.4) | ... | ... | ... | ... | ... |
| [0.4 – 0.6) | ... | ... | ... | ... | ... |
| [0.6 – 0.8) | ... | ... | ... | ... | ... |
| [0.8 – 1.0] | ... | ... | ... | ... | ... |
| Column total | ... | ... | ... | ... | N |
Exam tip: The chi‑square test is performed exactly as for categorical data – compute expected frequencies under independence, calculate and compare to the critical value with degrees of freedom.
Hypotheses
- : Satisfaction level is independent of levels of average monthly hours (i.e., the row and column categories are independent).
- : Satisfaction level is not independent of levels of average monthly hours (dependence exists).
Connection to Pearson correlation
In an ideal setting, the conclusions from the chi‑square test should align with the sign and strength of the Pearson correlation coefficient. For this dataset, the (linear) correlation between satisfaction and average monthly hours is mildly negative – employees who work more tend to be less satisfied. If the chi‑square test also rejects independence, it reinforces that conclusion. A discrepancy would suggest a non‑linear relationship that correlation missed.
| Aspect | Pearson correlation | Chi‑square test (discretised) |
|---|---|---|
| Assumptions | Bivariate normality | No distributional assumptions |
| Relationship detected | Linear only | Any form (linear or non‑linear) |
| Sensitivity to outliers | High | Low (after binning) |
| Output | (‑1 to +1) | + p‑value |
Key takeaways
- Discretise continuous variables into bins (judgment call based on range and sample size) and then apply the standard chi‑square test of independence.
- The test works without normality, captures non‑linear associations, and is robust to data quality issues.
- Interpret results in the context of the bins – dependence means that the distribution of one variable differs across categories of the other.
- Compare findings with the Pearson correlation to check for consistency and to identify possible non‑linearity.
Why Goodness of Fit?
Many standard parametric methods — the t‑test, ANOVA, linear regression — assume the data follow a particular distribution (most often the normal). Other common assumptions include the exponential distribution for product lifetimes or the Poisson distribution for count data (e.g., number of defects). Because the validity of these methods hinges on the assumption being true, an objective way to check it is essential. The goodness of fit test provides that check: it determines whether a sample of observed data matches a specified theoretical distribution.
Core Idea: Observed vs. Expected Frequencies
The test compares the observed frequencies (how often each outcome actually occurred in the sample) with the expected frequencies derived from the hypothesised distribution. The key question: are the deviations between observed and expected large enough to conclude the data does not follow the assumed distribution, or can they be attributed to random chance?
This logic is the same as that of the χ² test of independence (covered earlier in the module). Both tests use a chi‑square statistic to evaluate the gap between observed and expected counts.
Assumed distribution (e.g., normal, Poisson)
↓
Compute expected frequencies for each category
↓
Collect observed frequencies from the sample
↓
Test: Are deviations (observed – expected) significantly large?
↓
If yes → reject assumption | If no → data consistent with the distribution
Exam tip: The goodness of fit test is a nonparametric method — it makes no assumption about the distribution of the test statistic itself, only about the data being tested.
Common Distributions Tested in Business Contexts
| Variable type | Example distributions | Typical application |
|---|---|---|
| Discrete | Binomial, Poisson | Number of defects in a batch (Poisson); number of successes in trials (binomial) |
| Continuous | Normal, Exponential | Customer purchase amounts (normal); product lifetime (exponential) |
Selecting the correct distribution for a given business problem is a critical first step; the goodness of fit test then validates that choice.
Business Applications
Numerous real‑world uses share the same pattern: an assumed distribution is embedded in a decision model, and the test confirms or rejects that assumption.
| Domain | Specific use | Distribution typically assumed |
|---|---|---|
| Marketing & consumer behaviour | Check if customers choose brands equally (uniform) or if preferences follow a normal distribution. | Uniform, normal |
| Retail & sales | Validate that product demand follows a seasonal or normal pattern; check if sales across store locations follow a similar distribution (identify outlier stores). | Normal, seasonal patterns |
| Product demand forecasting | Assess whether past sales conform to the distribution used for future forecasts. | Varies (e.g., normal, exponential) |
| Quality control | Test if number of defects follows a Poisson distribution (stable process); verify that product weights/dimensions match the expected distribution. | Poisson, normal |
| Risk management | Check whether actual default rates match the binomial or normal model used for credit pricing. | Binomial, normal |
| HR analytics | Evaluate if employee attrition matches historical turnover distribution to guide workforce planning. | Historical distribution (any) |
| Supply chain & inventory | Validate demand patterns for inventory policy decisions. | Various |
Why It Matters
Goodness of fit tests act as a gatekeeper for downstream analytics. If the distributional assumption is wrong, any conclusions drawn from parametric methods (confidence intervals, hypothesis tests, forecasts) may be unreliable. By confirming or rejecting the assumption, the test ensures that business decisions are based on accurate and trustworthy models.
Key takeaways
- Goodness of fit test checks whether observed data follows a specific theoretical distribution (normal, Poisson, exponential, etc.).
- It compares observed frequencies to expected frequencies from the hypothesised distribution.
- The test is nonparametric and conceptually similar to the χ² test of independence.
- Widely used across marketing, retail, quality control, risk management, HR analytics, and supply chain to validate assumptions.
- A significant result (large deviations) indicates the data does not match the assumed distribution; a non‑significant result means the data is consistent with it.
- Always first identify the appropriate distribution for the business context, then use the test to verify.
General Principles of Goodness of Fit Test
A goodness of fit test checks whether a sample of data follows a specified probability distribution. Intuitively: does the observed data "look like" it came from the assumed distribution (uniform, normal, Poisson, etc.), or is the difference too large to be chance? The test quantifies the discrepancy between what we expect under the hypothesis and what we observe in the sample, using the chi‑square distribution.
Procedure: step‑by‑step
1. Formulate hypotheses
- Null hypothesis (): The data follow a specific distribution (e.g., uniform, normal, Poisson).
- Alternative hypothesis (): The data do not follow that distribution. (No claim about what other distribution holds.)
2. Build intervals (for continuous variables)
- Divide the range into bins or categories.
- Bins need not be equal in width, but expected frequencies should be balanced (neither too large nor too small). For uniform distributions, equal‑width bins are natural.
- Example: For satisfaction level (range 0–1), use 10 bins of width 0.1.
3. Obtain observed frequencies ()
Count how many sample observations fall into each bin.
4. Compute expected frequencies () under
- Use the cumulative distribution function (CDF) of the assumed distribution.
- For any bin, , where is the total sample size.
- For a uniform distribution: .
- For other distributions (normal, binomial, Poisson), you must first estimate the parameters from the data (e.g., for normal, for Poisson). Then use the CDF (e.g.,
NORM.DISTin Excel).
5. Calculate the chi‑square test statistic
A larger value indicates greater deviation from .
6. Degrees of freedom (df)
- = number of intervals
- = number of parameters estimated from the data (e.g., 2 for normal, 0 for uniform, 1 for Poisson)
7. Decision
- Compute the p‑value using the chi‑square distribution with the appropriate df (e.g.,
CHISQ.DIST.RTin Excel). - If p‑value < significance level (typically 0.05), reject – the data do not fit the distribution.
Exam tip: The goodness of fit test only detects departure from ; it does not tell you what the true distribution is. Also, each estimated parameter reduces df by 1 – forgetting this is a common error.
Worked example: Testing uniform distribution of employee satisfaction
Data: Satisfaction level (0–1) of employees.
Hypotheses:
- : Satisfaction is uniformly distributed over [0,1].
- : Satisfaction is not uniformly distributed.
Step 2 – intervals: 10 equal‑width bins: [0,0.1], [0.1,0.2], …, [0.9,1.0].
Step 3 – observed frequencies ():
| Interval | Observed () |
|---|---|
| [0, 0.1) | 553 |
| [0.1, 0.2) | … |
| … | … |
| [0.9, 1.0] | … |
| Total | 14,999 |
(Only the first interval’s count is shown; the others sum to 14,999.)
Step 4 – expected frequencies: Under uniform distribution, each bin has probability 0.1, so
Step 5 – chi‑square statistic: The contribution for the first bin:
Summing over all 10 bins gives:
Step 6 – degrees of freedom: , (uniform requires no estimated parameters), so
Step 7 – p‑value: Using a chi‑square distribution with 9 df, the p‑value is effectively 0 (extremely small). Therefore, reject – satisfaction level is not uniformly distributed. The company should investigate factors (department, workload, etc.) driving this non‑uniform pattern.
Degrees of freedom – general rule (examples)
| Distribution | Parameters estimated () | Example with intervals |
|---|---|---|
| Uniform | 0 | |
| Normal | 2 () | |
| Poisson | 1 () | |
| Binomial | 1 () |
Exam tip: Always count how many parameters were estimated from the sample to compute the expected frequencies. This directly affects df and the critical value.
Key takeaways
- Goodness of fit tests whether sample data come from a specified distribution.
- Use the chi‑square statistic: , comparing observed vs. expected counts.
- For continuous distributions, bin the data and compute expected frequencies via the CDF.
- Degrees of freedom: (intervals minus parameters estimated).
- A very small p‑value rejects ; the data do not fit the assumed distribution.
- The test is widely applicable – uniform, normal, Poisson, binomial – but requires careful interval construction and parameter estimation.
Goodness of Fit Test for Poisson Distribution
Intuition: The Poisson distribution models the count of rare, independent events over a fixed interval (e.g., customer arrivals, equipment failures). A goodness of fit test checks whether observed frequency counts (e.g., number of projects per employee) plausibly come from a Poisson process. If the data fit, the events appear random and independent at a constant average rate. If not, systematic factors (e.g., uneven workload) may be present.
Application context: In the HR dataset (14,999 employees), the only integer-valued variable suitable for Poisson modelling is number of projects (satisfaction level is continuous; years/hours are continuous even if reported as integers). A goodness of fit test here informs whether project distribution is random — or whether some employees are overburdened.
Step-by-step procedure
1. Hypotheses
2. Observed frequencies
- Raw counts show zero employees with 0 or 1 projects, and none with >7.
- Cell merging (required to avoid zero expected frequencies): aggregate into six intervals.
- Final observed frequencies (after merging):
| Interval (projects) | Observed frequency |
|---|---|
| 2,388 | |
| (computed in Excel) | |
| (computed in Excel) | |
| (computed in Excel) | |
| (computed in Excel) | |
| (computed in Excel) | |
| Total | 14,999 |
Exam tip: Always merge cells so that no expected count is < 1 and at most 20% of cells have expected counts < 5. For Poisson, low-probability tails (0,1,7+) are typical candidates for merging.
3. Expected frequencies under
The Poisson probability mass function:
- Estimate using the sample mean: (from 14,999 observations).
- Compute probabilities for each merged interval using the Poisson distribution with .
- Multiply each probability by to get expected frequencies .
Example – first cell ( projects):
In Excel use:
POISSON.DIST(x, mean, cumulative)- For exact probability:
FALSE - For cumulative probability:
TRUE
- For exact probability:
- Last cell ():
Resulting table (partial):
| Interval | ||
|---|---|---|
| 2,388 | 4,025.79 | |
| ... | ... | |
| ... | ... | |
| ... | ... | |
| ... | ... | |
| ... | ... | |
| Total | 14,999 | 14,999.00 |
4. Chi-square test statistic
For this dataset, the sum of all six contributions yields:
5. Degrees of freedom and p-value
Using CHISQ.DIST.RT(2780.09, 4) in Excel gives a p-value ≈ 0 (extremely small).
6. Inference
Because the p-value is virtually zero, we reject . There is overwhelming evidence that the number of projects does not follow a Poisson distribution. The company should investigate why some employees are overburdened while others have too few projects.
Brief extension: Testing normality
The same logic applies to continuous variables. For a goodness of fit test for normality:
- : Data follow a normal distribution.
- Estimate both parameters: and from the sample.
- Create intervals (bins) over the continuous range.
- Compute expected frequencies using
NORM.DIST(x, mean, stdev, cumulative). - Calculate and compare to with .
Example variables from the HR dataset to test: satisfaction level, last evaluation, average monthly hours.
Key takeaways
- Goodness of fit tests whether observed frequencies match a specified theoretical distribution (Poisson, normal, etc.).
- For Poisson: estimate from the sample mean; merge low-frequency cells to satisfy expected count conditions.
- Degrees of freedom = (#cells) – 1 – (#estimated parameters).
- A very large (e.g., 2780 on 4 df) leads to rejection of the Poisson assumption.
- The same framework extends to any distribution (normal, binomial, etc.) by using the appropriate probability function.
- Real business use: validating randomness of events (arrivals, defects) or detecting systematic imbalances (workload, resource allocation).
Wilcoxon Signed Rank Test
The Wilcoxon signed rank test is a non‑parametric procedure that relies solely on the relative ordering (ranks) of observations rather than their raw values. Unlike a parametric t-test, it makes no assumption about the underlying distribution — making it ideal for small samples, skewed data, or when normality fails.
Core idea
- One‑sample: Are the data systematically higher or lower than a specified median?
- Two‑sample (paired): Is the median of the paired differences significantly different from zero?
The test ranks the absolute deviations (or differences) and then examines the signs of those deviations. If the positive deviations are consistently larger than the negative ones (or vice versa), the test infers a significant shift away from the hypothesised median.
When to use it
| Scenario | Parametric alternative | Why Wilcoxon wins |
|---|---|---|
| Small sample size, normality unclear | One‑sample t-test / paired t-test | No normality required; robust to outliers |
| Data heavily skewed (e.g., financial returns, delivery times) | Same | Ranks neutralise extreme values |
| Before‑after intervention (same subjects) | Paired t-test | Works even if effect magnitude is irregular |
| Matched pairs (treatment vs control) | Paired t-test | Handles unknown distribution of differences |
One‑sample signed rank test
Intuition: A company sets a target delivery time of 5 days. Observed times deviate above or below that target. The test checks whether the deviations are balanced around zero (i.e., typical delivery = target) or are systematically positive/negative (i.e., median delivery ≠ 5 days).
Procedure (conceptual – details follow in later lectures):
- Compute for each observation.
- Rank the absolute differences (ignore signs).
- Assign the original signs (+ or –) to the ranks.
- Sum the ranks of the positive differences () and negative differences ().
- Compare the smaller sum to a critical value.
Because ranks are used, a few extreme values do not dominate the result — the test focuses on the ordering of deviations.
Exam tip: The one‑sample signed rank test does not test the mean; it tests the median. This is a common source of confusion in exams.
Two‑sample (paired) signed rank test
When observations come in pairs — either from the same subject before/after, or from matched subjects (treatment vs control) — the test examines whether the median of the within‑pair differences is zero.
Example: A pharmaceutical company gives a drug to one group and a placebo to a matched control group. The paired difference (drug outcome – placebo outcome) is computed for each pair. If the drug has no effect, the differences should centre around zero. A signed rank test can detect if the median difference is significantly positive (drug better) or negative.
Business applications:
- Marketing: Compare sales before and after a campaign.
- Quality control: Product weight vs. a standard weight.
- Finance: Stock returns before vs. after a market event, or returns of two portfolios over the same period.
- Operations: Production efficiency before and after a process change.
Why it is attractive
- No distributional assumptions — works when a goodness‑of‑fit test rejects normality.
- Robust to outliers — ranks cap the influence of extreme values.
- Handles small samples — where parametric tests lose power or cannot be validated.
- Applicable to paired data — common in business experiments and A/B testing.
Key takeaways
- Wilcoxon signed rank test is a non‑parametric alternative to the one‑sample and paired t-tests.
- It tests the median (not mean) of differences from a hypothesised value or of paired differences.
- Ranks are based on absolute deviations; signs determine the direction of deviation.
- No normality assumption — safe for skewed data, small samples, and outlier‑prone data.
- Common in business for before/after comparisons, quality control, and financial analysis.
Wilcoxon Signed Rank Test – II
The Wilcoxon signed rank test (also called the one-sample Wilcoxon rank sum test) is a non-parametric procedure for determining whether the median of a single sample differs significantly from a hypothesized value . It is an alternative to the one-sample t-test when the normality assumption is violated or when the sample size is too small for parametric methods. Because the test works on ranks rather than raw values, it is robust to skewness and outliers.
Intuition: Instead of asking whether the mean differs from a target, we ask whether the typical observation (the median) is systematically higher or lower than a claimed value. By ranking the deviations and summing the ranks for positive deviations, we can detect a consistent directional shift.
When to use
- Data are not normally distributed (e.g., skewed, heavy tails).
- Sample size is very small (making normality check unreliable).
- Interest lies in the population median, not the mean.
- Presence of outliers would distort a t-test.
Test procedure
-
Compute differences: For each observation , calculate , where is the hypothesized median.
-
Rank absolute differences: Rank from smallest to largest (ignore sign). If ties occur, assign the average rank to tied values.
-
Assign signs: Attach the sign (+ or –) of to each rank. Ranks of zero differences () are discarded.
-
Sum positive signed ranks: Similarly, define as the sum of negative signed ranks. Always: where is the number of non‑zero differences.
-
Test statistic: For large samples (), use the z‑approximation: with Under the null hypothesis (population median ), has this known mean and standard error.
Exam tip: The approximation works because is the sum of independent (under equally likely +/–) signed ranks; by the Central Limit Theorem it becomes normal as grows.
For small samples (), use exact critical values from the Wilcoxon signed-rank distribution.
-
Decision rule: Compare to a critical value from the standard normal distribution.
- At : reject if .
- At : reject if .
Alternatively, compute the p‑value:
using Excel
NORM.DISTor similar.
Worked example: Average monthly hours
Data: HR analytics dataset with employees.
Hypothesis: hours (a believed norm) vs. .
Steps:
- Differences: For employee 1, , so .
- Rank absolute differences: . Using
RANK.AVG, ties produce average ranks (e.g., 7134.5). - Signs: Negative differences get sign –1; positive get +1; exact 200 get 0 (discarded).
- Positive signed ranks: Extract ranks for all positive signs.
- : Sum of positive signed ranks = 57,539,873.
- Z‑score:
Decision:
- At , critical value = 1.96 → , reject . The median is significantly different from 200 hours.
- At , critical value = 2.58 → , fail to reject . The evidence is not strong enough at the 1% level.
Further exploration: Testing medians 199 and 201 yields:
- 199: significant evidence of deviation.
- 201: no evidence of deviation.
Thus the organization’s median lies between 200 and 201 hours.
Comparison with one‑sample t-test
Run a one‑sample t-test on the same data (testing mean = 200) and compare the results. Because the data likely violate normality, the t-test may give a different conclusion. Always accompany such tests with exploratory plots (e.g., a histogram of monthly hours) to build a complete picture.
Exam tip: The Wilcoxon signed rank test is not a test of the mean; it is a test of the median. When the population is symmetric and normally distributed, the t-test is more powerful. When normality fails, the Wilcoxon test is more reliable.
Key takeaways
- The Wilcoxon signed rank test assesses whether the population median differs from a hypothesized value.
- It is a non‑parametric alternative to the one‑sample t-test, robust to non‑normality and outliers.
- Procedure: compute differences, rank absolute differences, assign signs, sum positive ranks ().
- For , use the z‑approximation: .
- Reject if exceeds the normal critical value (1.96 at 5%).
- Always report exploratory analysis (histograms, boxplots) alongside the test.
Paired Wilcoxon Signed‑Rank Test
The paired Wilcoxon signed‑rank test is a non‑parametric alternative to the paired t‑test. It answers: Is the median difference between two related samples significantly different from zero?
Use it when data come in pairs (same subject before/after, matched pairs) and the normality assumption fails.
Intuition
We don’t care about the absolute values of the two groups; we care only about the direction and magnitude of the difference within each pair. By ranking the absolute differences and then restoring the sign, we test whether positive and negative differences are balanced – if they are, the median difference is zero.
Hypothesis
For a two‑tailed test (is there any difference?):
- : median difference between paired observations = 0
- : median difference
For a one‑tailed test (e.g., did the campaign increase visits?), change to “median difference ” or “”.
Procedure
- Compute differences for each pair: .
- Ignore signs and rank the absolute values . Assign rank 1 to the smallest absolute difference; average ranks for ties.
- Restore signs to the ranks → signed ranks.
- Sum the positive ranks – denote this as .
- Under , should be close to half of total rank sum.
- Decision rule depends on sample size:
| Sample size | Method | Test statistic | Critical value source |
|---|---|---|---|
| Large‑sample normal approximation | Standard normal | ||
| Small‑sample exact table | Wilcoxon signed‑rank table (e.g., for , , critical = 8) |
Exam tip: The large‑sample formula uses the mean and variance . These come from the null distribution of signed ranks.
Worked Example (Marketing Campaign)
A company tracks website visits of 10 customers before and after a campaign.
Only the first four pairs are shown below (full data not given).
| Customer | Before | After | Difference | | Rank of | Signed rank | |----------|--------|-------|----------------|-------|----------------|-------------| | 1 | 50 | 55 | +5 | 5 | 4 | +4 | | 2 | 60 | 65 | +5 | 5 | 4 | +4 | | 3 | 45 | 48 | +3 | 3 | 2 | +2 | | 4 | 80 | 75 | –5 | 5 | 4 | –4 | | … | … | … | … | … | … | … |
Note: ties in absolute difference (e.g., three values of 5) receive average rank in a full ranking; this partial table uses a simplified rank for illustration.
Complete the ranking for all 10 customers, then compute = sum of positive signed ranks.
Result (per lecturer): For at (two‑tailed), the critical value from the Wilcoxon table is 8. The computed exceeds 8, so we reject – there is a statistically significant difference in website visits before vs. after the campaign.
Interpretation: The marketing campaign changed website traffic. To determine whether it increased visits, conduct a one‑tailed test (right‑tail: : median difference ).
When to Choose This Test
Key Takeaways
- Paired Wilcoxon signed‑rank test compares median difference in matched/paired data.
- Difference is computed; absolute values are ranked; signs restored; positive ranks summed → .
- For , use exact Wilcoxon table; for , use normal approximation .
- No normality assumption required – robust for small or skewed samples.
- The test tells you whether a difference exists; direction requires a one‑tailed version.
- Do not confuse with Wilcoxon rank‑sum test (for independent samples).
Summary of Module 1: Non-Parametric Methods
Non-parametric methods provide robust alternatives to parametric tests when underlying assumptions (e.g., normality, homoscedasticity) are violated. This module covered three core techniques: chi-square test of independence, goodness-of-fit test, and Wilcoxon signed-rank test. Each is widely applied in business for categorical data, distribution checks, and paired comparisons.
1. Chi-Square Test of Independence
Determines whether two categorical variables are associated. Intuition: If the variables are independent, the observed frequency of each combination should be close to what we’d expect by chance.
Test statistic:
where = observed count in cell , = expected count under independence .
Process:
- Build a contingency table of observed frequencies.
- Compute expected frequencies assuming no relationship.
- Calculate and compare to a critical value (degrees of freedom = ).
Example: A business wants to know if product type purchased is related to customer geographic region. A chi-square test on the contingency table of product × region can reveal regional preferences without any distributional assumptions.
Applications:
- Customer demographics, product categories, payment methods, satisfaction ratings.
- Can also be used as a non-parametric alternative to the correlation coefficient for continuous data (by discretizing into categories).
Key takeaways
- Only requires categorical data; no normality assumption.
- Tests independence between two variables.
- Calculated from observed vs. expected frequencies in a contingency table.
- Widely used in retail, marketing, healthcare, manufacturing.
2. Chi-Square Goodness-of-Fit Test
Assesses whether observed data follow a specific theoretical distribution (e.g., normal, Poisson). Intuition: Does the actual frequency distribution match the expected pattern?
Test statistic:
where = observed count in category , = expected count from the theoretical distribution.
Example:
- Retail sales forecasting: A company assumes customer spending follows a normal distribution. The test compares observed sales bins against normal expectations. If significant deviation is found, the forecasting model must be adjusted (better inventory, less waste/stockout).
- Quality control: Defects in manufacturing are expected to follow a Poisson distribution (constant average rate). The test verifies alignment; a significant result indicates a process shift.
Other applications:
- Finance: Check if stock returns follow a normal distribution (common in risk modeling).
- Marketing: Verify whether customer purchase patterns match expected distributions to fine-tune promotions.
Key takeaways
- Tests whether data match a specific distribution (normal, Poisson, etc.).
- Critical for validating assumptions behind forecasting, quality control, and risk models.
- Significant deviation prompts model adjustment or process investigation.
- Uses same formula as test of independence, but with one-way categories.
3. Wilcoxon Signed-Rank Test
A non-parametric alternative to the paired t-test (or one-sample t-test) when normality is not satisfied. It operates on ranks of the differences, making it robust to outliers and skewed data.
When to use:
- One sample vs. a hypothetical median.
- Two related (paired) samples: before/after, matched pairs.
Idea: Rank the absolute differences, then sum the ranks for positive and negative differences. The test statistic compares the smaller sum against a critical value.
Example: A company launches a new product feature and measures customer satisfaction for the same group before and after. Even if the satisfaction ratings are not normally distributed, the Wilcoxon test can reliably tell if the median difference is significant.
Business scenarios:
- Marketing campaign effectiveness – small sample, volatile sales figures.
- Employee productivity after a training program.
- Financial returns before/after an interest rate change – returns are often skewed with extreme outliers.
Advantages over paired t-test:
- No normality assumption required.
- Focuses on median difference rather than mean.
- Less influenced by extreme values.
Key takeaways
- Ranks differences between paired observations; does not require normality.
- Ideal for small samples, skewed data, or outliers.
- Use for one-sample or paired comparisons (before/after).
- Commonly applied in HR, marketing, and finance.
Why Non-Parametric Methods Matter
When assumptions like normality or homoscedasticity fail, these three tools offer reliable, assumption-light alternatives:
| Method | Data Type | Question Answered | Parametric Alternative |
|---|---|---|---|
| Chi-square independence | Categorical (two variables) | Are they related? | – (no direct param.) |
| Chi-square goodness-of-fit | Categorical (one variable vs. distribution) | Does the data fit a distribution? | – |
| Wilcoxon signed-rank | Paired continuous/ordinal | Is the median different? | Paired t-test |
Key takeaway for business analytics: Non-parametric methods handle real-world data (non‑normal, categorical, small samples) without sacrificing rigor. They enable confident decision‑making when parametric assumptions are untenable.