Covariance
Covariance quantifies the direction of a linear relationship between two quantitative variables. It indicates whether an increase in one variable tends to be associated with an increase (positive) or decrease (negative) in the other.
For a sample with observation pairs , the sample covariance is:
where and are sample means. The division by (rather than ) follows the same degrees‑of‑freedom reasoning used for sample variance.
Intuition via quadrants
Plot the data with vertical/horizontal lines at and :
- Quadrant I () → product
- Quadrant III () → product
- Quadrant II () → product
- Quadrant IV () → product
If most influential points lie in Quadrants I and III → (positive linear association).
If they lie in Quadrants II and IV → (negative linear association).
If points are evenly spread → (no linear association).
Relation to variance
When , covariance reduces to variance of :
Thus covariance generalises variance to pairs of variables.
Unit dependence – a major flaw
Covariance changes with the units of measurement. For example, measuring income in rupees instead of lakhs scales by , inflating even though the underlying relationship is unchanged. This makes covariance unsuitable for comparing the strength of association across different datasets.
Example: Hanumantha Pai’s credit card data
| Variable pair | Interpretation | |
|---|---|---|
| Monthly spend vs. annual income | 16,109 | Positive association |
| Monthly spend vs. household size | 1,610 | Positive association (smaller magnitude) |
The numerical values are not comparable because units differ (lakhs vs. number of people).
Exam tip: Never compare covariance values across different pairs to judge which relationship is “stronger”. Use the correlation coefficient instead.
Key takeaways – Covariance
- Measures the direction (positive/negative) but not the strength of linear association.
- Formula: .
- Positive → x and y tend to move together; negative → they move opposite.
- Unit‑dependent – not useful for comparing strength across different variable pairs.
Correlation
The Pearson product‑moment correlation coefficient (sample correlation coefficient) overcomes covariance’s unit dependence by standardising:
where and are the sample standard deviations of and .
Properties
- iff all points lie on a positively‑sloped straight line (perfect positive linear).
- iff all points lie on a negatively‑sloped straight line (perfect negative linear).
- indicates no linear relationship (non‑linear relationships may still exist).
- Closer to → stronger linear association.
Worked example – Hanumantha Pai
Income vs. Monthly spend
, lakhs, rupees.
→ Strong positive linear association.
Household size vs. Monthly spend
, , .
→ Moderate positive linear association (weaker than income).
| Correlation | Strength | Example value |
|---|---|---|
| Perfect | — | |
| Strong | 0.81 (income) | |
| – | Moderate | 0.63 (household) |
| Weak | — | |
| None (linear) | — |
Correlation is NOT causation
A high correlation does not prove that a change in one variable causes a change in the other. In the credit‑card example, higher income does not force higher spending – other factors (preferences, debt) could be at play.
Exam tip: The mantra “correlation does not imply causation” is a frequently tested concept. Always remember that a third (lurking) variable or reverse causality could explain the association.
Limitation – linear only
The Pearson correlation only measures linear relationships. A near‑zero can occur even when a strong non‑linear pattern exists (e.g., U‑shaped relationship between utility spending and outdoor temperature). Always inspect a scatter diagram alongside the correlation coefficient.
When variables are not quantitative
If one or both variables are nominal or ordinal, other measures (e.g., Spearman’s rank correlation) are used – these are discussed later.
Key takeaways – Correlation
- , dimensionless and scale‑invariant.
- Lies between (perfect negative) and (perfect positive).
- means no linear relationship – non‑linear patterns may still exist.
- Correlation is not causation.
- Always check a scatter plot to avoid being misled by a near‑zero from a non‑linear relationship.
Covariance and Correlation
Covariance measures the direction of the linear relationship between two variables: positive → both move together; negative → one goes up as the other goes down. Correlation standardises covariance onto a scale, giving a unitless measure of linear association strength.
Computing in Excel
- Sample covariance (divides by ):
=COVARIANCE.S(array1, array2) - Sample correlation (Pearson’s ):
=CORREL(array1, array2)
Applied to Hanumantha’s credit card data:
| Variables | Covariance | Correlation |
|---|---|---|
| Monthly spend & household income | 16,109 | 0.806 |
| Monthly spend & household size | 1,610 | 0.629 |
The Pearson correlation measures only linear association
Worked example (nonlinear relationship):
, .
Excel gives , . But is a perfect nonlinear relationship.
Exam tip: does not mean “no association” – it means no linear association. Always inspect a scatter plot.
Spurious vs. Explainable Correlation
- Explainable correlation: logical reason connects the variables.
- Spurious correlation: apparent relationship due to a hidden (latent) variable or pure coincidence.
| Example | Observed correlation | Hidden cause |
|---|---|---|
| Ice cream sales ↔ crime rate | Positive | Temperature (summer → more ice cream + more vacant homes) |
| Number of doctors ↔ number of deaths in villages | Positive | Poor public health → more doctors needed + more deaths |
| Divorce rate in Maine ↔ margarine consumption | (1999–2009) | No known hidden variable – purely spurious |
| Skirt length ↔ economic activity (Scotland theory) | Skirts shorter in boom, longer in bust | Proposed reason: affordability of stockings; widely considered spurious |
Key takeaways
- Covariance and Pearson correlation quantify linear association only.
- does not imply independence; check for nonlinear patterns.
- Spurious correlation can arise from hidden variables or random chance – never assume causation.
- Sample correlation is a point estimate; a hypothesis test (using a -distribution) can assess if the population correlation is nonzero.
Simple Linear Regression – I
Simple linear regression (SLR) predicts a continuous dependent (response) variable using one independent (predictor) variable . It assumes the relationship can be approximated by a straight line.
Origin and terminology
- Francis Galton (1886) studied heights of 905 children from 205 families. He multiplied female heights by 1.08 to make them male-equivalent. He observed regression toward the mean: tall parents had slightly shorter children, short parents had slightly taller children. He coined the term regression.
- = dependent / response / outcome variable
- = independent / predictor / explanatory variable / feature
The model
- (intercept), (slope) – regression coefficients (parameters)
- = predicted value (conditional expected value of given )
- = random error (residual/unexplained variation)
What “linear” really means
Linearity is defined with respect to the parameters , not the variable .
- → linear regression (parameters are linear)
- → linear regression
- → nonlinear regression (parameter appears in exponent)
Exam tip: A model is linear if it can be written as a linear combination of the parameters. Transformations of are allowed; transformations of are not (unless the model can be re‑linearised).
Examples of SLR applications
- Monthly credit card spend vs. household income or household size
- Government policy impact on wages
- Restaurant waiting time vs. popularity rating
- Hospital treatment cost vs. patient weight
- Website visits vs. e‑commerce sales revenue
- Unemployment rate vs. bank non‑performing assets
Nine steps of building a regression model
- Data preprocessing: handle outliers, missing values; derive new features (e.g., Galton’s 1.08 factor).
- Training/validation split: random sampling; typical 75% training, 25% validation.
- Ordinary Least Squares (OLS): finds that minimise . OLS gives the best linear unbiased estimates (BLUE).
- Diagnostics: check assumptions (linearity, normality of errors, homoscedasticity, independence).
- Overfitting: model performs well on training data but poorly on validation data. An ideal model has low bias and low variance.
Key takeaways
- SLR models a continuous using one ; parameters are linear in .
- “Linear” refers to coefficients, not the shape of .
- OLS minimises sum of squared residuals.
- Model building is iterative: collect → preprocess → split → describe → specify → estimate → diagnose → validate → deploy.
- Spurious correlation does not imply regression causation; always use domain logic and diagnostics.
The Regression Equation and Its Interpretation
The regression equation for a sample is written as
where
- : observed value of the dependent variable for the th observation,
- : observed value of the independent variable for the th observation,
- : random error (also called residual or unexplained variation),
- : regression parameters (coefficients).
What the equation really means.
The population can be viewed as a collection of subpopulations — one for each distinct value. For example, among Hanumanthappa’s bank customers, all customers with an annual income of ₹20 lakh form one subpopulation, those with ₹30 lakh another, and so on. Each subpopulation has its own distribution of monthly spend and its own expected value . The regression equation describes how this expected value depends on :
The error term disappears when taking the expectation because .
Estimation Using Sample Data
Population parameters and are unknown. They are estimated from a sample (the training data) by sample statistics and . Substituting these estimates gives the estimated regression equation:
where is the point estimator of:
- the mean value of for a given (e.g., mean monthly spend for all customers with ₹20 lakh income), and
- the predicted value of for an individual observation at that (e.g., predicted spend for one particular customer with ₹20 lakh income).
Exam tip: The same formula serves both for estimating a conditional mean and for predicting an individual outcome. The difference matters when constructing confidence/prediction intervals (later).
Ordinary Least Squares (OLS) Estimation
The ordinary least squares (OLS) method finds and that minimize the sum of squared deviations between observed and predicted :
Using calculus, the solution is:
- , are sample means of and .
- The slope estimates the expected change in for a one-unit increase in .
- The intercept estimates when (meaningful only if is within the data range).
Caution: Predictions using values outside the range of the training data are unreliable — the linear relationship may not hold beyond the observed data.
Worked Example: Hanumanthappa’s Bank Customers
Data: monthly spend () vs. annual income in lakhs () for 55 customers.
| Statistic | Value |
|---|---|
| 32.89 lakhs | |
| ₹3,852 | |
| 85.27 | |
| 1,047.37 |
Estimated regression equation:
Interpretation of slope: An increase of ₹1 lakh in annual income is associated with an increase of ₹85.27 in monthly spend.
Prediction: For a customer with annual income ₹35 lakhs:
Predicted monthly spend: ₹4,032.
Computing OLS in Excel
- Create a scatter plot (y vs. x).
- Right-click a data point → Add Trendline → Linear.
- Check “Display Equation on chart” and “Display R-squared value on chart”.
Excel outputs the same , , and also .
Coefficient of Determination ()
(or coefficient of determination) measures the goodness of fit of the estimated regression equation. It is the percentage of the total variation in that is explained by the linear relationship with :
- Ranges from 0 to 1 (or 0% to 100%).
- Higher indicates a better fit.
- In simple linear regression, equals the square of the Pearson correlation .
Exam tip: tells you how much of y’s variability the model explains, not whether the model is correct. Always inspect residual plots before trusting .
Key takeaways
- Regression models ; estimated by .
- OLS minimizes ; formulas: , .
- Slope interpretation: = estimated change in per one‑unit increase in .
- Do not extrapolate beyond the range of in the sample.
- = proportion of variation in explained by the model; closer to 1 → better linear fit.
Residuals and Sums of Squares
For the (i)-th observation, the residual is the difference between the observed value (y_i) and the predicted value (\hat{y}_i):
[ \text{residual}_i = y_i - \hat{y}_i ]
Residuals capture the error in using the regression line to predict the sample. The least‑squares method minimises the sum of squared residuals, called the sum of squares due to error (SSE):
[ \text{SSE} = \sum_{i=1}^{n} (y_i - \hat{y}_i)^2 ]
SSE measures the unexplained variation in (y) after fitting the model.
If no predictor (x) were available, the best guess for any (y_i) is the sample mean (\bar{y}). The total sum of squares (SST) measures the total variation in (y):
[ \text{SST} = \sum_{i=1}^{n} (y_i - \bar{y})^2 ]
The sum of squares due to regression (SSR) measures how much the regression predictions (\hat{y}_i) deviate from (\bar{y}):
[ \text{SSR} = \sum_{i=1}^{n} (\hat{y}_i - \bar{y})^2 ]
These three quantities are linked by a fundamental identity:
[ \text{SST} = \text{SSR} + \text{SSE} ]
Total variation = explained variation + unexplained variation.
Worked example (Hanumantha’s data, (n=55)): [ \text{SSE} = 40,017,408,\quad \text{SSR} = 74,180,516,\quad \text{SST} = 114,197,923 ] Check: (74,180,516 + 40,017,408 = 114,197,923).
Coefficient of Determination ( (R^2) )
The coefficient of determination (R^2) is the proportion of total variation explained by the regression:
[ R^2 = \frac{\text{SSR}}{\text{SST}} ]
- (R^2 = 1) when (\text{SSE}=0) (perfect fit).
- (R^2 = 0) when (\text{SSR}=0) (model explains nothing).
- (0 \le R^2 \le 1).
Interpreted as a percentage, (R^2) tells us the percentage of variability in (y) accounted for by the linear relationship with (x). For Hanumantha’s data:
[ R^2 = \frac{74,180,516}{114,197,923} \approx 0.650 \quad\Rightarrow\quad 65% ]
Thus, 65% of the variation in monthly spend is explained by annual income.
Exam tip: A high (R^2) indicates good fit in the sample, but does not guarantee the model is appropriate – assumptions must still be checked.
Key takeaways (decomposition & (R^2))
- (\text{SST} = \text{SSR} + \text{SSE}) – total variation splits into explained and unexplained parts.
- (R^2 = \text{SSR}/\text{SST}) – fraction of variation explained by the regression.
- Higher (R^2) implies a better linear fit, but (R^2) alone does not validate the model.
Correlation Coefficient vs. Coefficient of Determination
The sample correlation coefficient (r) measures the strength and direction of a linear relationship between (x) and (y) ((-1 \le r \le 1)). It relates to (R^2) in simple linear regression:
[ r = \sqrt{R^2} \times \text{sign}(b_1) ]
where (b_1) is the estimated slope from the regression.
- If (b_1 > 0), (r = +\sqrt{R^2}).
- If (b_1 < 0), (r = -\sqrt{R^2}).
In Hanumantha’s example, (b_1 = 85.27 > 0) and (R^2 = 0.65), so:
[ r = \sqrt{0.65} \approx 0.806 ]
This indicates a strong positive linear association.
Exam tip: Unlike (r), (R^2) can be used for nonlinear models and multiple regression – it is a more general measure of fit.
Assumptions of the Error Term
For valid inference in simple linear regression, four key assumptions are made about the error term (\varepsilon):
- Zero mean: (E(\varepsilon) = 0) for all (x). Hence (E(y) = \beta_0 + \beta_1 x).
- Constant variance: (\text{Var}(\varepsilon) = \sigma^2) for all (x) (homoscedasticity).
- Independence: The errors are independent; no autocorrelation.
- Normality: (\varepsilon) is normally distributed for every (x).
Together, these imply that for any given (x), (y) is normally distributed with mean (\beta_0 + \beta_1 x) and constant variance (\sigma^2).
Test for Significance of the Linear Relationship
We test whether the slope (\beta_1) is different from zero (i.e., whether a linear relationship exists).
[ H_0: \beta_1 = 0,\quad H_a: \beta_1 \neq 0 ]
Estimating (\sigma^2)
The mean square error (MSE) provides an unbiased estimate of (\sigma^2):
[ \text{MSE} = s^2 = \frac{\text{SSE}}{n-2} ]
With (n-2) degrees of freedom (two parameters estimated: (\beta_0, \beta_1)). The standard error of the estimate is (s = \sqrt{\text{MSE}}).
Hanumantha’s data: (n=55), (\text{SSE} = 40,017,408), so
[ \text{MSE} = \frac{40,017,408}{53} \approx 755,045,\quad s = \sqrt{755,045} \approx 868.9 ]
t‑Test
The standard error of the slope (b_1) is:
[ \text{se}(b_1) = \frac{s}{\sqrt{\sum (x_i - \bar{x})^2}} ]
For Hanumantha’s data, (\sum (x_i - \bar{x})^2 = 10,201), so
[ \text{se}(b_1) = \frac{868.9}{\sqrt{10,201}} \approx 8.603 ]
The test statistic follows a t distribution with (n-2) degrees of freedom:
[ t = \frac{b_1 - \beta_1}{\text{se}(b_1)} = \frac{85.27 - 0}{8.603} \approx 9.91 ]
The corresponding p‑value (two‑tailed) is extremely small ((\approx 1.1 \times 10^{-13})), providing strong evidence to reject (H_0). We conclude that a significant linear relationship exists.
Exam tip: For simple linear regression, the t‑test and an F‑test (using SSR/MSR vs MSE) are equivalent. In multiple regression they diverge.
Key takeaways (assumptions & significance test)
- Error term assumptions: zero mean, constant variance, independence, normality.
- (\text{MSE} = \text{SSE}/(n-2)) estimates (\sigma^2); (s = \sqrt{\text{MSE}}) is the standard error of the estimate.
- Test (H_0: \beta_1=0) using (t = b_1/\text{se}(b_1)) with (n-2) df.
- A very small p‑value (e.g., (<0.05)) indicates a statistically significant linear relationship.
SLR - Summary Statistics
The Excel regression output for Hanumantha's data provides a compact summary of all key statistics. The R square value is 0.65, meaning 65% of the variation in monthly spend is explained by annual income. The correlation between the two variables is the square root of R square: , indicating a strong positive linear relationship.
The standard error is the estimate of (the standard deviation of the error term ). Its square estimates , the variance of .
ANOVA Table
| Source | SS | df | MS | F | p-value |
|---|---|---|---|---|---|
| Regression (SSR) | 74,180,516 | 1 | 74,180,516 | 98.25 | |
| Residual (SSE) | 40,017,408 | 53 | 755,045 | ||
| Total | 114,197,924 | 54 |
- Sum of squares regression SSR = 74,180,516.
- Sum of squares error SSE = 40,017,408.
- Mean square regression MSR = SSR / 1 = 74,180,516.
- Mean square error MSE = SSE / 53 = 755,045.
- F statistic = MSR / MSE = 98.25.
- The p-value (Significance F) is near zero, so we reject and conclude the linear relationship is statistically significant.
Coefficients and Estimated Regression Equation
| Coefficient | Estimate | Standard Error | t Statistic | p-value |
|---|---|---|---|---|
| Intercept () | 10,470.37 | — | — | — |
| Annual Income () | 85.27 | 8.6 | 9.91 |
Estimated regression equation:
- standard error: .
- t test for : , with p-value → reject .
Exam tip: Excel’s regression tool outputs all these values at once. The key interpretation steps: check R square, then overall F test, then individual coefficient t tests. The F and t tests yield the same conclusion for simple linear regression (since ).
Key takeaways
- R square = 0.65, correlation = 0.806.
- Standard error estimates .
- F test (, p ≈ 0) confirms a significant linear relationship.
- Estimated equation: .
- Slope’s t test (, p ≈ 0) matches F test result.
SLR - Residual Analysis
Residual for observation : (observed minus predicted). Residuals are the best empirical proxy for the error term . Three model assumptions about must be checked:
- Expected value (automatically satisfied by least squares).
- Constant variance — homoscedasticity — for all .
- Independence of errors.
- Normality — .
If assumptions fail, hypothesis tests (t and F) may be invalid.
Residual Plots
Three diagnostic plots are commonly examined:
- Residual vs. independent variable (provided by Excel).
- Residual vs. predicted values (can be created from output).
- Standardized residual plot (also provided by Excel).
Standardized residual: (mean = 0, standard deviation ≈ 1). For normally distributed errors, about 95% of standardized residuals lie between −2 and +2.
Interpreting Residual Plot Patterns
- Panel A (horizontal band): no obvious violation → model assumptions appear satisfied.
- Panel B (funnel): variance increases with → homoscedasticity assumption violated.
- Panel C (U-shape): systematic curvature → linear model not appropriate.
For Hanumantha’s data, the residual plot shows a horizontal band pattern (like Panel A), supporting the assumptions. The same pattern appears in the residual vs. plot (for simple linear regression, both plots give identical information; for multiple regression, the plot is preferred).
Checking Normality with Standardized Residuals
Excel provides standardized residuals. In Hanumantha’s dataset, nearly all standardized residuals lie within ±2, consistent with a standard normal distribution. Thus the normality assumption is not violated.
Exam tip: Residual analysis is qualitative. Look for clear deviations (funnel, U-shape, many points outside ±2). A few points near the boundaries do not automatically invalidate the model.
Key takeaways
- Residuals estimate and are used to validate model assumptions.
- Three key plots: residuals vs. , residuals vs. , standardized residuals.
- Desired pattern: horizontal band (constant variance, linearity).
- About 95% of standardized residuals should fall in [−2, +2] for normality.
- Hanumantha’s data satisfies all assumptions (band pattern, no extreme standardized residuals).
Outliers
An outlier is a data point that does not follow the trend of the rest of the data. In simple linear regression, outliers can be detected via:
- Scatter plot of vs. .
- Standardized residuals — any observation with (roughly) is a potential outlier.
Handling Outliers
| Situation | Action |
|---|---|
| Erroneous data (entry/recording error) | Correct the value and re-run regression; R square usually improves. |
| Violation of model assumptions | Consider a different model (e.g., curvilinear, transformation). |
| Genuine unusual value | Retain the observation — it may be a chance occurrence. |
Influential Observations
An influential observation dramatically changes the estimated regression equation (slope, intercept, or both) when removed. Detection methods:
- Scatter diagram — look for points that are:
- An outlier in (far from the trend line),
- An extreme value (far from ),
- A combination of both.
Handling Influential Observations
- Check for errors in data collection/recording — correct if possible.
- If valid, treat the observation as providing additional insight. Consider collecting more data at intermediate values to better understand the relationship.
Exam tip: Outliers and influential points are not automatically bad. Always investigate the reason before discarding. In small datasets, a single influential point can dominate the regression — report it transparently.
Key takeaways
- Outliers: standardized residual; may be erroneous, model violation, or genuine.
- Influential observation: extreme , outlying , or both; can shift the regression line substantially.
- First step for both: check data accuracy.
- Valid influential observations may motivate collecting additional data at intermediate values.
Test of Significance and Residual Analysis
A fitted regression line is not enough — we must ask whether the observed linear relationship is statistically significant or merely a fluke of sampling. Two hypothesis tests (t-test and F-test) answer this question. Both rely on assumptions about the error term , which are checked through residual analysis.
The Regression Model and Estimation (Review)
The simple linear regression model is
Using ordinary least squares (OLS), we obtain estimates and , giving the estimated regression equation:
OLS minimises the sum of squares error (SSE):
Variation in is decomposed as
where
- (total sum of squares)
- (regression sum of squares)
The coefficient of determination measures how well the model explains the data:
The correlation between and is with the sign of .
Assumptions for Hypothesis Testing
The error term is assumed to be:
- Normally distributed (for exact p‑values)
- Mean zero,
- Constant variance (homoscedasticity)
- Independent across observations
These assumptions are critical for the validity of the tests. The variance is estimated by the mean square error (MSE):
The standard error of is .
t‑Test for
- Null hypothesis : (no linear relationship)
- Alternative :
Test statistic:
Under , follows a -distribution with degrees of freedom. The p-value is computed; if small (e.g., ), reject and conclude the linear relationship is statistically significant.
F‑Test for Overall Significance
For simple linear regression, the F‑test tests the same null hypothesis ().
Test statistic:
If is true, is approximately 1. If , . The -statistic has 1 and degrees of freedom; a small p‑value leads to rejecting .
Exam tip: In simple linear regression, the t‑test and F‑test are equivalent — they give the same p‑value. In multiple linear regression, the F‑test tests the overall model while t‑tests test individual coefficients.
Residual Analysis
Residuals are used to check the assumptions on .
- Standardised residuals (or simpler ). Values outside suggest possible outliers.
- Plot residuals vs. — should be randomly scattered around zero with constant spread (no funnel shape, no curvilinear pattern).
- Plot residuals vs. (or vs. ) — same expectation. A clear trend indicates that the model is missing relevant variables or that a linear fit is inadequate.
Key takeaways
- Both t‑test and F‑test test ; reject → linear relationship is significant.
- estimates error variance ; is the standard error of the estimate.
- Residual plots must be examined for violations of constant variance, independence, and linearity.
- Standardised residuals beyond ±2 may indicate outliers.
Application: Monthly Spend vs. Household Size
Using the same dataset (Hanumantha’s), we now model monthly spend () as a function of household size ().
Estimated Regression Equation
From Excel’s regression output:
Model Fit
| Measure | Value |
|---|---|
| 0.396 | |
| Correlation | |
| Standard error | 1140.71 |
Only 39.6% of the variation in monthly spend is explained by household size — much weaker than the model using annual income ().
ANOVA and F‑Test
| Source | SS | df | MS | p‑value | |
|---|---|---|---|---|---|
| Regression (SSR) | 45,233,346 | 1 | 45,233,346 | 34.7 | ≈ 0 |
| Error (SSE) | 68,964,576 | 51 | 1,301,218 | ||
| Total (SST) | 114,197,922 | 52 |
with a very small p‑value → reject → linear relationship between monthly spend and household size is statistically significant.
Coefficients and t‑Test
| Coefficient | Value | Std. Error | p‑value | |
|---|---|---|---|---|
| Intercept | 2083.7 | |||
| Household size | 520.1 | 88.2 | 5.89 | ≈ 0 |
The t‑test confirms significance (same conclusion as F‑test).
Residual Diagnostics
- Standardised residuals: three values exceed +2 in absolute value → potential outliers.
- Residuals vs. : no funnel or curvilinear shape; band is acceptable.
- Residuals vs. : shows a slight trend (not a perfect random band). This suggests that other factors (e.g., income) also influence monthly spend — a multiple regression model may improve fit.
Exam tip: Always check residual plots, not just and p‑values. A trend in residuals vs. fitted values is a red flag that important predictors are missing.
Key takeaways
- means household size alone explains only 39.6% of spend variation.
- Despite low , the linear relationship is statistically significant (p ≈ 0 for both t and F).
- Three observations are flagged as possible outliers (standardised residual > 2).
- Residual pattern suggests the model is incomplete; include other variables like income in multiple regression.
Simple Linear Regression — Applications: Examples II
Two worked examples demonstrate the full SLR workflow: check scatter plot, fit model, interpret R², evaluate significance via F-test and t-test, examine residuals for violations and missing predictors.
Example 1: Buzzer Rogers Customer Satisfaction
Context — 36 survey responses. Question: Is price a significant determinant of satisfaction score?
- Dependent variable = satisfaction score (0–100)
- Independent variable = price paid (₹ 1000s)
Step 1 — Scatter plot
Price on x‑axis, score on y‑axis. A linear trend is visible, though with spread.
Step 2 — Regression equation & R²
From the trend line (and later from the tool pack):
- R² = 0.36 — only 36 % of score variation is explained by price. The plot showed a fair trend, so the low R² hints that other factors (brand, RAM, etc.) matter.
- Correlation (positive, moderate).
Step 3 — ANOVA and overall significance
| Source | SS | df | MS | ||
|---|---|---|---|---|---|
| Regression (SSR) | 510.8 | 1 | 510.8 | = MSR/MSE = 18.9 | ≈ 0 |
| Error (SSE) | 914.8 | 34 | 26.9 | ||
| Total | 1425.6 | 35 |
- = 18.9, reject . The linear relationship is statistically significant.
Step 4 — Coefficient table
| Predictor | Coefficient | Std. Error | ||
|---|---|---|---|---|
| Intercept () | 48.44 | — | — | — |
| Price () | 0.132 | 0.03 | 4.36 | ≈ 0 |
- matches (; rounding). small → reject .
- Interpretation: For every ₹ 1000 increase in price, the predicted satisfaction score rises by 0.132 points.
⚠️ Exam tip: This slope is counter‑intuitive (higher price → higher satisfaction). Regression describes the sample; the data may reflect that higher‑priced computers satisfy certain customers. Always check real‑world plausibility separately.
Step 5 — Residual diagnostics
- Standardized residuals: Two values below –2 (–2.48, –2.12) — possible outliers (points far below the regression line).
- Residuals vs. x: No funnel or curve, but not a perfect band. Some pattern suggests omitted variables (e.g., brand, features).
- Residuals vs. y: Shows a trend — again hints that more predictors are needed.
Key takeaways — Example 1
- R² = 0.36: price alone explains little of satisfaction.
- Relationship is significant despite low R² ( test rejects ).
- Slope is positive — data says higher price was associated with higher satisfaction in this sample.
- Residual analysis flags outliers and missing predictors → multiple regression needed.
Example 2: Manjula Nayak — Household Debt (99 Acres)
Context — ~500 Bengaluru households. Question: Is first income a significant determinant of household debt?
- Dependent variable = household debt (₹)
- Independent variable = first income (₹)
Step 1 — Scatter plot
First income vs. debt. A linear trend is barely visible; large variation in debt for a given income.
Step 2 — Regression equation & R²
- R² = 0.31 — income only explains 31 % of debt variation.
- Correlation (positive, moderate).
Step 3 — ANOVA and overall significance
| Source | SS | df | MS | ||
|---|---|---|---|---|---|
| Regression (SSR) | 1 | = 228 | ≈ 0 | ||
| Error (SSE) | 498 | ||||
| Total | 499 |
- = 228, → reject . Linear relationship is significant.
Step 4 — Coefficient table
| Predictor | Coefficient | Std. Error | ||
|---|---|---|---|---|
| Intercept () | 32 976 | — | — | — |
| First income () | 0.04 | 0.002 | 15.1 | ≈ 0 |
- (very large). tiny → strong evidence against .
- Interpretation: For every ₹ 1 increase in first income, debt rises by ₹ 0.04.
Step 5 — Residual diagnostics
- Standardized residuals: Few values outside ±2, but mostly within range.
- Residuals vs. x: Oval shape — not a perfect band, no funnel or curve.
- Residuals vs. y: Shows a trend, again hinting that other variables (ownership, monthly payments, etc.) affect debt.
Key takeaways — Example 2
- R² = 0.31: income alone explains little of debt.
- Relationship is highly significant (, ).
- Slope 0.04 means a 1‑rupee increase in income is associated with 0.04 rupee more debt.
- Residual patterns signal that a multiple regression (next module) will better capture debt drivers.
Common workflow across both examples
Exam tip: When R² is low but the overall ‑test is significant, the model is still statistically useful but practically weak — other variables are needed. Residual plots that show trends (not random scatter) are the main clue.
SLR Applications: Monthly Payment vs. Debt
Intuition: Households with higher monthly payments (loans, credit cards, utilities) are expected to carry more debt. This second example uses monthly payment as the independent variable to predict debt , comparing results with the earlier income-based model.
Estimated Regression Equation
Using the same 500-household dataset, the least‑squares regression yields:
- : for every one‑rupee increase in monthly payment, predicted debt increases by 2.94 rupees.
- Caution: With only one predictor, this coefficient captures both the direct effect and correlations with omitted variables; it is not a pure marginal effect.
Goodness‑of‑Fit: and Correlation
| Measure | Value | Interpretation |
|---|---|---|
| 0.37 | Monthly payment explains 37% of the variability in debt. | |
| 0.60 | Positive correlation between monthly payment and debt. |
Comparison with earlier model (Debt vs. First Income):
| Model | ||
|---|---|---|
| Debt vs. Income | 0.31 | 0.56 |
| Debt vs. Monthly Payment | 0.37 | 0.60 |
The monthly‑payment model is marginally better, but both leave >60% of variance unexplained – strong evidence that debt depends on multiple factors.
ANOVA and Hypothesis Tests
Null hypothesis: (no linear relationship).
Alternative: .
| Source | Sum of Squares | df | Mean Square | -value | |
|---|---|---|---|---|---|
| Regression (SSR) | 1 | 288 | |||
| Error (SSE) | 498 | ||||
| Total (SST) | 499 |
- Standard error of estimate: rupees – typical prediction error.
- , → reject ; the linear relationship is statistically significant.
Coefficient details for :
| Coefficient | Std. Error | statistic | -value | |
|---|---|---|---|---|
| Monthly payment () | 2.94 | 0.174 |
- The ‑test confirms , identical conclusion to the ‑test.
Residual Analysis
Standardized residuals: Several observations lie below −2 or above +2, suggesting potential outliers.
Residual vs. (monthly payment): Forms an oval shape – not a perfect random band, but no obvious funnel or curvature. This points to heteroscedasticity or non‑linearity that may be addressed by including additional predictors.
Residual vs. : Similar pattern reinforces that the model fails to capture all systematic variation.
Exam tip: An oval‑shaped residual plot (instead of a horizontal band) is a classic sign that a simple linear regression is misspecified – often resolved by moving to multiple regression.
Key takeaways
- Monthly payment explains 37% of debt variability (, ), slightly better than income (31%).
- Both slope and overall model are highly significant (, , ).
- Residual analysis reveals several outliers and a non‑random pattern, indicating omitted variables.
- Coefficient cannot be interpreted as a pure marginal effect in simple regression.
Module Summary
This module introduced the fundamentals of measuring and modelling linear relationships between two variables.
Key concepts covered:
-
Covariance and correlation – quantify the direction and strength of linear association.
- Distinguish explainable vs. spurious correlation.
-
Simple linear regression model with .
-
Least‑squares estimation – sample statistics and estimate population parameters.
-
Coefficient of determination – proportion of ’s variability explained by .
-
Assumptions about errors – independence, constant variance, normality – required for valid inference.
-
Hypothesis tests – ‑test for and ‑test for overall significance (equivalent in SLR).
-
Residual analysis – plots and standardized residuals used to check assumptions and detect outliers/influential points.
The next module extends this framework to multiple linear regression, where depends on several variables.
Key takeaways
- Covariance and correlation measure association; regression models the relationship.
- is the key measure of fit; low indicates omitted variables.
- Hypothesis tests (, ) assess statistical significance of the linear relationship.
- Residual plots are essential for model validation.