Term 1 · Module 7 of 8

Multiple Linear Regression - I

Business Statistics for Entrepreneurs

Multiple Linear Regression - I

Multiple linear regression (MLR) extends simple linear regression to model a single dependent variable yy using two or more independent variables x1,x2,…,xpx_1, x_2, \dots, x_p. The intuition: real-world outcomes rarely depend on just one factor. By including more relevant predictors, MLR typically yields better predictions than simple regression.

The MLR Model

The multiple regression model describes how yy relates to the independent variables plus a random error term:

y=β0+β1x1+β2x2+⋯+βpxp+ϵy = \beta_0 + \beta_1 x_1 + \beta_2 x_2 + \dots + \beta_p x_p + \epsilon

where:

  • β0,β1,…,βp\beta_0, \beta_1, \dots, \beta_p are parameters (unknown population coefficients).
  • ϵ\epsilon is the error term – a random variable that captures variability in yy not explained by the linear combination of the xx's.
  • pp denotes the number of independent variables.

Key assumption

The expected value of the error term is zero: E(ϵ)=0E(\epsilon) = 0. Consequently, the mean or expected value of yy is:

E(y)=β0+β1x1+β2x2+⋯+βpxpE(y) = \beta_0 + \beta_1 x_1 + \beta_2 x_2 + \dots + \beta_p x_p

This equation is called the multiple regression equation. If the parameters were known, we could compute E(y)E(y) for any given values of x1,…,xpx_1, \dots, x_p.

Estimated regression equation

In practice, parameters are unknown. Using a sample (training data), we obtain point estimators b0,b1,…,bpb_0, b_1, \dots, b_p for β0,β1,…,βp\beta_0, \beta_1, \dots, \beta_p. The estimated multiple regression equation is:

y^=b0+b1x1+b2x2+⋯+bpxp\hat{y} = b_0 + b_1 x_1 + b_2 x_2 + \dots + b_p x_p

y^\hat{y} is the predicted value of the dependent variable.

Exam tip: The structure of MLR is a direct generalization of simple linear regression. Expect to see the same core concepts (estimators, residuals, hypothesis tests) extended to multiple predictors.

Key takeaways – MLR model

  • pp = number of independent variables.
  • Model: y=β0+β1x1+⋯+βpxp+ϵy = \beta_0 + \beta_1 x_1 + \dots + \beta_p x_p + \epsilon with E(ϵ)=0E(\epsilon)=0.
  • Regression equation: E(y)=β0+β1x1+⋯+βpxpE(y) = \beta_0 + \beta_1 x_1 + \dots + \beta_p x_p.
  • Estimated equation: y^=b0+b1x1+⋯+bpxp\hat{y} = b_0 + b_1 x_1 + \dots + b_p x_p, where bb's are sample estimates.

Steps for Developing an MLR Model

Model building is an iterative process; many steps may need revisiting. The following ten steps are typical:

  1. Collect data – Identify relevant independent variables and gather data from secondary sources (e.g., ERP systems, databases, government census) or primary sources (surveys, interviews).

  2. Preprocess data – Assess completeness and correctness. Handle missing observations and transform qualitative/categorical variables (e.g., through dummy coding).

  3. Descriptive analytics – Compute basic statistics and use visualizations (scatter plots, box plots) to examine correlations, flag potential multicollinearity and overfitting. Identify proxy variables if original variables cannot be used (ethical or regulatory reasons).

  4. Modeling strategy – Select the set of independent variables. Desirable properties:

    • Variables should be independent of each other.
    • Model performance should be consistent across training and test data.
    • Variables should be controllable by the decision maker (e.g., price is controllable; weather is not). A model with uncontrollable variables may be less actionable.
  5. Split data – Divide into training and validation sets. Multiple subsets can be used for cross-validation.

  6. Define functional form – Start with a linear relationship between yy and the xx's. Nonlinear transformations (e.g., x2x^2) are allowed and do not make the regression non‑linear in parameters – they are still estimated with OLS.

  7. Estimate parameters – Use ordinary least squares (OLS) to find b0,b1,…,bpb_0, b_1, \dots, b_p that minimize the sum of squared residuals. OLS provides the Best Linear Unbiased Estimator (BLUE) under standard assumptions.

  8. Model diagnostics – Check statistical significance:

    • F‑test – tests overall model significance (whether all coefficients except β0\beta_0 are jointly zero).
    • t‑test – tests significance of individual coefficients. Additional diagnostics for MLR include checking normality of residuals, multicollinearity, and heteroscedasticity. If assumptions are violated, take remedial measures.
  9. Validate model – Use the validation data to assess performance. Common metrics:

    • R2R^2 (coefficient of determination)
    • Adjusted R2R^2 (penalizes for number of predictors)
    • Mean absolute percentage error (MAPE)
    • Root mean square error (RMSE) Models with low bias and low variance on both training and validation sets are preferred.
  10. Deploy model – Generate actionable insights and create an implementation plan. Continuously monitor performance; models may need to be updated over time.

Exam tip: The F‑test and t‑test in MLR serve the same roles as in simple regression, but now the F‑test checks the entire set of predictors simultaneously. Multicollinearity is a new issue unique to MLR – watch for high correlation among independent variables.

Key takeaways – model development

  • MLR development is iterative; steps are often repeated.
  • Data quality, preprocessing, and descriptive analytics are critical.
  • OLS yields BLUE estimates.
  • Diagnostics include F‑test (overall), t‑tests (individual), normality, multicollinearity, heteroscedasticity.
  • Validation uses R2R^2, adjusted R2R^2, MAPE, RMSE; aim for consistent performance.
  • Deployment requires ongoing monitoring.

OLS in Multiple Linear Regression

Ordinary Least Squares (OLS) extends from simple regression to the multiple regression setting. The principle is identical: minimize the sum of squared deviations between observed and predicted values of the dependent variable.

Minimization criterion

Minimize ∑i=1n(yi−y^i)2\text{Minimize } \sum_{i=1}^{n} (y_i - \hat{y}_i)^2

  • yiy_i = observed value of the dependent variable for observation ii
  • y^i\hat{y}_i = predicted value from the estimated multiple regression equation

Estimated regression equation

y^=b0+b1x1+b2x2+⋯+bpxp\hat{y} = b_0 + b_1 x_1 + b_2 x_2 + \dots + b_p x_p

The coefficients b0,b1,…,bpb_0, b_1, \dots, b_p are obtained by applying OLS to sample data. In simple regression closed-form formulas exist; in multiple linear regression (MLR) the solution uses matrix algebra (covered via software). Interpretation of output is the primary focus.

Interpreting coefficients in MLR vs. simple regression

In simple regression, b1b_1 is the estimated change in yy for a one‑unit change in xx. In MLR, the interpretation becomes conditional:

bib_i represents the estimated change in yy corresponding to a one‑unit change in xix_i when all other independent variables are held constant.

Worked example: Hanumantha’s credit card data

  • n=55n = 55 observations
  • yy = monthly credit card spend (Rs. \text{Rs. })
  • x1x_1 = annual household income (₹ lakh)
  • x2x_2 = household size (persons)

Simple regression with only income: y^=1047.37+85.27x1\hat{y} = 1047.37 + 85.27 x_1

  • F=98.24F = 98.24, p≈0p \approx 0 → significant
  • R2=0.65R^2 = 0.65

Simple regression with only household size: y^=2083.67+520.13x2\hat{y} = 2083.67 + 520.13 x_2

  • F=34.76F = 34.76, p≈0p \approx 0 → significant
  • R2=0.39R^2 = 0.39

Multiple regression with both predictors: y^=355.68+71.16x1+339.9x2\hat{y} = 355.68 + 71.16 x_1 + 339.9 x_2

  • R2=0.80R^2 = 0.80

Notice that b1b_1 dropped from 85.2785.27 (simple) to 71.1671.16 (MLR) because the effect of income is now estimated holding household size fixed.

CoefficientSimple (only income)MLR (income + size)Interpretation in MLR
b1b_1 (income)85.2785.2771.1671.16For a ₹1 lakh increase in income, monthly spend increases by ₹71.16, assuming household size remains constant.
b2b_2 (size)520.13520.13339.90339.90For an increase of one person in household size, monthly spend increases by ₹339.90, assuming income remains constant.

Exam tip: The conditional interpretation (“holding other variables constant”) is the key distinction between simple and multiple regression coefficients. Always phrase it explicitly.

Key takeaways – OLS in MLR

  • OLS minimises ∑(yi−y^i)2\sum (y_i - \hat{y}_i)^2, whether for one or many predictors.
  • Coefficients are computed via matrix algebra; focus on interpretation of software output.
  • In MLR, bib_i measures the change in yy per unit change in xix_i with all other xx’s held constant.
  • Adding a relevant predictor can change the magnitude (and sometimes sign) of existing coefficients – do not expect them to stay the same.

Multiple Coefficient of Determination

The multiple coefficient of determination R2R^2 measures the overall fit of the estimated regression equation.

Sum of squares decomposition

SST=SSR+SSESST = SSR + SSE

TermDefinitionFormula
SSTSSTTotal sum of squares (total variability in yy)∑(yi−yˉ)2\sum (y_i - \bar{y})^2
SSRSSRSum of squares due to regression (explained variability)∑(y^i−yˉ)2\sum (\hat{y}_i - \bar{y})^2
SSESSESum of squares due to error (unexplained variability)∑(yi−y^i)2\sum (y_i - \hat{y}_i)^2

SSTSST depends only on yy, not on the model. Adding predictors typically increases SSRSSR and decreases SSESSE, improving fit.

R2R^2 in simple vs. multiple regression

R2=SSRSSTR^2 = \frac{SSR}{SST}

Interpretation: proportion of the variability in yy explained by the estimated regression equation (multiply by 100 for percentage).

Hanumantha’s example – R2R^2 comparison:

ModelSSRSSRSSESSESSTSSTR2R^2
Simple (income only)74,180,51640,017,408114,197,9230.65 (65%)
MLR (income + size)91,468,85222,729,070114,197,9230.80 (80%)

R2R^2 increased from 0.65 to 0.80 – the two predictors together explain 80% of the variation in monthly spend.

Why R2R^2 always increases (or stays the same) with more predictors

Adding an independent variable cannot worsen the fit because OLS can always set its coefficient to zero, leaving R2R^2 unchanged. In practice, R2R^2 almost always increases.

Adjusted R2R^2 – penalizing for extra variables

The adjusted R2R^2 modifies R2R^2 to account for the number of predictors, preventing overestimation of explanatory power.

Adjusted R2=1−(1−R2)n−1n−p−1\text{Adjusted } R^2 = 1 - (1 - R^2)\frac{n-1}{n-p-1}

  • nn = number of observations
  • pp = number of independent variables

If pp is large relative to nn, the adjusted R2R^2 can become negative – in that case set it to 0.

Hanumantha’s example – adjusted R2R^2:

  • Simple regression (p=1): 1−(1−0.65)55−155−2=1−0.35×5453≈0.641 - (1 - 0.65)\frac{55-1}{55-2} = 1 - 0.35 \times \frac{54}{53} \approx 0.64
  • MLR (p=2): 1−(1−0.80)55−155−3=1−0.20×5452≈0.7931 - (1 - 0.80)\frac{55-1}{55-3} = 1 - 0.20 \times \frac{54}{52} \approx 0.793

The adjusted R2R^2 (0.793) is only slightly lower than R2R^2 (0.80), suggesting the extra variable is genuinely useful.

Exam tip: R2R^2 alone can mislead – always report adjusted R2R^2 when comparing models with different numbers of predictors.

Key takeaways – Multiple coefficient of determination

  • R2=SSR/SSTR^2 = SSR/SST; it is the proportion of yy‑variation explained by the model.
  • SSTSST is unaffected by the model; adding predictors increases SSRSSR and decreases SSESSE, pushing R2R^2 upward.
  • Adjusted R2R^2 penalises for adding useless predictors; use it to avoid overfitting.
  • Formula: Adj R2=1−(1−R2)n−1n−p−1\text{Adj }R^2 = 1 - (1-R^2)\frac{n-1}{n-p-1} (set to 0 if negative).

Test of Significance in Multiple Linear Regression

Before using an estimated multiple regression equation for decision making, we must test whether the assumed linear relationship is statistically significant. This requires two types of tests: an F-test for overall significance and t-tests for individual significance.

Assumptions for Inference

All hypothesis tests rest on four assumptions about the error term ε\varepsilon (identical to simple linear regression):

  • Mean zero: E(ε)=0E(\varepsilon) = 0 for all combinations of x1,x2,…,xpx_1, x_2, \dots, x_p. Hence E(y)=β0+β1x1+β2x2+⋯+βpxpE(y) = \beta_0 + \beta_1 x_1 + \beta_2 x_2 + \dots + \beta_p x_p.
  • Constant variance: Var(ε)=σ2\text{Var}(\varepsilon) = \sigma^2 for all xx values.
  • Independence: The error terms (and therefore the yy values) for different observations are independent.
  • Normality: ε∼N(0,σ2)\varepsilon \sim N(0,\sigma^2) for all xx values. Since yy is a linear function of ε\varepsilon, yy is also normally distributed.

These assumptions are often satisfied in practice but should be checked before relying on the model.

Overall Significance: The F-Test

The F-test asks: Does the entire set of independent variables collectively explain a significant portion of the variation in yy?

  • Hypotheses H0:β1=β2=⋯=βp=0H_0: \beta_1 = \beta_2 = \dots = \beta_p = 0 Ha:At least one βi≠0H_a: \text{At least one } \beta_i \neq 0

  • Mean squares Recall that a mean square is a sum of squares divided by its degrees of freedom.

    SourceSum of SquaresdfMean Square
    RegressionSSRppMSR=SSRp\displaystyle \text{MSR} = \frac{\text{SSR}}{p}
    ErrorSSEn−p−1n-p-1MSE=SSEn−p−1\displaystyle \text{MSE} = \frac{\text{SSE}}{n-p-1}
    TotalSSTn−1n-1—

    MSE is an unbiased estimator of σ2\sigma^2. Under H0H_0, MSR also estimates σ2\sigma^2; their ratio should be near 1. Under HaH_a, MSR overestimates σ2\sigma^2, so the ratio becomes larger.

  • Test statistic F=MSRMSE∼F(p, n−p−1) under H0F = \frac{\text{MSR}}{\text{MSE}} \quad\sim\quad F(p,\, n-p-1) \text{ under }H_0

  • Decision rule Reject H0H_0 if p-value≤α\text{p-value} \le \alpha.

Worked Example: Hanumantha’s Data

Hanumantha models monthly credit card spend (yy) using annual income (x1x_1) and household size (x2x_2); p=2p = 2. The ANOVA output gives:

  • SSR=91,468,852\text{SSR} = 91,468,852, df =2=2 → MSR=45,734,426\text{MSR} = 45,734,426
  • SSE=22,729,071\text{SSE} = 22,729,071, df =52=52 → MSE=437,098\text{MSE} = 437,098
  • F=45,734,426437,098=104.63F = \frac{45,734,426}{437,098} = 104.63
  • p-value≈0\text{p-value} \approx 0 (at α=0.05\alpha = 0.05)

Reject H0H_0: there is a significant overall linear relationship between monthly spend and the two predictors.

Exam tip: The square root of MSE is the standard error of estimate, s=MSE=661.13s = \sqrt{\text{MSE}} = 661.13 — the estimated standard deviation of the error term. It is often reported in regression output.

Individual Significance: The t-Test

Once the F-test confirms overall significance, we test each coefficient separately to see which predictors matter individually.

  • Hypotheses (for a given βi\beta_i) H0:βi=0Ha:βi≠0H_0: \beta_i = 0 \qquad H_a: \beta_i \neq 0

  • Test statistic t=bis(bi)∼t(n−p−1) under H0t = \frac{b_i}{s(b_i)} \quad\sim\quad t(n-p-1) \text{ under }H_0 where bib_i is the estimated coefficient and s(bi)s(b_i) its standard error.

  • Decision rule Reject H0H_0 if p-value≤α\text{p-value} \le \alpha. Rejection indicates that xix_i contributes significantly to explaining yy after accounting for the other predictors.

Worked Example: Hanumantha’s Data (continued)

The regression output provides:

PredictorCoefficient bib_iStd. Error s(bi)s(b_i)tt statisticp-value
Annual Income (x1x_1)71.176.9271.176.92=10.28\frac{71.17}{6.92}=10.28≈ 0
Household Size (x2x_2)339.9254.05339.9254.05=6.29\frac{339.92}{54.05}=6.29≈ 0

Both p-values are essentially zero → reject both null hypotheses. Conclusion: Both annual income and household size have a statistically significant individual relationship with monthly credit card spend.

Putting It All Together: Testing Workflow

In simple linear regression the F-test and t-test give identical conclusions. In multiple regression they have distinct roles: F tests the whole set; t tests isolate individual contributions.

Key Takeaways

  • Assumptions about ε\varepsilon (mean 0, constant variance, independence, normality) are required for inference.
  • The F-test for overall significance evaluates H0:β1=⋯=βp=0H_0: \beta_1=\cdots=\beta_p=0 using F=MSR/MSEF = \text{MSR}/\text{MSE}.
  • Rejecting the F-test means at least one predictor is linearly related to yy.
  • t-tests for individual significance evaluate H0:βi=0H_0: \beta_i=0 using t=bi/s(bi)t = b_i / s(b_i).
  • In MLR, the F-test and t-tests serve different purposes; both are necessary to fully evaluate the model.
  • MSE provides an unbiased estimate of σ2\sigma^2, and its square root ss is the standard error of estimate.
  • Always interpret the F-test first; only proceed to t-tests if overall significance is found.

Multicollinearity

Multicollinearity refers to the presence of correlation among the independent variables in a multiple linear regression model. Although these variables are called “independent,” in practice they are rarely statistically independent; some degree of correlation is normal. Multicollinearity becomes a problem only when that correlation is high.

Intuition: When two predictors carry nearly the same information, the regression cannot reliably separate their individual effects. The model still fits the data, but the estimated coefficients become unstable and their individual significance tests become misleading.

Diagnosing multicollinearity

The simplest diagnostic is the sample correlation coefficient between pairs of independent variables. A common rule of thumb: if ∣r∣>0.7|r| > 0.7, multicollinearity may cause trouble.

Pair of IVsCorrelation rrTrouble?
Annual income & household size0.320.32No (low)
Monthly income & annual income0.830.83Yes (high)

Consequences of high multicollinearity

Consider the modified dataset with two highly correlated IVs: monthly income (x1x_1) and annual income (x2x_2). The regression output:

  • R2=0.66R^2 = 0.66, overall F-test significant (small pp-value) → model explains some variation.
  • t-tests for individual coefficients:
    • x1x_1 (monthly income): t≈0.18t \approx 0.18, pp large → cannot reject H0:β1=0H_0: \beta_1=0.
    • x2x_2 (annual income): tt large, p≈0p \approx 0 → significant.
  • Coefficient signs:
    • b1=−202.85b_1 = -202.85 (negative, counterintuitive)
    • b2=102.74b_2 = 102.74 (positive, expected)

Interpretation: the negative sign for monthly income is nonsensical — both income measures should positively affect spending. This happens because the model cannot distinguish the separate effect of each correlated variable. Adding monthly income to a model that already contains annual income adds no new information, yet the OLS estimates become distorted.

Key points:

  • Multicollinearity inflates the standard errors of individual coefficients, making t-tests unreliable. A coefficient may appear insignificant even though the variable is genuinely related to YY.
  • In severe cases, coefficients can have the wrong sign.
  • The overall F-test for the regression can remain significant even when none of the individual coefficients are significant.

Exam tip: If the F-test is significant but all or most t-tests are not, suspect multicollinearity. Check pairwise correlations between IVs (∣r∣>0.7|r|>0.7 is a warning). The problem is with individual interpretation, not prediction — multicollinearity does not bias the overall fit or forecasts.

Handling multicollinearity

  • Avoid including highly correlated independent variables in the model. Remove one of the offending variables.
  • In this example, the original model with annual income and household size (r=0.32r=0.32) is preferred.
  • When removal is not feasible (e.g., variables are theoretically important), more advanced remedies exist (e.g., ridge regression, principal components) — but these are beyond this module.

Key takeaways

  • Multicollinearity means independent variables are highly correlated (∣r∣>0.7|r|>0.7 is a common threshold).
  • Main consequence: individual t-tests become unreliable, coefficients can have wrong signs; the F-test and overall R2R^2 are affected much less.
  • It does not invalidate the model for prediction, but it does destroy interpretation of individual effects.
  • Best solution: remove one of the correlated variables from the model.

Residual Analysis in MLR

Residual analysis checks whether the assumptions made about the error term ε\varepsilon (zero mean, constant variance, normality, independence) are valid. In MLR, three types of plots are used:

  1. Residuals vs. each independent variable xix_i
  2. Residuals vs. predicted values y^\hat{y}
  3. Standardized residuals vs. independent variables (or vs. y^\hat{y})

Residual Plots Against Each Independent Variable xix_i

A scatter plot with xix_i on the horizontal axis and the residual (ei=yi−y^ie_i = y_i - \hat{y}_i) on the vertical axis. One plot is generated for every independent variable.

What to look for:

PatternInterpretationAssumption status
Horizontal band (Panel A)Constant variance, model adequateHomoscedasticity holds; model is appropriate
Funnel shape (Panel B)Variability increases with xxHeteroscedasticity — constant variance violated
U‑shape / curve (Panel C)Systematic curvature remains in residualsModel is misspecified — linear form inadequate

Example – Hanumantha’s dataset (monthly spend on income x1x_1 and household size x2x_2):

  • Plot of residuals vs. annual income (x1x_1): resembles a horizontal band, no funnel or U‑shape.
  • Plot of residuals vs. household size (x2x_2): even more clearly a band.

Conclusion: Assumptions of constant variance and linearity appear satisfied.

Exam tip: A band pattern is the “gold standard.” Funnel ⇒ heteroscedasticity (fix via weighted least squares or transformation). U‑shape ⇒ try including x2x^2 or interaction terms.


Residual Plot Against Predicted Values y^\hat{y}

More widely used in MLR because it summarises all independent variables in one plot. y^\hat{y} on the horizontal axis, residuals on the vertical axis. Excel’s Analysis Toolpak does not produce this automatically, but it can be constructed from the outputs.

Hanumantha example: The plot shows a random, band‑like pattern, consistent with the individual xix_i plots — no concern.


Standardized Residual Plot

Standardized residual for observation ii:

standardized residuali=eisei\text{standardized residual}_i = \frac{e_i}{s_{e_i}}

Because residuals from OLS have mean zero, standardisation is done by dividing by an estimate of their standard deviation. Statistical software (including Excel) provides these values.

Why it matters: If the errors ε\varepsilon are normally distributed, the standardized residuals should approximately follow a standard normal distribution. Consequently, about 95% of the standardized residuals should lie between −2-2 and +2+2.

Hanumantha example: The standardized residual plot confirms that almost all points fall inside (−2,+2)(-2, +2); a few are slightly above +2+2 (possible mild outliers). This supports the normality assumption.

Exam tip: A common exam question: “What percentage of standardized residuals do you expect between ±2 under normality?” Answer: ≈95%.

Key Takeaways – Residual Analysis

  • Three plots: residuals vs. xix_i, residuals vs. y^\hat{y}, standardized residuals.
  • Horizontal band ⇒ assumptions hold; funnel ⇒ heteroscedasticity; U‑shape ⇒ model misspecification.
  • Residuals vs. y^\hat{y} is the primary diagnostic in MLR.
  • Standardized residuals outside ±2\pm 2 may indicate outliers or normality violation.
  • Excel provides residuals and standardized residuals automatically.

Outliers and Influential Observations

Outliers

An outlier is a data point that does not follow the pattern of the rest of the data.

Detection:

  • Scatter plots (preliminary visual check)
  • Standardized residuals: observations with ∣standardized residual∣>2|\text{standardized residual}| > 2 are potential outliers (since 95% are inside ±2\pm 2 under normality).

Hanumantha example: A few standardized residuals slightly exceed +2+2.

Action when an outlier is found:

CauseAction
Erroneous data (recording/collection error)Correct the data and re‑run regression
Violation of model assumptionsConsider a different model (e.g., non‑linear, transformations)
Unusual but valid value (by chance)Retain the observation — don’t remove arbitrarily

Influential Observations

An influential observation is one whose removal would substantially change the estimated regression coefficients (slopes and intercept).

Sources of influence:

  • Outlier in the yy direction (large residual)
  • Extreme value of an independent variable xix_i (far from its mean)
  • A combination of both

Detection: Advanced measures (e.g., Cook’s distance, leverage) exist but are not covered in this module. For now, rely on careful examination of any observation flagged as an outlier or extreme.

Action:

  1. Check for data errors → correct if found.
  2. If the observation is valid, it may provide valuable insight into the relationship at the extremes.
  3. Do not remove valid influential observations unless there is clear evidence they do not belong to the population being studied.
  4. If possible, collect additional data at intermediate xx values to better understand the relationship.

Exam tip: “Influential” is not the same as “outlier.” An outlier may have little influence if xx is near the mean; an extreme xx can be influential even if its residual is small.

Key Takeaways – Outliers & Influential Observations

  • Outliers: ∣standardized residual∣>2|\text{standardized residual}| > 2 is a common threshold.
  • Influential observations can change regression results dramatically.
  • Always verify data integrity first.
  • Retain valid observations — do not delete merely to improve fit.
  • Advanced detection methods (Cook’s distance) exist but are beyond this module.

Recap: MLR Fundamentals (as applied to Hanumantha’s Dataset)

Model and Estimation

Multiple linear regression model:

y=β0+β1x1+β2x2+⋯+βpxp+εy = \beta_0 + \beta_1 x_1 + \beta_2 x_2 + \cdots + \beta_p x_p + \varepsilon

Estimated regression equation (using OLS, minimising SSE =∑(yi−y^i)2= \sum (y_i - \hat{y}_i)^2):

y^=b0+b1x1+b2x2+⋯+bpxp\hat{y} = b_0 + b_1 x_1 + b_2 x_2 + \cdots + b_p x_p

Hanumantha example (yy = monthly spend, x1x_1 = annual income (lakhs), x2x_2 = household size):

y^=355.7+71.17 x1+339.92 x2\hat{y} = 355.7 + 71.17\,x_1 + 339.92\,x_2


Goodness of Fit

  • Coefficient of determination R2=SSRSSTR^2 = \dfrac{SSR}{SST}, where SST=SSE+SSRSST = SSE + SSR.
  • Adjusted R2R^2 penalises for extra predictors; in this example R2=0.80R^2 = 0.80, Adjusted R2=0.79R^2 = 0.79.
  • Standard error of estimate S=661.1S = 661.1.

Interpretation: The model with both predictors explains about 80% of the variation in monthly spend — an improvement over simple regression with only income (R2=0.65R^2 = 0.65).


Significance Tests

  • F‑test for overall significance H0:β1=β2=⋯=βp=0H_0: \beta_1 = \beta_2 = \cdots = \beta_p = 0 Test statistic F=MSRMSEF = \dfrac{MSR}{MSE}. Hanumantha: F=104.6F = 104.6, very small p‑value → reject H0H_0; at least one predictor is significant.

  • t‑tests for individual coefficients H0:βi=0H_0: \beta_i = 0, test statistic t=biSE(bi)t = \dfrac{b_i}{SE(b_i)}, df =n−p−1= n-p-1. Both p‑values near zero → both x1x_1 and x2x_2 have statistically significant linear relationships with yy.


Assumptions and Diagnostics

  • Assumptions about ε\varepsilon: mean 0, constant variance (homoscedasticity), normal distribution, independence.
  • Multicollinearity: high correlation among xix_i’s makes interpretation difficult. In Hanumantha’s data, no problem with x1x_1 and x2x_2; but adding a third variable (monthly income) illustrated multicollinearity in action.
  • Residual analysis: plots showed no violations; standardized residuals mostly inside ±2\pm 2.

Key Takeaways – Recap

  • MLR uses OLS to estimate pp slope coefficients.
  • R2R^2 and adjusted R2R^2 measure explanatory power; FF‑test checks overall significance; tt‑tests check individual predictors.
  • Always verify residual assumptions (plots, standardized residuals).
  • Multicollinearity can distort signs and p‑values; examine correlation among predictors.
  • Hanumantha’s example: clean model with good fit, significant predictors, and valid residuals.

Applications with Examples

Multiple linear regression (MLR) extends simple linear regression by using two or more independent variables to predict a dependent variable. The goal is to capture more of the variation in YY by including additional predictors.

Model form: Y=β0+β1X1+β2X2+εY = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + \varepsilon

The estimated regression equation is: Y^=b0+b1X1+b2X2\hat{Y} = b_0 + b_1 X_1 + b_2 X_2

Interpretation of coefficients:

  • b1b_1: expected change in YY for a one-unit increase in X1X_1, holding X2X_2 constant.
  • b2b_2: expected change in YY for a one-unit increase in X2X_2, holding X1X_1 constant.

Two key hypothesis tests are used:

  1. Overall F-test – tests whether at least one coefficient is non-zero (H0:β1=β2=0H_0: \beta_1 = \beta_2 = 0).
  2. Individual t-tests – tests whether a specific coefficient is zero (H0:βi=0H_0: \beta_i = 0).

Exam tip: The phrase “holding constant” is the core of MLR interpretation. Always include it when explaining a slope coefficient.

Example 1: Predicting Household Debt from First Income and Monthly Payment

Context: Manjula Nayak’s Bengaluru household survey (n ≈ 500). Dependent variable: debt (₹). Independent variables: X1X_1 = first income, X2X_2 = monthly payment.

Regression output:

MeasureValue
R2R^20.45
Adjusted R2R^20.448
Standard error25644
b0b_013081
b1b_1 (income)0.025
b2b_2 (monthly payment)2.09
SE(b1)SE(b_1)0.003
SE(b2)SE(b_2)0.188
tt-stat for b1b_18.74 (p ≈ 0)
tt-stat for b2b_211.1 (p ≈ 0)
F-statistic203 (p ≈ 0)
SSRSSR2.68×10112.68 \times 10^{11}
SSESSE3.27×10113.27 \times 10^{11}
SSTSST5.94×10115.94 \times 10^{11}
MSRMSR1.34×10111.34 \times 10^{11}
MSEMSE6.576×1086.576 \times 10^8

Interpretation:

  • R2R^2 = 0.45: The model explains 45% of the variation in debt. This is higher than the simple regression R2R^2 for income alone (0.31) or monthly payment alone (0.37).
  • Adjusted R2R^2 ≈ R2R^2 because of large sample size (n=500).
  • Standard error = 25644 – smaller than the simple regression standard errors (~28,619 for income; ~27,518 for monthly payment).
  • Coefficient b1b_1 = 0.025: For every ₹1 increase in first income, holding monthly payment constant, debt increases by ₹0.025.
  • Coefficient b2b_2 = 2.09: For every ₹1 increase in monthly payment, holding income constant, debt increases by ₹2.09.
  • Both t-tests reject H0H_0 (p ≈ 0): both predictors have a significant linear relationship with debt.
  • F-test rejects H0H_0 (p ≈ 0): at least one predictor is significant; the overall model is useful.

Residual analysis:

  • Standardized residuals: most within ±2, a few outside.
  • Residual vs. X1X_1 and residual vs. X2X_2 plots show an oval pattern – not a perfect random band, but no funnel shape or curvature (acceptable).
  • Residual vs. fitted plot similar – no major violations.

Key takeaways – Model 1

  • Two predictors explain more variation than either alone (R2R^2 increased).
  • Both income and monthly payment are significant individual predictors.
  • Model still leaves 55% of variation unexplained – debt depends on other factors.
  • Residual plots are adequate; no strong evidence of heteroscedasticity or nonlinearity.

Example 2: Predicting Household Debt from Monthly Payment and Utilities

Context: Same dataset, different predictors: X1X_1 = monthly payment, X2X_2 = utilities.

Regression output:

MeasureValue
R2R^20.67
Adjusted R2R^2≈0.67
Standard error19839
b0b_0-140771
b1b_1 (monthly payment)1.44
b2b_2 (utilities)168.91
SE(b1)SE(b_1)0.143
SE(b2)SE(b_2)7.87
tt-stat for b1b_110.02 (p ≈ 0)
tt-stat for b2b_221.47 (p ≈ 0)
F-statistic506.9 (p ≈ 0)
SSRSSR3.99×10113.99 \times 10^{11}
SSESSE1.96×10111.96 \times 10^{11}
SSTSST5.94×10115.94 \times 10^{11}
MSRMSR1.99×10111.99 \times 10^{11}
MSEMSE3.936×1083.936 \times 10^8

Interpretation:

  • R2R^2 = 0.67: Model explains 67% of debt variation – a clear improvement over Model 1 (0.45) and over the simple regressions (0.31 for income, 0.37 for monthly payment).
  • Standard error = 19839 – lower than Model 1 (25644), indicating more precise predictions.
  • Coefficient b1b_1 = 1.44: For every ₹1 increase in monthly payment, holding utilities constant, debt increases by ₹1.44.
  • Coefficient b2b_2 = 168.91: For every ₹1 increase in utilities, holding monthly payment constant, debt increases by ₹168.91.
  • Both t-tests significant (p ≈ 0).
  • F-statistic = 506.9 (p ≈ 0) – overall model highly significant.

Comparison of SSR and SSE:

  • SSRSSR increased from 2.68×10112.68 \times 10^{11} (Model 1) to 3.99×10113.99 \times 10^{11} (Model 2).
  • SSESSE decreased from 3.27×10113.27 \times 10^{11} to 1.96×10111.96 \times 10^{11}.
  • SSTSST remains the same (5.94×10115.94 \times 10^{11}) because total variation in debt is fixed.

Exam tip: SSTSST is constant across models for the same YY; changes in R2R^2 are driven entirely by SSRSSR (or equivalently SSESSE).

Residual analysis:

  • Standardized residuals: mostly within ±2, a few outliers.
  • Residual vs. X1X_1, vs. X2X_2, vs. fitted values all show random scatter – no patterns (oval, funnel, or curvature).
  • This model appears more reliable than Model 1 in terms of residual diagnostics.

Important caveat: Regression establishes association, not causality. A higher utility bill does not necessarily cause more debt; it may reflect a larger house or higher usage, which correlates with debt.

Key takeaways – Model 2

  • Model with monthly payment and utilities performs better (R2=0.67R^2=0.67, lower SESE) than the model with income and monthly payment.
  • Both predictors are highly significant individually and jointly.
  • Residual plots show random scatter – no serious violations of linear regression assumptions.
  • Better R2R^2 does not mean causality; always interpret associations cautiously.
  • The choice of independent variables strongly affects model performance.

Three-Variable Multiple Linear Regression (Household Debt Data)

After observing that two different two-variable models (monthly payment + utilities; first income + monthly payment) each yielded significant relationships, a model combining all three predictors is tested: debt yy as a function of first income x1x_1, monthly payment x2x_2, and utilities x3x_3.

The estimated regression equation from Excel (Analysis ToolPak) is:

y^=−140378+0.017x1+0.96x2+157.95x3\hat{y} = -140378 + 0.017 x_1 + 0.96 x_2 + 157.95 x_3

Model Fit Statistics

StatisticValueComparison to Best Two-Variable Model (x2,x3x_2,x_3)
R2R^20.708Higher (0.67 previously)
Adjusted R2R^2~0.70 (slightly less than R2R^2)Close to R2R^2
Standard Error18,699Lower (19,839 previously)
SSRSSR4.21×10114.21 \times 10^{11}Higher (3.99×10113.99 \times 10^{11})
SSESSE1.73×10111.73 \times 10^{11}Lower (1.96×10111.96 \times 10^{11})
SSTSST5.94×10115.94 \times 10^{11} (unchanged)—

The three-variable model explains 70% of the variation in household debt, the best fit so far.

ANOVA and Overall Significance

  • Degrees of freedom: regression = 3, error = 496.
  • MSR=SSR3=1.40×1011MSR = \frac{SSR}{3} = 1.40 \times 10^{11}, MSE=SSE496=349,653,451MSE = \frac{SSE}{496} = 349,653,451.
  • F=MSRMSE=401.6F = \frac{MSR}{MSE} = 401.6, p≈0p \approx 0.

Conclusion: Reject H0:β1=β2=β3=0H_0: \beta_1 = \beta_2 = \beta_3 = 0. There is significant overall linear relationship.

Coefficient Estimates and Individual t-Tests

VariableCoefficientStd. Errort-statisticp-value
Intercept (b0b_0)-140,378———
First income (b1b_1)0.0170.0027.96~0
Monthly payment (b2b_2)0.960.1486.5~0
Utilities (b3b_3)157.957.5420.95~0

All three coefficients are statistically significant (p ≈ 0). Interpretation (ceteris paribus):

  • For every ₹1 increase in first income, debt increases by ₹0.017.
  • For every ₹1 increase in monthly payment, debt increases by ₹0.96.
  • For every ₹1 increase in utilities, debt increases by ₹157.95.

Exam tip: In multiple regression, the interpretation of a coefficient is "holding all other predictors constant." This is the key distinction from simple regression.

Residual Analysis

  • Standardized residuals mostly within ±3 range; a few points exceed ±3 (possible outliers) — more visible because the model fits better and the standard error is lower.
  • Plots of residuals against x1x_1, x2x_2, x3x_3, and y^\hat{y} all show random scatter above and below zero → no obvious violation of linearity or homoscedasticity.

Overall, the three-variable model is the strongest so far for this dataset.


Simple Linear Regression on Real Estate Data (Property Prices)

Dataset: 30 flats in Jayalakshmi Puram, Mysuru. Variables include selling price yy (in lakhs), area in sq. ft. xx, number of bedrooms/bathrooms, premium locality, gated complex, amenities. Only area is used in this simple linear regression (SLR).

Scatter Plot and Trend Line

Rough positive linear trend with several potential outliers. Trend line equation:

y^=28.37+0.0968⋅x\hat{y} = 28.37 + 0.0968 \cdot x

R2=0.44R^2 = 0.44 — only 44% of variation in selling price explained by area.

Regression Output (ToolPak)

StatisticValue
R2R^20.44
Adjusted R2R^20.43
Standard Error (ss)70.57 lakhs
SSRSSR111,539
SSESSE139,462
SSTSST251,001
MSRMSR111,539
MSEMSE4,981
FF (1,28)22.39, p≈0p \approx 0

Coefficients and t-Test

VariableCoefficientStd. Errort-statisticp-value
Intercept (b0b_0)28.37———
Area (b1b_1)0.0960.024.73~0

Interpretation: For every 1 sq. ft. increase in area, selling price increases by 0.096 lakhs (₹9,600). The relationship is statistically significant.

Exam tip: The large coefficient magnitude (₹9,600 per sq. ft.) results from having only one predictor. Adding other relevant variables will likely change this coefficient (and reduce omitted variable bias).

Residual Analysis

  • Standardized residuals: a few points beyond ±3 (possible outliers).
  • Plots of residuals vs. xx and vs. y^\hat{y} show random scatter → no major pattern.

This SLR is a starting point. A multiple linear regression with additional predictors (bedrooms, bathrooms, locality, etc.) could improve the model.


Key Takeaways

  • Adding more relevant predictors can increase R2R^2 and lower standard error, as seen in the debt model (three variables outperformed two).
  • In multiple regression, each coefficient is interpreted "holding all other variables constant."
  • The FF-test checks overall significance; t-tests check individual significance.
  • Residual plots should be random; points beyond ±3 standardized residuals may be outliers.
  • Simple linear regression can underfit; multiple regression often provides better explanation and more reliable coefficient estimates.

Exam tip: When comparing models, prefer adjusted R2R^2 over plain R2R^2 because it penalizes adding unnecessary variables. Also, always check residual plots for violations.

MLR Application – Real Estate Example

We aim to improve the simple linear regression model that predicts flat selling price (₹ lakhs) using only area (sq ft) by adding other quantitative variables: number of bedrooms and number of bathrooms. The dataset contains 30 flats from Jayalakshmi Puram, Mysore. The baseline model:

y^=28.37+0.0968x(area)\hat{y} = 28.37 + 0.0968x \quad\text{(area)}

R2=0.44R^2 = 0.44: only 44% of variation in price is explained by area alone. Both the overall F‑test and the t‑test confirmed a significant linear relationship.

Baseline Simple Linear Regression Recap

MetricValue
R2R^20.44
Standard error70.57
SSR111,539
SSE139,461
SST251,001

Model 1: Price = f(Area, Bedrooms)

Estimated Equation

y^=−29.19+0.085×Area+23.46×Bedrooms\hat{y} = -29.19 + 0.085 \times \text{Area} + 23.46 \times \text{Bedrooms}

Regression Summary

MetricValue
R2R^20.46
Adjusted R2R^20.42
Standard error70.74
SSR115,892
SSE135,109
SST251,001
MSR57,946
MSE5,004
F-statistic11.58
p-value (F)0.0002

Coefficient Table

PredictorCoef.Std. Errort-statp-value
Intercept−29.19———
Area (sq ft)0.0850.0243.460.002
Bedrooms23.4625.150.9320.36

Interpretation

  • Overall F‑test: H0:β1=β2=0H_0: \beta_1 = \beta_2 = 0 is rejected (p=0.0002p=0.0002) → at least one predictor is useful.
  • t‑test for Bedrooms: p=0.36p=0.36 → cannot reject H0:β2=0H_0: \beta_2 = 0. Adding bedrooms does not significantly improve the model once area is already included.
  • Residual plots showed no alarming patterns (two standardized residuals >2, otherwise fine).

Why does bedrooms become insignificant?

The correlation between area and bedrooms is 0.542. Although this is below the common rule‑of‑thumb threshold of 0.7, multicollinearity still inflates the standard error of the bedrooms coefficient, making its t‑test non‑significant. The diagnostic tells us: given area, bedrooms add little unique information.


Model 2: Price = f(Area, Bathrooms)

Estimated Equation

y^=10.61+0.084×Area+17.80×Bathrooms\hat{y} = 10.61 + 0.084 \times \text{Area} + 17.80 \times \text{Bathrooms}

Regression Summary

MetricValue
R2R^20.45
Adjusted R2R^20.41
Standard error71.18
SSR114,185
SSE136,816
SST251,001
MSR57,092
MSE5,067
F-statistic11.26
p-value (F)0.0003

Coefficient Table

PredictorCoef.Std. Errort-statp-value
Intercept10.61———
Area (sq ft)0.084——0.005
Bathrooms17.80—0.7230.48

Use the reported t-statistics and p-values; the standard errors are not required for this comparison.

Observations

  • R2R^2 (0.45) is barely different from the simple model (0.44) and slightly worse than Model 1 (0.46).
  • Again the t‑test for bathrooms is non‑significant (p=0.48p=0.48), while the overall F‑test remains significant.
  • Correlation between area and bathrooms is 0.659 – higher than before, confirming multicollinearity.

Diagnosing Multicollinearity

Multicollinearity occurs when predictor variables are highly correlated, making it difficult to isolate their individual effects on the response.

PairCorrelation
Area ↔ Bedrooms0.542
Area ↔ Bathrooms0.659

Exam tip: A non‑significant t‑test for an individual coefficient in a multiple regression does not always mean the variable is irrelevant. When predictors are correlated, the shared information can cause inflated standard errors. The overall F‑test may still be significant, indicating the model as a whole is useful.

Key Takeaways from the two models:

  • Adding bedrooms or bathrooms to a model that already includes area yields only marginal improvement in R2R^2 (0.44 → 0.46 or 0.45).
  • The F‑test for overall significance is significant in both models, confirming that at least one predictor (area) is important.
  • The individual t‑tests for bedrooms and bathrooms are non‑significant because of multicollinearity with area.
  • Multicollinearity can exist even with correlations below 0.7; diagnostics (t‑tests, VIF) are needed.
  • With only these quantitative variables exhausted, the next step is to incorporate qualitative variables (premium locality, gated community, amenities) – covered in the next module.

What did we learn?

  • Adding more variables does not guarantee a better model; multicollinearity can limit the gain.
  • Use the combination of overall F‑test and individual t‑tests together to guide variable selection.
  • When a variable is redundant given others, remove it to keep the model parsimonious.

Intercept (β0\beta_0) – Why It’s Often Ignored

Most regression output includes a test for β0\beta_0: H0:β0=0vs.Ha:β0≠0H_0: \beta_0 = 0 \quad \text{vs.} \quad H_a: \beta_0 \neq 0 However, this hypothesis is rarely of practical interest.

  • Inferences on β0\beta_0 are special cases of inferences on the fitted value of the regression line as a whole – concepts like confidence intervals for predicted and estimated values of YY.
  • The pp-value for β0\beta_0 in standard computer output is usually not important and is often ignored.

Extrapolation risk: If the training data do not include cases where all Xj=0X_j = 0, then using the equation to predict YY at X=0X=0 is extrapolation – outside the range of observed data. Such predictions are unreliable.

Exam tip: Never interpret β0\beta_0 as a meaningful baseline unless the data actually contain observations near X=0X=0. The intercept is often a mathematical necessity, not a substantive parameter.

Omitted Variable Bias (OVB)

OVB occurs when a regression model excludes one or more relevant independent variables. The included variables then absorb some of the effect of the omitted variables, leading to biased coefficients (over‑ or underestimation of true contributions).

Conditions for OVB

For omitted variable bias to exist, both of the following must hold:

  1. The omitted variable is correlated with the dependent variable YY.
  2. The omitted variable is correlated with at least one included independent variable.

Example (household debt data):

  • YY = debt, X1X_1 = income, X2X_2 = monthly payments, X3X_3 = utilities.
  • Simple linear regression (debt ~ income only) yields R2=0.31R^2 = 0.31.
  • Correlations (from sample):
VariableCorrelation with Debt
Income (X1X_1)0.56
Monthly payments (X2X_2)0.61
Utilities (X3X_3)0.78
  • Correlations among XX variables:
    • Monthly payments & income: 0.55
    • Utilities & income: 0.39
    • Utilities & monthly payments: 0.49

Both conditions hold → OVB present.

How to detect OVB

  • R2R^2 increases significantly when omitted variables are added. (Example: 0.31 → 0.71)
  • Standard error drops substantially.
  • pp-values of included variables drop (become more significant) after adding omitted ones.

Quantifying OVB

Coefficient bias can be estimated (details beyond this module), but the key diagnostic is the improvement in model fit when the omitted variables are included.

Key takeaways

  • The intercept β0\beta_0 test (H0:β0=0H_0:\beta_0=0) is usually unimportant; avoid extrapolating to X=0X=0 outside the data range.
  • Omitted variable bias arises when an excluded variable correlates with both YY and included XXs.
  • Detect OVB by adding omitted variables and observing a large jump in R2R^2, a drop in standard errors, and improved pp-values.
  • OVB causes included coefficients to be biased – the included variable “soaks up” effects that belong to the omitted one.

Multiple Linear Regression – Module Recap

The module covered the full workflow for building and validating an MLR model.

Model form Y=β0+β1X1+β2X2+⋯+βpXp+εY = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + \cdots + \beta_p X_p + \varepsilon E(Y)=β0+β1X1+β2X2+⋯+βpXpE(Y) = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + \cdots + \beta_p X_p

Estimated regression equation Y^=b0+b1X1+b2X2+⋯+bpXp\hat{Y} = b_0 + b_1 X_1 + b_2 X_2 + \cdots + b_p X_p

Key steps covered in this module:

  • Least squares estimation – finding bjb_j to minimize ∑(Yi−Y^i)2\sum (Y_i - \hat{Y}_i)^2.
  • Multiple coefficient of determination (R2R^2) – measure of overall fit.
  • Assumptions on the error term ε\varepsilon – zero mean, constant variance, independence, normality.
  • F-test for overall significance of the entire regression.
  • t-tests for individual parameters βj\beta_j (for j≥1j \ge 1).
  • Multicollinearity – its detection and consequences.
  • Omitted variable bias – conditions, detection, impact.
  • Residual analysis – identifying outliers and influential observations.
  • Iterative model building – comparing and improving models using tools like Analysis ToolPak.

The next module will extend MLR to include categorical variables, followed by a course summary.

Key takeaways

  • MLR models a linear relationship between YY and multiple XXs.
  • The estimated coefficients bjb_j are sample statistics estimating population parameters βj\beta_j.
  • Model validation uses R2R^2, F-test, t-tests, and residual analysis.
  • Omitted variable bias and multicollinearity are critical pitfalls to check during model building.