Multiple Linear Regression - I
Multiple linear regression (MLR) extends simple linear regression to model a single dependent variable using two or more independent variables . The intuition: real-world outcomes rarely depend on just one factor. By including more relevant predictors, MLR typically yields better predictions than simple regression.
The MLR Model
The multiple regression model describes how relates to the independent variables plus a random error term:
where:
- are parameters (unknown population coefficients).
- is the error term – a random variable that captures variability in not explained by the linear combination of the 's.
- denotes the number of independent variables.
Key assumption
The expected value of the error term is zero: . Consequently, the mean or expected value of is:
This equation is called the multiple regression equation. If the parameters were known, we could compute for any given values of .
Estimated regression equation
In practice, parameters are unknown. Using a sample (training data), we obtain point estimators for . The estimated multiple regression equation is:
is the predicted value of the dependent variable.
Exam tip: The structure of MLR is a direct generalization of simple linear regression. Expect to see the same core concepts (estimators, residuals, hypothesis tests) extended to multiple predictors.
Key takeaways – MLR model
- = number of independent variables.
- Model: with .
- Regression equation: .
- Estimated equation: , where 's are sample estimates.
Steps for Developing an MLR Model
Model building is an iterative process; many steps may need revisiting. The following ten steps are typical:
-
Collect data – Identify relevant independent variables and gather data from secondary sources (e.g., ERP systems, databases, government census) or primary sources (surveys, interviews).
-
Preprocess data – Assess completeness and correctness. Handle missing observations and transform qualitative/categorical variables (e.g., through dummy coding).
-
Descriptive analytics – Compute basic statistics and use visualizations (scatter plots, box plots) to examine correlations, flag potential multicollinearity and overfitting. Identify proxy variables if original variables cannot be used (ethical or regulatory reasons).
-
Modeling strategy – Select the set of independent variables. Desirable properties:
- Variables should be independent of each other.
- Model performance should be consistent across training and test data.
- Variables should be controllable by the decision maker (e.g., price is controllable; weather is not). A model with uncontrollable variables may be less actionable.
-
Split data – Divide into training and validation sets. Multiple subsets can be used for cross-validation.
-
Define functional form – Start with a linear relationship between and the 's. Nonlinear transformations (e.g., ) are allowed and do not make the regression non‑linear in parameters – they are still estimated with OLS.
-
Estimate parameters – Use ordinary least squares (OLS) to find that minimize the sum of squared residuals. OLS provides the Best Linear Unbiased Estimator (BLUE) under standard assumptions.
-
Model diagnostics – Check statistical significance:
- F‑test – tests overall model significance (whether all coefficients except are jointly zero).
- t‑test – tests significance of individual coefficients. Additional diagnostics for MLR include checking normality of residuals, multicollinearity, and heteroscedasticity. If assumptions are violated, take remedial measures.
-
Validate model – Use the validation data to assess performance. Common metrics:
- (coefficient of determination)
- Adjusted (penalizes for number of predictors)
- Mean absolute percentage error (MAPE)
- Root mean square error (RMSE) Models with low bias and low variance on both training and validation sets are preferred.
-
Deploy model – Generate actionable insights and create an implementation plan. Continuously monitor performance; models may need to be updated over time.
Exam tip: The F‑test and t‑test in MLR serve the same roles as in simple regression, but now the F‑test checks the entire set of predictors simultaneously. Multicollinearity is a new issue unique to MLR – watch for high correlation among independent variables.
Key takeaways – model development
- MLR development is iterative; steps are often repeated.
- Data quality, preprocessing, and descriptive analytics are critical.
- OLS yields BLUE estimates.
- Diagnostics include F‑test (overall), t‑tests (individual), normality, multicollinearity, heteroscedasticity.
- Validation uses , adjusted , MAPE, RMSE; aim for consistent performance.
- Deployment requires ongoing monitoring.
OLS in Multiple Linear Regression
Ordinary Least Squares (OLS) extends from simple regression to the multiple regression setting. The principle is identical: minimize the sum of squared deviations between observed and predicted values of the dependent variable.
Minimization criterion
- = observed value of the dependent variable for observation
- = predicted value from the estimated multiple regression equation
Estimated regression equation
The coefficients are obtained by applying OLS to sample data. In simple regression closed-form formulas exist; in multiple linear regression (MLR) the solution uses matrix algebra (covered via software). Interpretation of output is the primary focus.
Interpreting coefficients in MLR vs. simple regression
In simple regression, is the estimated change in for a one‑unit change in . In MLR, the interpretation becomes conditional:
represents the estimated change in corresponding to a one‑unit change in when all other independent variables are held constant.
Worked example: Hanumantha’s credit card data
- observations
- = monthly credit card spend ()
- = annual household income (₹ lakh)
- = household size (persons)
Simple regression with only income:
- , → significant
Simple regression with only household size:
- , → significant
Multiple regression with both predictors:
Notice that dropped from (simple) to (MLR) because the effect of income is now estimated holding household size fixed.
| Coefficient | Simple (only income) | MLR (income + size) | Interpretation in MLR |
|---|---|---|---|
| (income) | For a ₹1 lakh increase in income, monthly spend increases by ₹71.16, assuming household size remains constant. | ||
| (size) | For an increase of one person in household size, monthly spend increases by ₹339.90, assuming income remains constant. |
Exam tip: The conditional interpretation (“holding other variables constant”) is the key distinction between simple and multiple regression coefficients. Always phrase it explicitly.
Key takeaways – OLS in MLR
- OLS minimises , whether for one or many predictors.
- Coefficients are computed via matrix algebra; focus on interpretation of software output.
- In MLR, measures the change in per unit change in with all other ’s held constant.
- Adding a relevant predictor can change the magnitude (and sometimes sign) of existing coefficients – do not expect them to stay the same.
Multiple Coefficient of Determination
The multiple coefficient of determination measures the overall fit of the estimated regression equation.
Sum of squares decomposition
| Term | Definition | Formula |
|---|---|---|
| Total sum of squares (total variability in ) | ||
| Sum of squares due to regression (explained variability) | ||
| Sum of squares due to error (unexplained variability) |
depends only on , not on the model. Adding predictors typically increases and decreases , improving fit.
in simple vs. multiple regression
Interpretation: proportion of the variability in explained by the estimated regression equation (multiply by 100 for percentage).
Hanumantha’s example – comparison:
| Model | ||||
|---|---|---|---|---|
| Simple (income only) | 74,180,516 | 40,017,408 | 114,197,923 | 0.65 (65%) |
| MLR (income + size) | 91,468,852 | 22,729,070 | 114,197,923 | 0.80 (80%) |
increased from 0.65 to 0.80 – the two predictors together explain 80% of the variation in monthly spend.
Why always increases (or stays the same) with more predictors
Adding an independent variable cannot worsen the fit because OLS can always set its coefficient to zero, leaving unchanged. In practice, almost always increases.
Adjusted – penalizing for extra variables
The adjusted modifies to account for the number of predictors, preventing overestimation of explanatory power.
- = number of observations
- = number of independent variables
If is large relative to , the adjusted can become negative – in that case set it to 0.
Hanumantha’s example – adjusted :
- Simple regression (p=1):
- MLR (p=2):
The adjusted (0.793) is only slightly lower than (0.80), suggesting the extra variable is genuinely useful.
Exam tip: alone can mislead – always report adjusted when comparing models with different numbers of predictors.
Key takeaways – Multiple coefficient of determination
- ; it is the proportion of ‑variation explained by the model.
- is unaffected by the model; adding predictors increases and decreases , pushing upward.
- Adjusted penalises for adding useless predictors; use it to avoid overfitting.
- Formula: (set to 0 if negative).
Test of Significance in Multiple Linear Regression
Before using an estimated multiple regression equation for decision making, we must test whether the assumed linear relationship is statistically significant. This requires two types of tests: an F-test for overall significance and t-tests for individual significance.
Assumptions for Inference
All hypothesis tests rest on four assumptions about the error term (identical to simple linear regression):
- Mean zero: for all combinations of . Hence .
- Constant variance: for all values.
- Independence: The error terms (and therefore the values) for different observations are independent.
- Normality: for all values. Since is a linear function of , is also normally distributed.
These assumptions are often satisfied in practice but should be checked before relying on the model.
Overall Significance: The F-Test
The F-test asks: Does the entire set of independent variables collectively explain a significant portion of the variation in ?
-
Hypotheses
-
Mean squares Recall that a mean square is a sum of squares divided by its degrees of freedom.
Source Sum of Squares df Mean Square Regression SSR Error SSE Total SST — MSE is an unbiased estimator of . Under , MSR also estimates ; their ratio should be near 1. Under , MSR overestimates , so the ratio becomes larger.
-
Test statistic
-
Decision rule Reject if .
Worked Example: Hanumantha’s Data
Hanumantha models monthly credit card spend () using annual income () and household size (); . The ANOVA output gives:
- , df →
- , df →
- (at )
Reject : there is a significant overall linear relationship between monthly spend and the two predictors.
Exam tip: The square root of MSE is the standard error of estimate, — the estimated standard deviation of the error term. It is often reported in regression output.
Individual Significance: The t-Test
Once the F-test confirms overall significance, we test each coefficient separately to see which predictors matter individually.
-
Hypotheses (for a given )
-
Test statistic where is the estimated coefficient and its standard error.
-
Decision rule Reject if . Rejection indicates that contributes significantly to explaining after accounting for the other predictors.
Worked Example: Hanumantha’s Data (continued)
The regression output provides:
| Predictor | Coefficient | Std. Error | statistic | p-value |
|---|---|---|---|---|
| Annual Income () | 71.17 | 6.92 | ≈ 0 | |
| Household Size () | 339.92 | 54.05 | ≈ 0 |
Both p-values are essentially zero → reject both null hypotheses. Conclusion: Both annual income and household size have a statistically significant individual relationship with monthly credit card spend.
Putting It All Together: Testing Workflow
In simple linear regression the F-test and t-test give identical conclusions. In multiple regression they have distinct roles: F tests the whole set; t tests isolate individual contributions.
Key Takeaways
- Assumptions about (mean 0, constant variance, independence, normality) are required for inference.
- The F-test for overall significance evaluates using .
- Rejecting the F-test means at least one predictor is linearly related to .
- t-tests for individual significance evaluate using .
- In MLR, the F-test and t-tests serve different purposes; both are necessary to fully evaluate the model.
- MSE provides an unbiased estimate of , and its square root is the standard error of estimate.
- Always interpret the F-test first; only proceed to t-tests if overall significance is found.
Multicollinearity
Multicollinearity refers to the presence of correlation among the independent variables in a multiple linear regression model. Although these variables are called “independent,” in practice they are rarely statistically independent; some degree of correlation is normal. Multicollinearity becomes a problem only when that correlation is high.
Intuition: When two predictors carry nearly the same information, the regression cannot reliably separate their individual effects. The model still fits the data, but the estimated coefficients become unstable and their individual significance tests become misleading.
Diagnosing multicollinearity
The simplest diagnostic is the sample correlation coefficient between pairs of independent variables. A common rule of thumb: if , multicollinearity may cause trouble.
| Pair of IVs | Correlation | Trouble? |
|---|---|---|
| Annual income & household size | No (low) | |
| Monthly income & annual income | Yes (high) |
Consequences of high multicollinearity
Consider the modified dataset with two highly correlated IVs: monthly income () and annual income (). The regression output:
- , overall F-test significant (small -value) → model explains some variation.
- t-tests for individual coefficients:
- (monthly income): , large → cannot reject .
- (annual income): large, → significant.
- Coefficient signs:
- (negative, counterintuitive)
- (positive, expected)
Interpretation: the negative sign for monthly income is nonsensical — both income measures should positively affect spending. This happens because the model cannot distinguish the separate effect of each correlated variable. Adding monthly income to a model that already contains annual income adds no new information, yet the OLS estimates become distorted.
Key points:
- Multicollinearity inflates the standard errors of individual coefficients, making t-tests unreliable. A coefficient may appear insignificant even though the variable is genuinely related to .
- In severe cases, coefficients can have the wrong sign.
- The overall F-test for the regression can remain significant even when none of the individual coefficients are significant.
Exam tip: If the F-test is significant but all or most t-tests are not, suspect multicollinearity. Check pairwise correlations between IVs ( is a warning). The problem is with individual interpretation, not prediction — multicollinearity does not bias the overall fit or forecasts.
Handling multicollinearity
- Avoid including highly correlated independent variables in the model. Remove one of the offending variables.
- In this example, the original model with annual income and household size () is preferred.
- When removal is not feasible (e.g., variables are theoretically important), more advanced remedies exist (e.g., ridge regression, principal components) — but these are beyond this module.
Key takeaways
- Multicollinearity means independent variables are highly correlated ( is a common threshold).
- Main consequence: individual t-tests become unreliable, coefficients can have wrong signs; the F-test and overall are affected much less.
- It does not invalidate the model for prediction, but it does destroy interpretation of individual effects.
- Best solution: remove one of the correlated variables from the model.
Residual Analysis in MLR
Residual analysis checks whether the assumptions made about the error term (zero mean, constant variance, normality, independence) are valid. In MLR, three types of plots are used:
- Residuals vs. each independent variable
- Residuals vs. predicted values
- Standardized residuals vs. independent variables (or vs. )
Residual Plots Against Each Independent Variable
A scatter plot with on the horizontal axis and the residual () on the vertical axis. One plot is generated for every independent variable.
What to look for:
| Pattern | Interpretation | Assumption status |
|---|---|---|
| Horizontal band (Panel A) | Constant variance, model adequate | Homoscedasticity holds; model is appropriate |
| Funnel shape (Panel B) | Variability increases with | Heteroscedasticity — constant variance violated |
| U‑shape / curve (Panel C) | Systematic curvature remains in residuals | Model is misspecified — linear form inadequate |
Example – Hanumantha’s dataset (monthly spend on income and household size ):
- Plot of residuals vs. annual income (): resembles a horizontal band, no funnel or U‑shape.
- Plot of residuals vs. household size (): even more clearly a band.
Conclusion: Assumptions of constant variance and linearity appear satisfied.
Exam tip: A band pattern is the “gold standard.” Funnel ⇒ heteroscedasticity (fix via weighted least squares or transformation). U‑shape ⇒ try including or interaction terms.
Residual Plot Against Predicted Values
More widely used in MLR because it summarises all independent variables in one plot. on the horizontal axis, residuals on the vertical axis. Excel’s Analysis Toolpak does not produce this automatically, but it can be constructed from the outputs.
Hanumantha example: The plot shows a random, band‑like pattern, consistent with the individual plots — no concern.
Standardized Residual Plot
Standardized residual for observation :
Because residuals from OLS have mean zero, standardisation is done by dividing by an estimate of their standard deviation. Statistical software (including Excel) provides these values.
Why it matters: If the errors are normally distributed, the standardized residuals should approximately follow a standard normal distribution. Consequently, about 95% of the standardized residuals should lie between and .
Hanumantha example: The standardized residual plot confirms that almost all points fall inside ; a few are slightly above (possible mild outliers). This supports the normality assumption.
Exam tip: A common exam question: “What percentage of standardized residuals do you expect between ±2 under normality?” Answer: ≈95%.
Key Takeaways – Residual Analysis
- Three plots: residuals vs. , residuals vs. , standardized residuals.
- Horizontal band ⇒ assumptions hold; funnel ⇒ heteroscedasticity; U‑shape ⇒ model misspecification.
- Residuals vs. is the primary diagnostic in MLR.
- Standardized residuals outside may indicate outliers or normality violation.
- Excel provides residuals and standardized residuals automatically.
Outliers and Influential Observations
Outliers
An outlier is a data point that does not follow the pattern of the rest of the data.
Detection:
- Scatter plots (preliminary visual check)
- Standardized residuals: observations with are potential outliers (since 95% are inside under normality).
Hanumantha example: A few standardized residuals slightly exceed .
Action when an outlier is found:
| Cause | Action |
|---|---|
| Erroneous data (recording/collection error) | Correct the data and re‑run regression |
| Violation of model assumptions | Consider a different model (e.g., non‑linear, transformations) |
| Unusual but valid value (by chance) | Retain the observation — don’t remove arbitrarily |
Influential Observations
An influential observation is one whose removal would substantially change the estimated regression coefficients (slopes and intercept).
Sources of influence:
- Outlier in the direction (large residual)
- Extreme value of an independent variable (far from its mean)
- A combination of both
Detection: Advanced measures (e.g., Cook’s distance, leverage) exist but are not covered in this module. For now, rely on careful examination of any observation flagged as an outlier or extreme.
Action:
- Check for data errors → correct if found.
- If the observation is valid, it may provide valuable insight into the relationship at the extremes.
- Do not remove valid influential observations unless there is clear evidence they do not belong to the population being studied.
- If possible, collect additional data at intermediate values to better understand the relationship.
Exam tip: “Influential” is not the same as “outlier.” An outlier may have little influence if is near the mean; an extreme can be influential even if its residual is small.
Key Takeaways – Outliers & Influential Observations
- Outliers: is a common threshold.
- Influential observations can change regression results dramatically.
- Always verify data integrity first.
- Retain valid observations — do not delete merely to improve fit.
- Advanced detection methods (Cook’s distance) exist but are beyond this module.
Recap: MLR Fundamentals (as applied to Hanumantha’s Dataset)
Model and Estimation
Multiple linear regression model:
Estimated regression equation (using OLS, minimising SSE ):
Hanumantha example ( = monthly spend, = annual income (lakhs), = household size):
Goodness of Fit
- Coefficient of determination , where .
- Adjusted penalises for extra predictors; in this example , Adjusted .
- Standard error of estimate .
Interpretation: The model with both predictors explains about 80% of the variation in monthly spend — an improvement over simple regression with only income ().
Significance Tests
-
F‑test for overall significance Test statistic . Hanumantha: , very small p‑value → reject ; at least one predictor is significant.
-
t‑tests for individual coefficients , test statistic , df . Both p‑values near zero → both and have statistically significant linear relationships with .
Assumptions and Diagnostics
- Assumptions about : mean 0, constant variance (homoscedasticity), normal distribution, independence.
- Multicollinearity: high correlation among ’s makes interpretation difficult. In Hanumantha’s data, no problem with and ; but adding a third variable (monthly income) illustrated multicollinearity in action.
- Residual analysis: plots showed no violations; standardized residuals mostly inside .
Key Takeaways – Recap
- MLR uses OLS to estimate slope coefficients.
- and adjusted measure explanatory power; ‑test checks overall significance; ‑tests check individual predictors.
- Always verify residual assumptions (plots, standardized residuals).
- Multicollinearity can distort signs and p‑values; examine correlation among predictors.
- Hanumantha’s example: clean model with good fit, significant predictors, and valid residuals.
Applications with Examples
Multiple linear regression (MLR) extends simple linear regression by using two or more independent variables to predict a dependent variable. The goal is to capture more of the variation in by including additional predictors.
Model form:
The estimated regression equation is:
Interpretation of coefficients:
- : expected change in for a one-unit increase in , holding constant.
- : expected change in for a one-unit increase in , holding constant.
Two key hypothesis tests are used:
- Overall F-test – tests whether at least one coefficient is non-zero ().
- Individual t-tests – tests whether a specific coefficient is zero ().
Exam tip: The phrase “holding constant” is the core of MLR interpretation. Always include it when explaining a slope coefficient.
Example 1: Predicting Household Debt from First Income and Monthly Payment
Context: Manjula Nayak’s Bengaluru household survey (n ≈ 500). Dependent variable: debt (₹). Independent variables: = first income, = monthly payment.
Regression output:
| Measure | Value |
|---|---|
| 0.45 | |
| Adjusted | 0.448 |
| Standard error | 25644 |
| 13081 | |
| (income) | 0.025 |
| (monthly payment) | 2.09 |
| 0.003 | |
| 0.188 | |
| -stat for | 8.74 (p ≈ 0) |
| -stat for | 11.1 (p ≈ 0) |
| F-statistic | 203 (p ≈ 0) |
Interpretation:
- = 0.45: The model explains 45% of the variation in debt. This is higher than the simple regression for income alone (0.31) or monthly payment alone (0.37).
- Adjusted ≈ because of large sample size (n=500).
- Standard error = 25644 – smaller than the simple regression standard errors (~28,619 for income; ~27,518 for monthly payment).
- Coefficient = 0.025: For every ₹1 increase in first income, holding monthly payment constant, debt increases by ₹0.025.
- Coefficient = 2.09: For every ₹1 increase in monthly payment, holding income constant, debt increases by ₹2.09.
- Both t-tests reject (p ≈ 0): both predictors have a significant linear relationship with debt.
- F-test rejects (p ≈ 0): at least one predictor is significant; the overall model is useful.
Residual analysis:
- Standardized residuals: most within ±2, a few outside.
- Residual vs. and residual vs. plots show an oval pattern – not a perfect random band, but no funnel shape or curvature (acceptable).
- Residual vs. fitted plot similar – no major violations.
Key takeaways – Model 1
- Two predictors explain more variation than either alone ( increased).
- Both income and monthly payment are significant individual predictors.
- Model still leaves 55% of variation unexplained – debt depends on other factors.
- Residual plots are adequate; no strong evidence of heteroscedasticity or nonlinearity.
Example 2: Predicting Household Debt from Monthly Payment and Utilities
Context: Same dataset, different predictors: = monthly payment, = utilities.
Regression output:
| Measure | Value |
|---|---|
| 0.67 | |
| Adjusted | ≈0.67 |
| Standard error | 19839 |
| -140771 | |
| (monthly payment) | 1.44 |
| (utilities) | 168.91 |
| 0.143 | |
| 7.87 | |
| -stat for | 10.02 (p ≈ 0) |
| -stat for | 21.47 (p ≈ 0) |
| F-statistic | 506.9 (p ≈ 0) |
Interpretation:
- = 0.67: Model explains 67% of debt variation – a clear improvement over Model 1 (0.45) and over the simple regressions (0.31 for income, 0.37 for monthly payment).
- Standard error = 19839 – lower than Model 1 (25644), indicating more precise predictions.
- Coefficient = 1.44: For every ₹1 increase in monthly payment, holding utilities constant, debt increases by ₹1.44.
- Coefficient = 168.91: For every ₹1 increase in utilities, holding monthly payment constant, debt increases by ₹168.91.
- Both t-tests significant (p ≈ 0).
- F-statistic = 506.9 (p ≈ 0) – overall model highly significant.
Comparison of SSR and SSE:
- increased from (Model 1) to (Model 2).
- decreased from to .
- remains the same () because total variation in debt is fixed.
Exam tip: is constant across models for the same ; changes in are driven entirely by (or equivalently ).
Residual analysis:
- Standardized residuals: mostly within ±2, a few outliers.
- Residual vs. , vs. , vs. fitted values all show random scatter – no patterns (oval, funnel, or curvature).
- This model appears more reliable than Model 1 in terms of residual diagnostics.
Important caveat: Regression establishes association, not causality. A higher utility bill does not necessarily cause more debt; it may reflect a larger house or higher usage, which correlates with debt.
Key takeaways – Model 2
- Model with monthly payment and utilities performs better (, lower ) than the model with income and monthly payment.
- Both predictors are highly significant individually and jointly.
- Residual plots show random scatter – no serious violations of linear regression assumptions.
- Better does not mean causality; always interpret associations cautiously.
- The choice of independent variables strongly affects model performance.
Three-Variable Multiple Linear Regression (Household Debt Data)
After observing that two different two-variable models (monthly payment + utilities; first income + monthly payment) each yielded significant relationships, a model combining all three predictors is tested: debt as a function of first income , monthly payment , and utilities .
The estimated regression equation from Excel (Analysis ToolPak) is:
Model Fit Statistics
| Statistic | Value | Comparison to Best Two-Variable Model () |
|---|---|---|
| 0.708 | Higher (0.67 previously) | |
| Adjusted | ~0.70 (slightly less than ) | Close to |
| Standard Error | 18,699 | Lower (19,839 previously) |
| Higher () | ||
| Lower () | ||
| (unchanged) | — |
The three-variable model explains 70% of the variation in household debt, the best fit so far.
ANOVA and Overall Significance
- Degrees of freedom: regression = 3, error = 496.
- , .
- , .
Conclusion: Reject . There is significant overall linear relationship.
Coefficient Estimates and Individual t-Tests
| Variable | Coefficient | Std. Error | t-statistic | p-value |
|---|---|---|---|---|
| Intercept () | -140,378 | — | — | — |
| First income () | 0.017 | 0.002 | 7.96 | ~0 |
| Monthly payment () | 0.96 | 0.148 | 6.5 | ~0 |
| Utilities () | 157.95 | 7.54 | 20.95 | ~0 |
All three coefficients are statistically significant (p ≈ 0). Interpretation (ceteris paribus):
- For every ₹1 increase in first income, debt increases by ₹0.017.
- For every ₹1 increase in monthly payment, debt increases by ₹0.96.
- For every ₹1 increase in utilities, debt increases by ₹157.95.
Exam tip: In multiple regression, the interpretation of a coefficient is "holding all other predictors constant." This is the key distinction from simple regression.
Residual Analysis
- Standardized residuals mostly within ±3 range; a few points exceed ±3 (possible outliers) — more visible because the model fits better and the standard error is lower.
- Plots of residuals against , , , and all show random scatter above and below zero → no obvious violation of linearity or homoscedasticity.
Overall, the three-variable model is the strongest so far for this dataset.
Simple Linear Regression on Real Estate Data (Property Prices)
Dataset: 30 flats in Jayalakshmi Puram, Mysuru. Variables include selling price (in lakhs), area in sq. ft. , number of bedrooms/bathrooms, premium locality, gated complex, amenities. Only area is used in this simple linear regression (SLR).
Scatter Plot and Trend Line
Rough positive linear trend with several potential outliers. Trend line equation:
— only 44% of variation in selling price explained by area.
Regression Output (ToolPak)
| Statistic | Value |
|---|---|
| 0.44 | |
| Adjusted | 0.43 |
| Standard Error () | 70.57 lakhs |
| 111,539 | |
| 139,462 | |
| 251,001 | |
| 111,539 | |
| 4,981 | |
| (1,28) | 22.39, |
Coefficients and t-Test
| Variable | Coefficient | Std. Error | t-statistic | p-value |
|---|---|---|---|---|
| Intercept () | 28.37 | — | — | — |
| Area () | 0.096 | 0.02 | 4.73 | ~0 |
Interpretation: For every 1 sq. ft. increase in area, selling price increases by 0.096 lakhs (₹9,600). The relationship is statistically significant.
Exam tip: The large coefficient magnitude (₹9,600 per sq. ft.) results from having only one predictor. Adding other relevant variables will likely change this coefficient (and reduce omitted variable bias).
Residual Analysis
- Standardized residuals: a few points beyond ±3 (possible outliers).
- Plots of residuals vs. and vs. show random scatter → no major pattern.
This SLR is a starting point. A multiple linear regression with additional predictors (bedrooms, bathrooms, locality, etc.) could improve the model.
Key Takeaways
- Adding more relevant predictors can increase and lower standard error, as seen in the debt model (three variables outperformed two).
- In multiple regression, each coefficient is interpreted "holding all other variables constant."
- The -test checks overall significance; t-tests check individual significance.
- Residual plots should be random; points beyond ±3 standardized residuals may be outliers.
- Simple linear regression can underfit; multiple regression often provides better explanation and more reliable coefficient estimates.
Exam tip: When comparing models, prefer adjusted over plain because it penalizes adding unnecessary variables. Also, always check residual plots for violations.
MLR Application – Real Estate Example
We aim to improve the simple linear regression model that predicts flat selling price (₹ lakhs) using only area (sq ft) by adding other quantitative variables: number of bedrooms and number of bathrooms. The dataset contains 30 flats from Jayalakshmi Puram, Mysore. The baseline model:
: only 44% of variation in price is explained by area alone. Both the overall F‑test and the t‑test confirmed a significant linear relationship.
Baseline Simple Linear Regression Recap
| Metric | Value |
|---|---|
| 0.44 | |
| Standard error | 70.57 |
| SSR | 111,539 |
| SSE | 139,461 |
| SST | 251,001 |
Model 1: Price = f(Area, Bedrooms)
Estimated Equation
Regression Summary
| Metric | Value |
|---|---|
| 0.46 | |
| Adjusted | 0.42 |
| Standard error | 70.74 |
| SSR | 115,892 |
| SSE | 135,109 |
| SST | 251,001 |
| MSR | 57,946 |
| MSE | 5,004 |
| F-statistic | 11.58 |
| p-value (F) | 0.0002 |
Coefficient Table
| Predictor | Coef. | Std. Error | t-stat | p-value |
|---|---|---|---|---|
| Intercept | −29.19 | — | — | — |
| Area (sq ft) | 0.085 | 0.024 | 3.46 | 0.002 |
| Bedrooms | 23.46 | 25.15 | 0.932 | 0.36 |
Interpretation
- Overall F‑test: is rejected () → at least one predictor is useful.
- t‑test for Bedrooms: → cannot reject . Adding bedrooms does not significantly improve the model once area is already included.
- Residual plots showed no alarming patterns (two standardized residuals >2, otherwise fine).
Why does bedrooms become insignificant?
The correlation between area and bedrooms is 0.542. Although this is below the common rule‑of‑thumb threshold of 0.7, multicollinearity still inflates the standard error of the bedrooms coefficient, making its t‑test non‑significant. The diagnostic tells us: given area, bedrooms add little unique information.
Model 2: Price = f(Area, Bathrooms)
Estimated Equation
Regression Summary
| Metric | Value |
|---|---|
| 0.45 | |
| Adjusted | 0.41 |
| Standard error | 71.18 |
| SSR | 114,185 |
| SSE | 136,816 |
| SST | 251,001 |
| MSR | 57,092 |
| MSE | 5,067 |
| F-statistic | 11.26 |
| p-value (F) | 0.0003 |
Coefficient Table
| Predictor | Coef. | Std. Error | t-stat | p-value |
|---|---|---|---|---|
| Intercept | 10.61 | — | — | — |
| Area (sq ft) | 0.084 | — | — | 0.005 |
| Bathrooms | 17.80 | — | 0.723 | 0.48 |
Use the reported t-statistics and p-values; the standard errors are not required for this comparison.
Observations
- (0.45) is barely different from the simple model (0.44) and slightly worse than Model 1 (0.46).
- Again the t‑test for bathrooms is non‑significant (), while the overall F‑test remains significant.
- Correlation between area and bathrooms is 0.659 – higher than before, confirming multicollinearity.
Diagnosing Multicollinearity
Multicollinearity occurs when predictor variables are highly correlated, making it difficult to isolate their individual effects on the response.
| Pair | Correlation |
|---|---|
| Area ↔ Bedrooms | 0.542 |
| Area ↔ Bathrooms | 0.659 |
Exam tip: A non‑significant t‑test for an individual coefficient in a multiple regression does not always mean the variable is irrelevant. When predictors are correlated, the shared information can cause inflated standard errors. The overall F‑test may still be significant, indicating the model as a whole is useful.
Key Takeaways from the two models:
- Adding bedrooms or bathrooms to a model that already includes area yields only marginal improvement in (0.44 → 0.46 or 0.45).
- The F‑test for overall significance is significant in both models, confirming that at least one predictor (area) is important.
- The individual t‑tests for bedrooms and bathrooms are non‑significant because of multicollinearity with area.
- Multicollinearity can exist even with correlations below 0.7; diagnostics (t‑tests, VIF) are needed.
- With only these quantitative variables exhausted, the next step is to incorporate qualitative variables (premium locality, gated community, amenities) – covered in the next module.
What did we learn?
- Adding more variables does not guarantee a better model; multicollinearity can limit the gain.
- Use the combination of overall F‑test and individual t‑tests together to guide variable selection.
- When a variable is redundant given others, remove it to keep the model parsimonious.
Intercept () – Why It’s Often Ignored
Most regression output includes a test for : However, this hypothesis is rarely of practical interest.
- Inferences on are special cases of inferences on the fitted value of the regression line as a whole – concepts like confidence intervals for predicted and estimated values of .
- The -value for in standard computer output is usually not important and is often ignored.
Extrapolation risk: If the training data do not include cases where all , then using the equation to predict at is extrapolation – outside the range of observed data. Such predictions are unreliable.
Exam tip: Never interpret as a meaningful baseline unless the data actually contain observations near . The intercept is often a mathematical necessity, not a substantive parameter.
Omitted Variable Bias (OVB)
OVB occurs when a regression model excludes one or more relevant independent variables. The included variables then absorb some of the effect of the omitted variables, leading to biased coefficients (over‑ or underestimation of true contributions).
Conditions for OVB
For omitted variable bias to exist, both of the following must hold:
- The omitted variable is correlated with the dependent variable .
- The omitted variable is correlated with at least one included independent variable.
Example (household debt data):
- = debt, = income, = monthly payments, = utilities.
- Simple linear regression (debt ~ income only) yields .
- Correlations (from sample):
| Variable | Correlation with Debt |
|---|---|
| Income () | 0.56 |
| Monthly payments () | 0.61 |
| Utilities () | 0.78 |
- Correlations among variables:
- Monthly payments & income: 0.55
- Utilities & income: 0.39
- Utilities & monthly payments: 0.49
Both conditions hold → OVB present.
How to detect OVB
- increases significantly when omitted variables are added. (Example: 0.31 → 0.71)
- Standard error drops substantially.
- -values of included variables drop (become more significant) after adding omitted ones.
Quantifying OVB
Coefficient bias can be estimated (details beyond this module), but the key diagnostic is the improvement in model fit when the omitted variables are included.
Key takeaways
- The intercept test () is usually unimportant; avoid extrapolating to outside the data range.
- Omitted variable bias arises when an excluded variable correlates with both and included s.
- Detect OVB by adding omitted variables and observing a large jump in , a drop in standard errors, and improved -values.
- OVB causes included coefficients to be biased – the included variable “soaks up” effects that belong to the omitted one.
Multiple Linear Regression – Module Recap
The module covered the full workflow for building and validating an MLR model.
Model form
Estimated regression equation
Key steps covered in this module:
- Least squares estimation – finding to minimize .
- Multiple coefficient of determination () – measure of overall fit.
- Assumptions on the error term – zero mean, constant variance, independence, normality.
- F-test for overall significance of the entire regression.
- t-tests for individual parameters (for ).
- Multicollinearity – its detection and consequences.
- Omitted variable bias – conditions, detection, impact.
- Residual analysis – identifying outliers and influential observations.
- Iterative model building – comparing and improving models using tools like Analysis ToolPak.
The next module will extend MLR to include categorical variables, followed by a course summary.
Key takeaways
- MLR models a linear relationship between and multiple s.
- The estimated coefficients are sample statistics estimating population parameters .
- Model validation uses , F-test, t-tests, and residual analysis.
- Omitted variable bias and multicollinearity are critical pitfalls to check during model building.