Review of Simple Linear Regression
Simple linear regression models the relationship between one dependent variable (outcome) and one independent variable (predictor). The goal is to find a straight line that best describes how changes in the predictor affect the outcome – a best‑fit line.
where is the predicted dependent variable, the independent variable, the intercept, and the slope (the estimated change in per unit change in ).
Intuition: Quantify and predict a one‑factor influence. Business examples:
- Predict sales based on advertising spend.
- Forecast monthly revenue from foot traffic.
- Estimate fuel consumption from distance driven.
Multiple Linear Regression
Multiple linear regression extends the idea to multiple independent variables simultaneously:
Why it matters: Real‑world outcomes are rarely driven by a single factor. Multiple regression disentangles the joint effect of several predictors.
Business examples:
- Real estate: House price depends on location, bedrooms, amenities.
- Finance: Credit score assessment using income, employment status, debt level.
- Sales prediction: Price, economic conditions, number of competitors all matter.
Exam tip: Simple regression is a special case of multiple regression (). All assumptions and diagnostic checks apply to both.
Comparison of Simple vs. Multiple Linear Regression
| Aspect | Simple Linear Regression | Multiple Linear Regression |
|---|---|---|
| Number of predictors | One | Two or more |
| Model equation | ||
| Use case | Single‑factor influence | Multi‑factor influence |
| Business complexity | Low (e.g., ad spend → sales) | High (e.g., house pricing) |
Key takeaways – Simple & Multiple Regression
- Both provide a structured way to quantify relationships and predict outcomes.
- Simple regression handles one predictor; multiple regression handles many.
- Effective for forecasting, resource allocation, and identifying key drivers.
- Reliability depends on assumptions (next section).
Key Assumptions of Linear Regression
Validity of linear regression rests on several assumptions. When any is violated, predictions and inferences become biased.
1. Linearity
The relationship between predictors and the dependent variable must be linear – a constant change in produces a constant change in .
- Violation example: Advertising spend often shows diminishing returns; initial dollars yield steep sales increases, later dollars yield little. The true relationship is curved, not a straight line.
2. Independence of Residuals
Residuals (errors = observed – predicted) must be independent of each other.
- Violation example: Time‑series data – today’s stock price depends on yesterday’s, leading to correlated errors.
3. Homoscedasticity
The variance of residuals should be constant across all levels of the predictors.
- Violation (heteroscedasticity): As income increases, spending variability tends to rise – wealthier individuals have more diverse habits.
4. Normality of Error Terms
Residuals should follow a normal distribution (critical for confidence intervals and hypothesis tests).
- Violation: Many real‑world datasets are skewed or have outliers. For continuous non‑normal data, consider transformation; for binary outcomes, switch to logistic regression.
5. No Multicollinearity
Independent variables should not be highly correlated with each other.
- Violation example: Advertising spend and marketing spend are often strongly correlated, making it impossible to isolate each variable’s unique effect on sales.
| Assumption | Description | Business violation example |
|---|---|---|
| Linearity | Predictor–outcome relationship is a straight line | Advertising with diminishing returns |
| Independence | Residuals not correlated over time | Stock prices; time‑series data |
| Homoscedasticity | Constant residual variance | Income vs. spending; richer people show more variability |
| Normality | Residuals are normally distributed | Skewed data or binary outcomes |
| No multicollinearity | Predictors not highly correlated | Advertising spend & marketing spend |
Key takeaways – Assumptions
- Violations lead to biased estimates, invalid hypothesis tests, and poor predictions.
- Each assumption points to a specific advanced technique (e.g., logistic regression for binary outcomes; nonlinear regression for curvature).
- Always test assumptions before relying on linear regression results.
When Assumptions Fail: Introduction to Advanced Methods
When linear regression’s assumptions are not met, two common alternatives are introduced:
Logistic Regression
Used when the dependent variable is categorical – often binary (yes/no, success/failure).
- Linear regression is inappropriate because it assumes a continuous outcome and can predict probabilities outside [0,1].
- Logistic regression models the probability of an event using a logistic (S‑shaped) function, ensuring predictions fall between 0 and 1.
- Business applications: predict customer purchase (yes/no), loan default, employee retention.
Non‑linear Regression
Used when the relationship between variables is inherently non‑linear.
- Examples: price–demand curves (small price changes cause large demand shifts at certain points), drug dosage–response (diminishing effects at high doses).
- Common forms: polynomial, exponential, logarithmic models.
Exam tip: If the dependent variable is binary, always choose logistic regression over linear regression – even if other assumptions hold. The linear model will produce nonsense probabilities.
Key takeaways – Advanced Methods
- Logistic regression handles binary outcomes; models probability via a logistic curve.
- Non‑linear regression captures curvilinear relationships that linear models cannot fit.
- These methods preserve the interpretability and predictive power that linear regression offers, but under more realistic conditions.
Linear Regression for Two-Sample Comparison of Means
Intuition: Comparing the means of two independent groups (e.g., before/after a campaign, with/without an accident) is typically done with a two-sample t-test. The same test can be performed using linear regression with a binary dummy variable that encodes group membership. This may seem roundabout, but it pays off: once the comparison is embedded in a regression framework, we can easily add control variables, extend to multiple groups, and link to more advanced models like fixed effects and interaction effects.
Setup
- Dependent variable : continuous outcome (e.g., satisfaction level).
- Independent variable : binary (0 or 1) indicating group membership.
- : reference group (e.g., no accident at work).
- : comparison group (e.g., had an accident).
Model:
Interpretation of Coefficients
| value | Expected | Interpretation |
|---|---|---|
| Mean of reference group | ||
| Mean of comparison group | ||
| Difference | Mean difference (comparison group minus reference group) |
Thus:
- Testing whether is significantly different from zero is equivalent to testing whether the two group means differ.
- The sign of indicates which group has the higher mean.
Worked Example: Accidents and Job Satisfaction
Question: Do employees who experienced an accident at work have a different average satisfaction level than those who did not?
Data: HR dataset containing satisfaction_level (continuous) and any_accident (1 = yes, 0 = no).
Simple linear regression output (from Excel):
| Coefficient | Estimate | Std. Error | p-value |
|---|---|---|---|
| Intercept () | 0.607 | ... | <0.05 |
any_accident () | 0.041 | 0.006 | <0.05 |
- : mean satisfaction of employees without an accident.
- : employees with an accident have a mean satisfaction that is 0.041 (≈4%) higher.
- The coefficient is statistically significant → the mean difference is real.
Exam tip: A regression with a single binary predictor produces exactly the same p‑value as an independent-samples t-test for the mean difference. The regression approach becomes superior when you need to control for additional variables.
Why This Result Might Be Surprising
The positive sign (accident → higher satisfaction) is counter‑intuitive. A possible explanation is confounding: employees who have accidents may also work in different departments, have different workloads, or be treated differently by the company. The simple regression cannot separate the effect of the accident from these other factors.
Controlling for Confounders: Multiple Regression
Add covariates that capture workload and experience:
number_projectsandaverage_monthly_hours(workload)years_at_companyandpromotion_last_5years(experience)
Model:
Result: The coefficient for any_accident remains about 0.041 and statistically significant. Even after controlling for workload and experience, the positive relationship persists. This strengthens the evidence that the accident effect is genuine (perhaps due to company support policies), though further investigation is warranted.
Key advantage over a t-test: The regression framework naturally controls for multiple covariates, reducing omitted‑variable bias and yielding a more reliable estimate of the causal effect (under appropriate assumptions).
Connections to Advanced Topics
- Fixed effects: Using dummy variables for each individual (or group) to control for all time‑invariant unobservable characteristics. The regression shown is a simple fixed‑effects model at the group level (accident vs. no accident). In panel data, each subject gets its own dummy, isolating within‑subject variation.
- Interaction effects: The effect of the binary variable may differ across levels of another variable (e.g., salary). For example, an accident might lower satisfaction for low‑paid employees but not for high‑paid ones. This is modelled by adding an interaction term (e.g.,
accident × salary) to the regression.
Key Takeaways
- A simple linear regression with a binary dummy variable performs a two‑sample comparison of means.
- = mean of reference group; = mean difference (comparison minus reference).
- A significant indicates a statistically significant difference between groups.
- The sign of reveals which group has the higher mean.
- Multiple regression allows controlling for confounders, a major advantage over a standard t-test.
- Dummy variable regression is the foundation for fixed effects models and interaction effects.
From Two Samples to Multiple Groups
A linear regression model can compare means across more than two independent groups, just as it does for two samples. Intuitively: instead of running a one-way ANOVA, we regress the outcome on a set of binary (dummy) variables that encode group membership, and test whether the group differences are statistically significant.
This approach treats the outcome as a continuous variable (unlike a chi‑square test of independence, which would require binning the continuous outcome into categories). The regression model directly estimates group means and their differences.
Dummy Variable Setup for Groups
- Create binary columns, one per group (e.g.,
Low,Medium,High). - Choose one group as the baseline (reference) category – omit its dummy from the model.
- The intercept estimates the mean of the baseline group.
- Each other coefficient estimates the difference between the mean of group and the baseline mean.
Example: Salary Level and Satisfaction
A dataset of 14,999 employees has salary categorised as Low, Medium, or High. The satisfaction level (0–1) is the dependent variable. To test whether mean satisfaction differs by salary group:
- Create dummy variables:
Low(=1 if salary=Low),Medium,High. - Set baseline: e.g.,
Lowis the reference group – omit it from the model. - Run regression: satisfaction ~
Medium+High(plus intercept).
| Variable | Coefficient | Interpretation |
|---|---|---|
| Intercept | Mean satisfaction of low‑salary employees ≈ 0.600 (i.e., 60% satisfied). | |
Medium | (significant) | Medium‑salary employees are on average 2.1% more satisfied than low‑salary. |
High | (significant) | High‑salary employees are on average 3.7% more satisfied than low‑salary. |
Thus:
- Low salary mean = 60.0%
- Medium salary mean = 60.0% + 2.1% = 62.1%
- High salary mean = 60.0% + 3.7% = 63.7%
A coefficient that is significantly different from zero (via its ‑statistic / ‑value) indicates that the group mean differs from the baseline. The sign tells the direction.
Exam tip: Always set the baseline to a meaningful group (e.g., the most common or a control). The interpretation of the intercept and all other coefficients depends on that choice.
Key Takeaways
- Dummy variables convert categorical membership into numeric predictors.
- One category is always omitted – it becomes the baseline (intercept).
- Coefficients on the included dummies represent mean differences relative to baseline.
- A significant coefficient implies the group’s mean is statistically different from baseline.
- The same logic extends the two‑sample ‑test to multiple groups (equivalent to one‑way ANOVA).
Interaction Effects: Does Salary’s Impact Vary by Department?
Question: Do employees from different departments place equal emphasis on salary when reporting satisfaction? Test this with an interaction term, which captures the joint effect of the predictors.
Intuition
A simple additive model assumes the salary effect is the same across departments. An interaction model allows the slope (difference between salary groups) to differ by department. For example, sales employees might value salary more than R&D employees, for whom job security matters more.
How to Build an Interaction Model (in Excel)
- Create dummy variables for every combination of the two categorical variables.
- Example: 10 departments × 3 salary levels = 30 dummy columns (e.g.,
Sales_Low,Sales_Medium,Sales_High,HR_Low, …).
- Example: 10 departments × 3 salary levels = 30 dummy columns (e.g.,
- Choose a baseline combination – e.g.,
HR_Low– and omit it from the regression. - Run the regression with all remaining 29 dummy variables as predictors.
Practical Concern – Sparse Cells
If a combination has very few observations (e.g., only 2 management employees with low salary), its coefficient will be unreliable. A recommended workaround: drop dummy columns for those sparse interactions. You will still capture the overall effect for that department but sacrifice the ability to estimate the interaction precisely.
Interpreting the Output
- Each interaction coefficient measures the difference in mean satisfaction for that specific (department, salary) combination relative to the baseline combination.
- Joint significance test (e.g., ‑test on all interaction dummies) tells whether including them significantly improves the model.
- If the interaction effects are significant: salary’s influence on satisfaction depends on department.
- If not: salary has a similar impact across all departments – the additive model suffices.
Exam tip: Interaction terms answer the question “Does the effect of X on Y depend on Z?” When Z is categorical, you need dummies for every X‑by‑Z cell. Always check for sparse cells; drop them to maintain reliable estimates.
Key Takeaways
- Interaction effects capture how the relationship between the dependent variable and one categorical predictor changes across levels of another categorical predictor.
- Implementation: create dummy variables for every combination of the two categories, then regress on those dummies (minus one baseline).
- Sparse cells (few observations) produce unreliable coefficients – consider dropping those dummy columns.
- Significant interactions imply the effect of salary differs by department; non‑significant interactions imply uniform effect.
Logistic Regression Model - I
Logistic regression is a technique for modeling the probability of a binary outcome (e.g., yes/no, success/failure, leave/stay) as a function of one or more independent variables. It is the natural choice when the dependent variable is categorical – unlike linear regression, which assumes a continuous response.
Why linear regression fails for binary outcomes
- A binary outcome (coded or ) follows a Bernoulli distribution.
- Linear regression would model . Because a line is unbounded, predicted probabilities can fall outside – nonsensical for a probability.
- Logistic regression fixes this by applying the logistic (sigmoid) function, which maps any real-valued input to .
The logistic (sigmoid) function
- The curve is S‑shaped: for very low or very high , the probability flattens near or .
- The sign of determines the direction: positive → increasing probability as rises; negative → decreasing.
Foundation: Bernoulli distribution
- A binary variable follows a Bernoulli distribution: , .
- Logistic regression estimates as a function of the independent variables, ensuring .
Odds and log odds
Odds of success:
- If , odds (equal chance).
- If , odds ; if , odds .
Log odds (logit):
- Spans the entire real line: when , when .
- Logistic regression models log odds as a linear function of the independent variables:
Exam tip: Linear regression models the mean directly; logistic regression models the log odds linearly. This linearity in log odds is what makes coefficients interpretable as changes in log odds per unit predictor increase.
Worked example: Instagram users by gender
Data: 1069 survey respondents (537 men, 532 women). 328 women and 234 men have Instagram accounts.
| Gender | Users | Total | Proportion | Odds | Log odds |
|---|---|---|---|---|---|
| Women | 328 | 532 | |||
| Men | 234 | 537 |
Define for women, for men. Logistic regression assumes:
Plugging in the log odds:
- For men ():
- For women ():
Estimated model:
Convert back to probability:
- For women (): (matches data).
- For men (): .
Exam tip: The coefficient is the difference in log odds between women and men. A positive means higher log odds (and thus higher probability) for the group coded .
Business applications
- Marketing: Predict whether a customer buys a product (based on demographics, browsing data).
- Credit scoring: Classify loan applicants as likely to default or repay.
- HR: Model employee attrition (e.g., probability of leaving given satisfaction level).
In all cases, logistic regression outputs a probability between 0 and 1, which can be used for classification with a chosen threshold (e.g., >0.5 → “will leave”).
Key takeaways
- Logistic regression models binary outcomes using the logistic (sigmoid) function to keep predicted probabilities in .
- It is built on the Bernoulli distribution.
- Instead of predicting directly, it models the log odds linearly: .
- The odds are ; log odds are the natural log of that ratio.
- Coefficients represent the change in log odds for a one‑unit increase in .
- Linear regression fails because it can produce probabilities outside .
- The sigmoid curve ensures valid probabilities: S‑shaped, flat near 0 and 1.
Logistic Regression Model - II
Logistic regression predicts the probability of a binary outcome (success/failure, quit/stay, buy/not buy). Unlike linear regression, it models the log odds of the event as a linear function of predictors, ensuring predictions stay between 0 and 1.
From linear combination to log odds
For multiple predictors , the logistic regression equation expresses the log odds:
- = probability the event occurs (e.g., employee leaves)
- = intercept (log odds when all = 0)
- = coefficients measuring the effect of each predictor on log odds
Using log odds solves the range problem: probability is bounded , but log odds can take any real value , so the linear model can fit without constraints.
The logistic (sigmoid) function
Rearranging the equation gives probability directly:
This is the logistic function (or sigmoid). It produces an S‑shaped curve that smoothly increases from 0 to 1 as the linear combination changes.
Definition: The logistic function maps any real input to a probability between 0 and 1 — essential for binary classification.
Interpreting coefficients
- A positive coefficient means an increase in increases the log odds (and thus the probability) of the outcome.
- A negative coefficient means an increase in decreases the log odds.
- Because the relationship between and is non‑linear, the change in probability per unit change in is not constant — it depends on the current value of . However, the direction (positive/negative) is fixed.
Exam tip: Logistic regression coefficients refer to log odds, not probability. To get the effect on odds, exponentiate the coefficient: gives the odds ratio.
Estimating coefficients: Maximum Likelihood Estimation (MLE)
Linear regression uses least squares; logistic regression uses maximum likelihood estimation (MLE).
Intuition: The model tries many possible coefficient values. For each set, it computes the predicted probability for every observation. Then it compares these probabilities to the actual binary outcomes (0 or 1). The goal is to find the coefficient values that make the observed data most likely — i.e., that maximize the likelihood of seeing the actual outcomes.
Procedure:
- Start with initial guesses for s.
- Compute predicted probabilities via the logistic function.
- Calculate the likelihood (a measure of how well predictions match real outcomes).
- Adjust coefficients iteratively to increase the likelihood.
- Stop when improvement becomes negligible — the maximum likelihood estimates.
This iterative search can be performed in Excel using the Solver add‑in, by setting up the likelihood function and maximizing it.
Implementation in Excel
For a dataset with one binary outcome and predictors:
- Use the logistic function to compute predicted probabilities for each row.
- Construct the log‑likelihood formula (sum of log probabilities for observed outcomes).
- Run Solver to maximize the log‑likelihood by changing the coefficient cells.
Exam tip: Solver finds the same coefficients that statistical software (R, Python, SPSS) produces — it's a good way to see MLE in action without coding.
Key takeaways
- Logistic regression models the log odds of a binary outcome.
- The sigmoid function turns the linear combination into a probability .
- Coefficients indicate direction (positive/negative) on log odds; the effect on probability is non‑linear.
- Maximum likelihood estimation iteratively finds the best coefficients by maximizing the probability of observing the data.
- In Excel, Solver can be used to perform MLE for small problems.
Intuition and Problem Setup
Logistic regression models a binary outcome (e.g., leave / stay) as a function of one or more predictors. Here the goal is to predict employee attrition (1 = left, 0 = stayed) from satisfaction level (0 to 1). The relationship is non‑linear: probability of leaving is linked to the predictors via the logistic function:
Equivalently, the log‑odds (logit) of the event is linear:
The coefficients are estimated by maximum likelihood – an iterative search that finds the values making the observed data most probable.
Step‑by‑Step Estimation in Excel
The data (columns A, B) contain the binary outcome (A) and satisfaction (B). The Excel implementation proceeds as follows.
1. Compute Log‑Odds
Place initial guesses for and in cells (e.g., H2, I2). For each row :
In Excel: = $H$2 + $I$2 * B2 (copied down).
2. Convert Log‑Odds → Odds → Probability
Then the predicted probability depends on the actual outcome:
- If :
- If :
Excel: = IF(A2=1, D2/(1+D2), 1/(1+D2)).
3. Log‑Likelihood
Because probabilities multiply across independent observations (product → very small), we work with logs:
Sum all to get the total log‑likelihood (cell, e.g., J2).
4. Optimization with Solver
Goal: maximize the total log‑likelihood by changing and .
- Open Solver (Data → Solver).
- Set Objective: cell containing sum of log‑likelihood.
- To: Max.
- By Changing Variable Cells: the , cells.
- Select Solving Method: GRG Nonlinear.
- Click Solve.
Excel iterates through coefficient combinations and returns the maximum‑likelihood estimates.
Results and Interpretation
Single Variable Model
The solver yields:
| Coefficient | Estimate |
|---|---|
| (intercept) | 0.974 |
| (satisfaction) | –3.832 |
Thus the estimated model:
Interpretation:
-
Sign of is negative → higher satisfaction lowers the odds of quitting.
-
Magnitude cannot be directly read as a change in probability (non‑linear). Instead compute predicted probability at key satisfaction values:
Satisfaction Predicted probability of leaving 0 (not at all satisfied) (72 %) 1 (fully satisfied) (4–5 %) -
The decrease is non‑linear: probability drops steeply at low satisfaction and flattens at high satisfaction.
Multiple Variable Model (Extension)
The same Solver approach extends to several predictors. For a model including:
- satisfaction level
- last performance evaluation rating
- average monthly hours
- years at company
- promotion in last 5 years (coded 1/0)
| Variable | Coefficient (approx.) | Sign | Interpretation |
|---|---|---|---|
| Satisfaction | –3.7 | – | More satisfied → less likely to leave |
| Promotion in last 5 years | large negative | – | Promoted employees much less likely to leave |
| Years at company | positive (largest magnitude) | + | Longer tenure → higher chance of leaving (retirement / better offers) |
| Average monthly hours | positive | + | Higher workload → higher chance of leaving |
| Last evaluation rating | positive | + | Higher rating → more likely to leave (counterintuitive – possible reasons: felt under‑rewarded, better external opportunities) |
Exam tip: The sign of a logistic coefficient tells the direction of effect on the log‑odds of the outcome. To express effect on probability, compute predicted at different values. Never interpret a coefficient as a linear change in probability.
Key takeaways
- Logistic regression models binary outcomes via the logit link: .
- Estimation uses maximum likelihood – iterative maximization of the sum of log‑likelihoods.
- In Excel: compute log‑odds → odds → probability (conditional on outcome) → log‑likelihood → maximize with Solver (GRG Nonlinear).
- Coefficient sign indicates direction: negative → predictor reduces odds of event.
- For multiple predictors, each coefficient is interpreted holding others constant; magnitude comparison gives relative importance.
- Counterintuitive signs (e.g., positive coefficient for performance rating) can reveal hidden dynamics – use domain knowledge to hypothesise.
Logistic Regression – Using Excel with Dummy Variables
When a categorical predictor (e.g., salary level) has ( k ) categories, it enters a logistic regression as ( k-1 ) dummy variables. One category is chosen as the baseline; the coefficients on the dummies measure the change in log‑odds relative to that baseline.
Setting Up Dummy Variables
- Salary has three categories: low, medium, high.
- Set low as baseline → create two dummies:
- ( x_{\text{medium}} = 1 ) if medium salary, else 0
- ( x_{\text{high}} = 1 ) if high salary, else 0
The logistic regression then includes these dummies alongside continuous predictors.
Model Specification
Eight independent variables are used:
| Variable | Type | Description |
|---|---|---|
| satisfaction_level | continuous (0–1) | Employee satisfaction |
| last_evaluation | continuous (0–1) | Last performance rating |
| average_monthly_hours | continuous | Workload (hours/month) |
| years_at_company | integer | Tenure |
| promotion | binary (0/1) | Received promotion in last year? |
| salary_medium | dummy | 1 if medium salary |
| salary_high | dummy | 1 if high salary |
| (plus intercept) | — | ( \beta_0 ) |
Parameters: ( \beta_0, \beta_1, \dots, \beta_7 ). Log‑odds for employee ( i ):
[ \text{logit}(p_i) = \ln\left(\frac{p_i}{1-p_i}\right) = \beta_0 + \beta_1 x_{i1} + \cdots + \beta_7 x_{i7} ]
The model is fitted by maximizing the sum of log‑likelihoods (using Excel Solver).
Interpretation of Coefficients
The exact values are not given, but the sign and relative magnitude are discussed.
| Variable | Sign | Interpretation |
|---|---|---|
| satisfaction_level | negative, large magnitude | Higher satisfaction → much less likely to leave |
| promotion | negative, large magnitude | Promotion → much less likely to leave |
| salary_medium | negative | vs low‑salary baseline: medium‑salary employees are less likely to leave |
| salary_high | negative, larger magnitude than medium | High‑salary employees are even less likely to leave than medium |
| years_at_company | positive, substantial magnitude | Longer tenure → more likely to leave (perhaps for better prospects) |
| last_evaluation | diminished after including salary/promotion | Becomes secondary once satisfaction and salary are controlled |
| average_monthly_hours | diminished | Similarly secondary |
Key insight: Satisfaction, salary, and promotion dominate the prediction; evaluation and hours matter less once these are accounted for.
Worked Example: What‑If Analysis
Employee profile:
| Variable | Value |
|---|---|
| satisfaction_level | 0.5 (50%) |
| last_evaluation | 0.7 |
| average_monthly_hours | 250 |
| years_at_company | 1 |
| promotion | 0 |
| salary | low → ( x_{\text{medium}}=0,; x_{\text{high}}=0 ) |
Prediction (using previously estimated coefficients):
Compute log‑odds:
[ \begin{aligned} \text{log‑odds} &= \beta_0 + \beta_1(0.5) + \beta_2(0.7) + \beta_3(250) + \beta_4(1) \ &\quad + \beta_5(0) + \beta_6(0) + \beta_7(0) \ &= -1.075 \end{aligned} ]
Convert to probability:
[ p = \frac{e^{-1.075}}{1+e^{-1.075}} \approx 0.254 \quad (25.4%) ]
Scenario Testing
Change one variable at a time and recompute ( p ) (hypothetical outcomes):
| Scenario | Change | New ( p ) | Interpretation |
|---|---|---|---|
| Raise salary to medium | ( x_{\text{medium}}=1 ) | 16.8% | Quit risk drops ◄ |
| Raise salary to high | ( x_{\text{high}}=1 ) | 4.7% | Quit risk very low ◄◄ |
| Give promotion | ( x_{\text{promotion}}=1 ) | (substantial reduction) | Promotion also strongly reduces leaving |
Business use: the HR department can estimate the impact of each intervention and choose the most cost‑effective strategy.
Extensions: Interaction Effects & Scenario Forecasting
- Interaction between salary level and department can reveal differential effects (e.g., a salary hike may retain HR staff more than management).
- This method is called scenario forecasting — the organisation tests multiple hypothetical changes and picks the one that best reduces turnover risk.
Exam tip: When building logistic models with categorical predictors, always create ( k-1 ) dummies. The baseline category (e.g., “low salary”) is absorbed into the intercept. Changing a dummy to 1 shifts the log‑odds relative to that baseline.
Key takeaways
- Categorical variables enter logistic regression as dummy variables (baseline omitted).
- Coefficients on dummies represent change in log‑odds compared to the baseline category.
- In the worked HR example: low → medium salary cut quit probability from 25.4% to 16.8%; low → high cut it to 4.7%.
- Satisfaction and promotion have large negative effects; years at company increases leaving propensity.
- Once key factors (salary, satisfaction, promotion) are controlled, variables like work hours and evaluation ratings lose predictive power.
- Use the fitted model for scenario forecasting: change one predictor at a time and compute the new probability to guide decisions.
The Curse of Dummy Variables
Fixed effects and interaction effects quickly explode the number of features. In the employee‑retention example:
- Salary – 3 categories → 2 dummy variables
- Department – 10 categories → 9 dummy variables
- Salary × Department interaction – 30 unique combinations → 29 dummy variables
Total: over 30 dummy columns. Estimation becomes unstable, convergence may fail, and interpretation suffers.
| Source | Categories | Dummy variables needed |
|---|---|---|
| Salary (fixed) | 3 | 2 |
| Department (fixed) | 10 | 9 |
| Salary × Department (interaction) | 30 | 29 |
| Total | ≥ 30 |
One partial remedy: drop interaction terms for categories with very few observations (e.g., if a department has only low‑salary employees, ignore the other salary‑department combos). But even then the model may stay too complex.
Exam tip: Whenever a categorical feature has many levels, or you include interactions, always check the total dummy count. A model with >20–30 dummy variables on a modest dataset is a red flag for overfitting and non‑convergence.
Why Variable Selection Matters
Variable selection is the process of identifying the subset of features that are most important for predicting the outcome, keeping the model simple without sacrificing predictive or explanatory power.
Three core reasons:
- Avoid overfitting – Adding more features always inflates (or pseudo‑), eventually approaching 1 on training data. But an close to 1 signals that the model fits noise, not signal – it will generalise poorly to new data.
- Interpretability – A model with 5–10 coefficients is far easier to understand than one with 50.
- Computational efficiency – Fewer variables means faster training and lower memory use.
Forward Selection
Start with an empty model (only the intercept). At each step:
- Test every candidate variable not yet in the model.
- Add the one that most improves model fit (e.g., highest increase).
- Use an F‑test (or deviance test in logistic regression) to check whether the improvement is significant.
- Repeat until no remaining variable yields a significant improvement.
Backward Elimination
Start with the full model (all variables). At each step:
- Identify the least important variable (e.g., highest ‑value, smallest drop if removed).
- Remove it.
- Repeat until removing any remaining variable significantly hurts model fit.
Stepwise Selection
A hybrid that combines forward and backward:
- Begin like forward selection: add the best variable.
- At every step, re‑evaluate all previously included variables. If any has become non‑significant (because the new variable took over its role), remove it.
- Continue until no variable can be added and none needs removal.
This handles the problem that some variables matter only in the presence of others – a variable that was significant at step 2 may become redundant after a later addition.
Ad‑hoc P‑Value Approach
A simpler, more manual method:
- Fit the full model once.
- Examine the ‑values of each coefficient.
- Remove all variables with non‑significant ‑values.
- (Optional) Refit and re‑check – significance can shift when variables are removed.
Caution: Because coefficients interact, a variable may appear non‑significant in the full model but become significant after dropping a correlated predictor. Always re‑evaluate after pruning.
Exam tip: The ad‑hoc p‑value method is fast but risky. It is acceptable only when you verify stability by fitting a few reduced models. The stepwise approach is more robust because it continuously checks for both additions and removals.
Key Takeaways
- Fixed/interaction effects with many categories generate dozens of dummy variables, causing overfitting, convergence issues, and poor interpretability.
- Variable selection finds a balance between model simplicity and accuracy.
- Common approaches: forward selection (add one by one), backward elimination (remove one by one), stepwise selection (add then re‑check removals), and ad‑hoc p‑value (prune non‑significant predictors).
- Always watch for overfitting: more features inflate , but hurt generalisation.
- In practice, use software‑built routines; but understand the logic to set sensible stopping criteria (e.g., significance level for entry/removal).
Non-Linear Regression Model
Non-linear regression models relationships where the change in the dependent variable is not proportional to the change in the independent variable. Unlike linear regression, it can capture curved, S‑shaped, or saturating patterns. This flexibility is essential for many real-world business problems where the underlying dynamics are inherently complex.
Why non‑linear? Common business applications
- Advertising & sales – Early ad spend may boost sales sharply, but additional spend eventually yields diminishing returns. Non‑linear models capture this plateau.
- Pricing & demand – Lowering price may initially spike demand, but further cuts produce only marginal gains. Complex price–response curves are better fitted with non‑linear terms.
- Growth models – Revenue or product adoption often follows an S‑shaped (sigmoid) pattern: rapid early growth, then slowdown as market saturates.
- Manufacturing & cost – Output costs may fall due to economies of scale, then rise again from inefficiencies. A non‑linear model can identify the optimal production level.
Key non‑linear regression methods
Polynomial regression
Generalizes linear regression by adding higher‑degree terms of the predictor(s). A quadratic (2nd‑degree) model with one predictor is:
Adding a cubic term () yields a cubic polynomial.
When to use – When the dependent variable rises/falls sharply at certain points and flattens elsewhere (e.g., product demand vs. price). Caution – Higher degrees increase complexity and risk overfitting; quadratic or cubic usually suffice.
Exponential regression
Models rapid acceleration (or deceleration) where the dependent variable changes at a rate proportional to its current value:
When to use – Viral campaign spread, rumor propagation, or any scenario where growth accelerates quickly.
Logarithmic regression
Models quick initial growth that then levels off – the mirror of exponential:
When to use – Situations with diminishing returns, e.g., advertising spend vs. brand awareness.
Power regression
Models a relationship where the rate of change in is a constant proportion of the change in :
When to use – Production efficiency vs. units produced (efficiency improves at a decreasing rate as scale increases).
Logistic regression (non‑linear form)
Uses the sigmoid (logistic) function to produce an S‑shaped curve:
(Here can be a continuous proportion or a binary probability.) When to use – New product adoption, market penetration, or any process that saturates.
Choosing the right method
Comparison of methods
| Method | Pattern captured | Typical equation | Example use |
|---|---|---|---|
| Polynomial (quadratic) | Curved, single bend | Price vs. demand | |
| Exponential | Rapid growth/decay | Viral campaign spread | |
| Logarithmic | Fast initial growth, then flattening | Ad spend vs. sales | |
| Power | Constant proportion rate change | Production efficiency scaling | |
| Logistic | S‑shaped saturation | Product adoption |
Exam tip – Overfitting is a real danger with polynomial regression. Stick to quadratic or cubic unless domain knowledge strongly suggests a higher degree. Cross‑validate to check generalisability.
Key takeaways
- Non‑linear regression models curved, non‑proportional relationships between variables.
- Common methods: polynomial (parabolic), exponential (rapid change), logarithmic (diminishing returns), power (constant proportion), logistic (S‑shaped).
- Choice depends on the observed pattern in the data – mis‑specification leads to poor predictions.
- Applications include advertising response, pricing optimisation, growth modelling, and manufacturing cost analysis.
- Higher‑degree polynomials increase flexibility but also the risk of overfitting; keep models as simple as the data allow.
Case Study: Bangalore House Prices
A case study integrating nonlinear regression and categorical-variable techniques to model house prices in Bangalore. The dataset contains 4340 rows and 10 variables.
Data Overview
| Variable | Type | Description |
|---|---|---|
posted_by | Categorical | Owner, dealer, or builder |
RERA_approval | Binary | 1 if approved, 0 otherwise |
bedrooms | Numeric | Number of bedrooms |
total_sqft | Numeric | Total area in square feet |
ready_to_move | Binary | 1 if ready to occupy, 0 otherwise |
resale | Binary | 1 if resale property, 0 otherwise |
address | Categorical (high-level) | General location |
longitude | Numeric | Longitude coordinate |
latitude | Numeric | Latitude coordinate |
price | Numeric (lakhs of ₹) | House price |
The response variable of interest is price, but due to its relationship with area and skewness, a transformation is needed.
Five Key Questions
The case study builds toward one main question through four preparatory ones:
- How is price distributed across the city? (exploratory mapping)
- Does price depend on specific streets/areas? (pattern identification)
- Is the price distribution normal? (normality check for regression)
- Does price differ by who posted (builder, owner, dealer)?
- What factors impact price, and how? (main modelling question)
These questions are cumulative: Q1–Q4 inform the model specification for Q5.
Exploratory Data Analysis (EDA)
Price Distribution and Normality
A histogram of price shows positive skew: high frequency at low prices, a sharp drop, then a long right tail with few observations. The distribution is clearly not normal.
Plotting price per square foot () reduces the effect of area but still shows positive skew – not symmetric.
Exam tip: In real data, variables like price, income, and rent almost always exhibit positive skew. The standard remedy is a log transformation.
Applying yields a near-symmetric histogram (some outliers on the left remain). Visually, the normality assumption becomes plausible. Formal confirmation would require a chi-square goodness-of-fit test.
Why Price per Sq Ft and Log Transformation
- Price per sq ft is the relevant metric for buyers because total price is proportional to area. A buyer compares cost per unit area.
- Log transformation is the appropriate fix for positive skew. It stabilises variance and makes the distribution more symmetric – both important for regression assumptions.
Nonlinear Regression Model
Model Specification
The regression is nonlinear because the response variable is transformed:
Independent variables (predictors):
RERA_approval(binary)bedrooms(numeric – initially treated as continuous)ready_to_move(binary)resale(binary)total_sqft(numeric)posted_by(categorical: owner, dealer, builder) – converted to dummy variables with builder as the baseline
This is a log-linear model: response is log-transformed, predictors remain untransformed.
Results and Interpretation
| Predictor | Sign | Significance (5% level) | Interpretation |
|---|---|---|---|
RERA_approval | Positive | Significant | Approval increases price per sq ft |
bedrooms | Positive | Significant | More bedrooms → higher price per sq ft (holding area fixed) |
total_sqft | Negative | Significant | Larger area → lower price per sq ft (economies of scale) |
ready_to_move | — | Not significant | No price premium for readiness |
resale | — | Not significant | Resale properties do not sell at a discount or premium |
posted_by = owner | Positive | Significant | Higher price per sq ft than builder |
posted_by = dealer | Positive (larger) | Significant | Even higher premium over builder |
Exam tip: In a log-linear model, coefficients on binary predictors represent the expected change in ; exponentiate to get the multiplicative effect on (e.g., gives the factor change in price per sq ft).
Because the dependent variable is logged, each coefficient indicates the change in for a one-unit increase in the predictor. For posted_by, both owner and dealer command a premium relative to builder, with dealer being the costliest.
Model Improvements and Further Analysis
Several ways to upgrade the model:
- Log-log model: Transform both response and
total_sqft– i.e., model as a function of . This captures proportional relationships better. - Bedrooms as categorical: The price increment per additional bedroom is likely nonlinear. Treating bedrooms as a categorical variable (1BR, 2BR, 3BR+) or applying a transformation (log, square root) may improve fit.
- Interaction effects:
resaleandposted_bymay interact – an owner selling a resale property might have different pricing than a builder selling a new one. - Location effect: Address/lat–lon data introduces many categories. Smart encoding (e.g., clustering neighbourhoods, using spatial coordinates, or creating a distance-from-center variable) is needed to avoid an explosion of dummy variables.
- Cross-validation for prediction: To compare models on predictive performance (not just ), split the data (e.g., 4000 training, 340 test). Train on one part, predict the other, and evaluate accuracy. This concept is explored further in the next module on time series and predictive modelling.
Key takeaways
- The case study demonstrates a complete workflow: EDA → transformation → nonlinear regression → interpretation → refinement.
- Log transformation is essential for right-skewed price data; log-linear models are a standard nonlinear regression tool.
- Categorical predictors require dummy variables; interpretation of coefficients changes when the response is logged.
- Model improvement avenues include log-log form, categorical treatment of numeric variables, interactions, and spatial effects.
- Cross-validation separates model fitting from evaluation – critical for predictive modelling.
Linear Regression Recap
Linear regression models the relationship between a continuous dependent variable and one or more independent variables. The goal is to quantify the strength and direction of these relationships: how changes in a predictor translate into increases or decreases in the outcome.
Because of its simplicity and interpretability, linear regression is widely used when the relationship between variables can be assumed linear. For example, a company can model the relationship between various factors (salary, workload, work environment) and employee satisfaction to identify what keeps employees happy.
Exam tip: Linear regression predicts a continuous outcome. The coefficients are directly interpretable: a one-unit increase in is associated with a change in , holding other predictors constant.
Key takeaways
- Linear regression models continuous dependent variables.
- Assumes a straight-line relationship.
- Provides clear, actionable insights in business contexts.
- Example: employee satisfaction analysis.
Logistic Regression Recap
Logistic regression is used when the dependent variable is binary (two possible outcomes). It models the probability of an event occurring, ensuring predicted values fall between 0 and 1.
Unlike linear regression, logistic regression uses the logit link function:
where is the probability of the event.
Business applications
- Employee attrition: Predict likelihood of an employee leaving based on job satisfaction, salary, workload, tenure. Enables proactive retention (promotions, salary increases, development).
- Customer segmentation & marketing: Predict whether a customer will purchase based on browsing behavior, purchase history, demographics. Helps target campaigns.
- Credit scoring & risk management: Predict probability of loan default using credit history, income, employment status. Informs approval/denial decisions.
Key takeaways
- Logistic regression for binary outcomes (0/1).
- Predicted values are probabilities bounded by 0 and 1.
- Used in HR analytics, marketing, financial risk.
Non‑Linear Regression Recap
Non‑linear regression is needed when the relationship between variables cannot be adequately captured by a straight line. Many real‑world business situations show curved patterns.
Example: House price analysis – the dependent variable (price) may not be normally distributed. Applying a logarithmic transformation (e.g., ) creates a non‑linear model that often improves performance.
Common scenarios
- Diminishing returns in marketing: Sales increase with ad spend, but after a point each additional dollar yields smaller sales increments. Non‑linear models capture this saturation effect, helping allocate budgets efficiently.
- Demand forecasting: The relationship between price and demand often follows a curve – small price changes may have large effects on demand in some ranges, little effect in others.
Exam tip: Non‑linear does not necessarily mean “complicated.” A log transformation of either the dependent or independent variable is a simple form of non‑linear regression that can linearise a curved relationship.
Key takeaways
- Non‑linear regression for curved relationships.
- Logarithmic transformations are a common tool.
- Business examples: marketing diminishing returns, price–demand curves.
Categorical predictors
Many business variables are categorical – gender, region, product type. These are incorporated into linear, logistic, or non‑linear regression using dummy variables (0/1 coding). This allows comparison of the impact of different categories on the outcome.
Interaction effects
Interaction effects occur when the relationship between one independent variable and the dependent variable changes depending on the value of another independent variable. For example:
- The effect of advertising on sales may depend on the pricing strategy.
- The impact of employee training on performance may vary by department.
Interaction terms are added to the regression model, often as the product of the involved variables (e.g., ). Including them leads to a more precise model that captures how combinations of factors influence outcomes.
Key takeaways
- Dummy variables enable categorical predictors in regression.
- Interaction effects capture dependency between predictors.
- Including interactions improves model accuracy for business decisions (e.g., attrition, house prices).