Term 1 · Module 6 of 8

Association between Random Variables and Simple Linear Regression

Business Statistics for Entrepreneurs

Covariance

Covariance quantifies the direction of a linear relationship between two quantitative variables. It indicates whether an increase in one variable tends to be associated with an increase (positive) or decrease (negative) in the other.

For a sample with nn observation pairs (x1,y1),(x2,y2),…,(xn,yn)(x_1, y_1), (x_2, y_2), \dots, (x_n, y_n), the sample covariance sxys_{xy} is:

sxy=∑i=1n(xi−xˉ)(yi−yˉ)n−1s_{xy} = \frac{\sum_{i=1}^{n} (x_i - \bar{x})(y_i - \bar{y})}{n - 1}

where xˉ\bar{x} and yˉ\bar{y} are sample means. The division by n−1n-1 (rather than nn) follows the same degrees‑of‑freedom reasoning used for sample variance.

Intuition via quadrants

Plot the data with vertical/horizontal lines at xˉ\bar{x} and yˉ\bar{y}:

  • Quadrant I (xi>xˉ,  yi>yˉx_i > \bar{x},\; y_i > \bar{y}) → product (xi−xˉ)(yi−yˉ)>0(x_i-\bar{x})(y_i-\bar{y}) > 0
  • Quadrant III (xi<xˉ,  yi<yˉx_i < \bar{x},\; y_i < \bar{y}) → product >0> 0
  • Quadrant II (xi<xˉ,  yi>yˉx_i < \bar{x},\; y_i > \bar{y}) → product <0< 0
  • Quadrant IV (xi>xˉ,  yi<yˉx_i > \bar{x},\; y_i < \bar{y}) → product <0< 0

If most influential points lie in Quadrants I and III → sxy>0s_{xy} > 0 (positive linear association).
If they lie in Quadrants II and IV → sxy<0s_{xy} < 0 (negative linear association).
If points are evenly spread → sxy≈0s_{xy} \approx 0 (no linear association).

Relation to variance

When x=yx = y, covariance reduces to variance of xx:

sxx=∑(xi−xˉ)2n−1=sx2s_{xx} = \frac{\sum (x_i - \bar{x})^2}{n - 1} = s_x^2

Thus covariance generalises variance to pairs of variables.

Unit dependence – a major flaw

Covariance changes with the units of measurement. For example, measuring income in rupees instead of lakhs scales (xi−xˉ)(x_i - \bar{x}) by 10510^5, inflating sxys_{xy} even though the underlying relationship is unchanged. This makes covariance unsuitable for comparing the strength of association across different datasets.

Example: Hanumantha Pai’s credit card data

Variable pairsxys_{xy}Interpretation
Monthly spend vs. annual income16,109Positive association
Monthly spend vs. household size1,610Positive association (smaller magnitude)

The numerical values are not comparable because units differ (lakhs vs. number of people).

Exam tip: Never compare covariance values across different pairs to judge which relationship is “stronger”. Use the correlation coefficient instead.

Key takeaways – Covariance

  • Measures the direction (positive/negative) but not the strength of linear association.
  • Formula: sxy=∑(xi−xˉ)(yi−yˉ)n−1s_{xy} = \frac{\sum (x_i-\bar{x})(y_i-\bar{y})}{n-1}.
  • Positive → x and y tend to move together; negative → they move opposite.
  • Unit‑dependent – not useful for comparing strength across different variable pairs.

Correlation

The Pearson product‑moment correlation coefficient (sample correlation coefficient) overcomes covariance’s unit dependence by standardising:

rxy=sxysx⋅syr_{xy} = \frac{s_{xy}}{s_x \cdot s_y}

where sxs_x and sys_y are the sample standard deviations of xx and yy.

Properties

  • −1≤rxy≤1-1 \leq r_{xy} \leq 1
  • rxy=+1r_{xy} = +1 iff all points lie on a positively‑sloped straight line (perfect positive linear).
  • rxy=−1r_{xy} = -1 iff all points lie on a negatively‑sloped straight line (perfect negative linear).
  • rxy=0r_{xy} = 0 indicates no linear relationship (non‑linear relationships may still exist).
  • Closer to ±1\pm 1 → stronger linear association.

Worked example – Hanumantha Pai

Income vs. Monthly spend
sxy=16,109s_{xy} = 16{,}109, sx=13.74s_x = 13.74 lakhs, sy=1,454s_y = 1{,}454 rupees.

r=16,10913.74×1,454=0.81r = \frac{16{,}109}{13.74 \times 1{,}454} = 0.81

→ Strong positive linear association.

Household size vs. Monthly spend
sxy=1,610s_{xy} = 1{,}610, sx=1.76s_x = 1.76, sy=1,454s_y = 1{,}454.

r=1,6101.76×1,454=0.63r = \frac{1{,}610}{1.76 \times 1{,}454} = 0.63

→ Moderate positive linear association (weaker than income).

CorrelationStrengthExample value
±1\pm 1Perfect—
>0.8> 0.8Strong0.81 (income)
0.50.5 – 0.80.8Moderate0.63 (household)
<0.5< 0.5Weak—
00None (linear)—

Correlation is NOT causation

A high correlation does not prove that a change in one variable causes a change in the other. In the credit‑card example, higher income does not force higher spending – other factors (preferences, debt) could be at play.

Exam tip: The mantra “correlation does not imply causation” is a frequently tested concept. Always remember that a third (lurking) variable or reverse causality could explain the association.

Limitation – linear only

The Pearson correlation only measures linear relationships. A near‑zero rr can occur even when a strong non‑linear pattern exists (e.g., U‑shaped relationship between utility spending and outdoor temperature). Always inspect a scatter diagram alongside the correlation coefficient.

When variables are not quantitative

If one or both variables are nominal or ordinal, other measures (e.g., Spearman’s rank correlation) are used – these are discussed later.

Key takeaways – Correlation

  • rxy=sxysxsyr_{xy} = \frac{s_{xy}}{s_x s_y}, dimensionless and scale‑invariant.
  • Lies between −1-1 (perfect negative) and +1+1 (perfect positive).
  • 00 means no linear relationship – non‑linear patterns may still exist.
  • Correlation is not causation.
  • Always check a scatter plot to avoid being misled by a near‑zero rr from a non‑linear relationship.

Covariance and Correlation

Covariance measures the direction of the linear relationship between two variables: positive → both move together; negative → one goes up as the other goes down. Correlation standardises covariance onto a [−1,1][-1, 1] scale, giving a unitless measure of linear association strength.

Computing in Excel

  • Sample covariance (divides by n−1n-1): =COVARIANCE.S(array1, array2)
  • Sample correlation (Pearson’s rr): =CORREL(array1, array2)

Applied to Hanumantha’s credit card data:

VariablesCovarianceCorrelation
Monthly spend & household income16,1090.806
Monthly spend & household size1,6100.629

The Pearson correlation measures only linear association

r=sxysxsy,sxy=1n−1∑i=1n(xi−xˉ)(yi−yˉ)r = \frac{s_{xy}}{s_x s_y}, \qquad s_{xy} = \frac{1}{n-1}\sum_{i=1}^n (x_i - \bar{x})(y_i - \bar{y})

Worked example (nonlinear relationship):
x={3,2,1,−1,−2,−3}x = \{3, 2, 1, -1, -2, -3\}, y=1/x2={19,14,1,1,14,19}y = 1/x^2 = \{\frac{1}{9}, \frac{1}{4}, 1, 1, \frac{1}{4}, \frac{1}{9}\}.
Excel gives sxy=0s_{xy} = 0, r=0r = 0. But y=1/x2y = 1/x^2 is a perfect nonlinear relationship.

Exam tip: r=0r = 0 does not mean “no association” – it means no linear association. Always inspect a scatter plot.

Spurious vs. Explainable Correlation

  • Explainable correlation: logical reason connects the variables.
  • Spurious correlation: apparent relationship due to a hidden (latent) variable or pure coincidence.
ExampleObserved correlationHidden cause
Ice cream sales ↔ crime ratePositiveTemperature (summer → more ice cream + more vacant homes)
Number of doctors ↔ number of deaths in villagesPositivePoor public health → more doctors needed + more deaths
Divorce rate in Maine ↔ margarine consumptionr≈0.99r \approx 0.99 (1999–2009)No known hidden variable – purely spurious
Skirt length ↔ economic activity (Scotland theory)Skirts shorter in boom, longer in bustProposed reason: affordability of stockings; widely considered spurious

Key takeaways

  • Covariance and Pearson correlation quantify linear association only.
  • r=0r = 0 does not imply independence; check for nonlinear patterns.
  • Spurious correlation can arise from hidden variables or random chance – never assume causation.
  • Sample correlation is a point estimate; a hypothesis test (using a tt-distribution) can assess if the population correlation is nonzero.

Simple Linear Regression – I

Simple linear regression (SLR) predicts a continuous dependent (response) variable yy using one independent (predictor) variable xx. It assumes the relationship can be approximated by a straight line.

Origin and terminology

  • Francis Galton (1886) studied heights of 905 children from 205 families. He multiplied female heights by 1.08 to make them male-equivalent. He observed regression toward the mean: tall parents had slightly shorter children, short parents had slightly taller children. He coined the term regression.
  • yy = dependent / response / outcome variable
  • xx = independent / predictor / explanatory variable / feature

The model

yi=β0+β1xi+εi,i=1,2,…,ny_i = \beta_0 + \beta_1 x_i + \varepsilon_i, \quad i = 1,2,\dots,n

  • β0\beta_0 (intercept), β1\beta_1 (slope) – regression coefficients (parameters)
  • β0+β1xi\beta_0 + \beta_1 x_i = predicted value (conditional expected value of yy given xix_i)
  • εi\varepsilon_i = random error (residual/unexplained variation)

What “linear” really means

Linearity is defined with respect to the parameters β0,β1\beta_0, \beta_1, not the variable xx.

  • yi=β0+β1xi2+εiy_i = \beta_0 + \beta_1 x_i^2 + \varepsilon_i → linear regression (parameters are linear)
  • yi=β0+β1log⁡(xi)+εiy_i = \beta_0 + \beta_1 \log(x_i) + \varepsilon_i → linear regression
  • yi=β0+eβ1xi+εiy_i = \beta_0 + e^{\beta_1} x_i + \varepsilon_i → nonlinear regression (parameter β1\beta_1 appears in exponent)

Exam tip: A model is linear if it can be written as a linear combination of the parameters. Transformations of xx are allowed; transformations of β\beta are not (unless the model can be re‑linearised).

Examples of SLR applications

  • Monthly credit card spend vs. household income or household size
  • Government policy impact on wages
  • Restaurant waiting time vs. popularity rating
  • Hospital treatment cost vs. patient weight
  • Website visits vs. e‑commerce sales revenue
  • Unemployment rate vs. bank non‑performing assets

Nine steps of building a regression model

  • Data preprocessing: handle outliers, missing values; derive new features (e.g., Galton’s 1.08 factor).
  • Training/validation split: random sampling; typical 75% training, 25% validation.
  • Ordinary Least Squares (OLS): finds β0,β1\beta_0, \beta_1 that minimise ∑(yi−(β0+β1xi))2\sum (y_i - (\beta_0 + \beta_1 x_i))^2. OLS gives the best linear unbiased estimates (BLUE).
  • Diagnostics: check assumptions (linearity, normality of errors, homoscedasticity, independence).
  • Overfitting: model performs well on training data but poorly on validation data. An ideal model has low bias and low variance.

Key takeaways

  • SLR models a continuous yy using one xx; parameters are linear in β0,β1\beta_0, \beta_1.
  • “Linear” refers to coefficients, not the shape of xx.
  • OLS minimises sum of squared residuals.
  • Model building is iterative: collect → preprocess → split → describe → specify → estimate → diagnose → validate → deploy.
  • Spurious correlation does not imply regression causation; always use domain logic and diagnostics.

The Regression Equation and Its Interpretation

The regression equation for a sample is written as

yi=β0+β1xi+εiy_i = \beta_0 + \beta_1 x_i + \varepsilon_i

where

  • yiy_i: observed value of the dependent variable for the iith observation,
  • xix_i: observed value of the independent variable for the iith observation,
  • εi\varepsilon_i: random error (also called residual or unexplained variation),
  • β0,β1\beta_0, \beta_1: regression parameters (coefficients).

What the equation really means.
The population can be viewed as a collection of subpopulations — one for each distinct xx value. For example, among Hanumanthappa’s bank customers, all customers with an annual income of ₹20 lakh form one subpopulation, those with ₹30 lakh another, and so on. Each subpopulation has its own distribution of monthly spend yy and its own expected value E(y)E(y). The regression equation describes how this expected value depends on xx:

E(y∣x)=β0+β1xE(y|x) = \beta_0 + \beta_1 x

The error term disappears when taking the expectation because E(εi)=0E(\varepsilon_i)=0.

Estimation Using Sample Data

Population parameters β0\beta_0 and β1\beta_1 are unknown. They are estimated from a sample (the training data) by sample statistics b0b_0 and b1b_1. Substituting these estimates gives the estimated regression equation:

y^=b0+b1x\hat{y} = b_0 + b_1 x

where y^\hat{y} is the point estimator of:

  • the mean value of yy for a given xx (e.g., mean monthly spend for all customers with ₹20 lakh income), and
  • the predicted value of yy for an individual observation at that xx (e.g., predicted spend for one particular customer with ₹20 lakh income).

Exam tip: The same formula y^=b0+b1x\hat{y}=b_0+b_1x serves both for estimating a conditional mean and for predicting an individual outcome. The difference matters when constructing confidence/prediction intervals (later).

Ordinary Least Squares (OLS) Estimation

The ordinary least squares (OLS) method finds b0b_0 and b1b_1 that minimize the sum of squared deviations between observed yiy_i and predicted y^i\hat{y}_i:

Minimize ∑i=1n(yi−y^i)2\text{Minimize } \sum_{i=1}^n (y_i - \hat{y}_i)^2

Using calculus, the solution is:

b1=∑i=1n(xi−xˉ)(yi−yˉ)∑i=1n(xi−xˉ)2b_1 = \frac{\sum_{i=1}^n (x_i - \bar{x})(y_i - \bar{y})}{\sum_{i=1}^n (x_i - \bar{x})^2}

b0=yˉ−b1xˉb_0 = \bar{y} - b_1 \bar{x}

  • xˉ\bar{x}, yˉ\bar{y} are sample means of xx and yy.
  • The slope b1b_1 estimates the expected change in yy for a one-unit increase in xx.
  • The intercept b0b_0 estimates E(y)E(y) when x=0x=0 (meaningful only if x=0x=0 is within the data range).

Caution: Predictions using xx values outside the range of the training data are unreliable — the linear relationship may not hold beyond the observed data.

Worked Example: Hanumanthappa’s Bank Customers

Data: monthly spend (yy) vs. annual income in lakhs (xx) for 55 customers.

StatisticValue
xˉ\bar{x}32.89 lakhs
yˉ\bar{y}₹3,852
b1b_185.27
b0b_01,047.37

Estimated regression equation:

y^=1, ⁣047.37+85.27 x\hat{y} = 1,\!047.37 + 85.27\,x

Interpretation of slope: An increase of ₹1 lakh in annual income is associated with an increase of ₹85.27 in monthly spend.

Prediction: For a customer with annual income ₹35 lakhs:

y^=1, ⁣047.37+85.27×35=4, ⁣032\hat{y} = 1,\!047.37 + 85.27 \times 35 = 4,\!032

Predicted monthly spend: ₹4,032.

Computing OLS in Excel

  • Create a scatter plot (y vs. x).
  • Right-click a data point → Add Trendline → Linear.
  • Check “Display Equation on chart” and “Display R-squared value on chart”.

Excel outputs the same b0b_0, b1b_1, and also R2R^2.

Coefficient of Determination (R2R^2)

R2R^2 (or coefficient of determination) measures the goodness of fit of the estimated regression equation. It is the percentage of the total variation in yy that is explained by the linear relationship with xx:

R2=variation explained by modeltotal variation in yR^2 = \frac{\text{variation explained by model}}{\text{total variation in } y}

  • Ranges from 0 to 1 (or 0% to 100%).
  • Higher R2R^2 indicates a better fit.
  • In simple linear regression, R2R^2 equals the square of the Pearson correlation rxyr_{xy}.

Exam tip: R2R^2 tells you how much of y’s variability the model explains, not whether the model is correct. Always inspect residual plots before trusting R2R^2.

Key takeaways

  • Regression models E(y∣x)=β0+β1xE(y|x) = \beta_0 + \beta_1 x; estimated by y^=b0+b1x\hat{y} = b_0 + b_1 x.
  • OLS minimizes ∑(yi−y^i)2\sum (y_i - \hat{y}_i)^2; formulas: b1=∑(xi−xˉ)(yi−yˉ)∑(xi−xˉ)2b_1 = \frac{\sum (x_i-\bar{x})(y_i-\bar{y})}{\sum (x_i-\bar{x})^2}, b0=yˉ−b1xˉb_0 = \bar{y} - b_1\bar{x}.
  • Slope interpretation: b1b_1 = estimated change in yy per one‑unit increase in xx.
  • Do not extrapolate beyond the range of xx in the sample.
  • R2R^2 = proportion of variation in yy explained by the model; closer to 1 → better linear fit.

Residuals and Sums of Squares

For the (i)-th observation, the residual is the difference between the observed value (y_i) and the predicted value (\hat{y}_i):

[ \text{residual}_i = y_i - \hat{y}_i ]

Residuals capture the error in using the regression line to predict the sample. The least‑squares method minimises the sum of squared residuals, called the sum of squares due to error (SSE):

[ \text{SSE} = \sum_{i=1}^{n} (y_i - \hat{y}_i)^2 ]

SSE measures the unexplained variation in (y) after fitting the model.

If no predictor (x) were available, the best guess for any (y_i) is the sample mean (\bar{y}). The total sum of squares (SST) measures the total variation in (y):

[ \text{SST} = \sum_{i=1}^{n} (y_i - \bar{y})^2 ]

The sum of squares due to regression (SSR) measures how much the regression predictions (\hat{y}_i) deviate from (\bar{y}):

[ \text{SSR} = \sum_{i=1}^{n} (\hat{y}_i - \bar{y})^2 ]

These three quantities are linked by a fundamental identity:

[ \text{SST} = \text{SSR} + \text{SSE} ]

Total variation = explained variation + unexplained variation.

Worked example (Hanumantha’s data, (n=55)): [ \text{SSE} = 40,017,408,\quad \text{SSR} = 74,180,516,\quad \text{SST} = 114,197,923 ] Check: (74,180,516 + 40,017,408 = 114,197,923).

Coefficient of Determination ( (R^2) )

The coefficient of determination (R^2) is the proportion of total variation explained by the regression:

[ R^2 = \frac{\text{SSR}}{\text{SST}} ]

  • (R^2 = 1) when (\text{SSE}=0) (perfect fit).
  • (R^2 = 0) when (\text{SSR}=0) (model explains nothing).
  • (0 \le R^2 \le 1).

Interpreted as a percentage, (R^2) tells us the percentage of variability in (y) accounted for by the linear relationship with (x). For Hanumantha’s data:

[ R^2 = \frac{74,180,516}{114,197,923} \approx 0.650 \quad\Rightarrow\quad 65% ]

Thus, 65% of the variation in monthly spend is explained by annual income.

Exam tip: A high (R^2) indicates good fit in the sample, but does not guarantee the model is appropriate – assumptions must still be checked.

Key takeaways (decomposition & (R^2))

  • (\text{SST} = \text{SSR} + \text{SSE}) – total variation splits into explained and unexplained parts.
  • (R^2 = \text{SSR}/\text{SST}) – fraction of variation explained by the regression.
  • Higher (R^2) implies a better linear fit, but (R^2) alone does not validate the model.

Correlation Coefficient vs. Coefficient of Determination

The sample correlation coefficient (r) measures the strength and direction of a linear relationship between (x) and (y) ((-1 \le r \le 1)). It relates to (R^2) in simple linear regression:

[ r = \sqrt{R^2} \times \text{sign}(b_1) ]

where (b_1) is the estimated slope from the regression.

  • If (b_1 > 0), (r = +\sqrt{R^2}).
  • If (b_1 < 0), (r = -\sqrt{R^2}).

In Hanumantha’s example, (b_1 = 85.27 > 0) and (R^2 = 0.65), so:

[ r = \sqrt{0.65} \approx 0.806 ]

This indicates a strong positive linear association.

Exam tip: Unlike (r), (R^2) can be used for nonlinear models and multiple regression – it is a more general measure of fit.

Assumptions of the Error Term

For valid inference in simple linear regression, four key assumptions are made about the error term (\varepsilon):

  1. Zero mean: (E(\varepsilon) = 0) for all (x). Hence (E(y) = \beta_0 + \beta_1 x).
  2. Constant variance: (\text{Var}(\varepsilon) = \sigma^2) for all (x) (homoscedasticity).
  3. Independence: The errors are independent; no autocorrelation.
  4. Normality: (\varepsilon) is normally distributed for every (x).

Together, these imply that for any given (x), (y) is normally distributed with mean (\beta_0 + \beta_1 x) and constant variance (\sigma^2).

Test for Significance of the Linear Relationship

We test whether the slope (\beta_1) is different from zero (i.e., whether a linear relationship exists).

[ H_0: \beta_1 = 0,\quad H_a: \beta_1 \neq 0 ]

Estimating (\sigma^2)

The mean square error (MSE) provides an unbiased estimate of (\sigma^2):

[ \text{MSE} = s^2 = \frac{\text{SSE}}{n-2} ]

With (n-2) degrees of freedom (two parameters estimated: (\beta_0, \beta_1)). The standard error of the estimate is (s = \sqrt{\text{MSE}}).

Hanumantha’s data: (n=55), (\text{SSE} = 40,017,408), so

[ \text{MSE} = \frac{40,017,408}{53} \approx 755,045,\quad s = \sqrt{755,045} \approx 868.9 ]

t‑Test

The standard error of the slope (b_1) is:

[ \text{se}(b_1) = \frac{s}{\sqrt{\sum (x_i - \bar{x})^2}} ]

For Hanumantha’s data, (\sum (x_i - \bar{x})^2 = 10,201), so

[ \text{se}(b_1) = \frac{868.9}{\sqrt{10,201}} \approx 8.603 ]

The test statistic follows a t distribution with (n-2) degrees of freedom:

[ t = \frac{b_1 - \beta_1}{\text{se}(b_1)} = \frac{85.27 - 0}{8.603} \approx 9.91 ]

The corresponding p‑value (two‑tailed) is extremely small ((\approx 1.1 \times 10^{-13})), providing strong evidence to reject (H_0). We conclude that a significant linear relationship exists.

Exam tip: For simple linear regression, the t‑test and an F‑test (using SSR/MSR vs MSE) are equivalent. In multiple regression they diverge.

Key takeaways (assumptions & significance test)

  • Error term assumptions: zero mean, constant variance, independence, normality.
  • (\text{MSE} = \text{SSE}/(n-2)) estimates (\sigma^2); (s = \sqrt{\text{MSE}}) is the standard error of the estimate.
  • Test (H_0: \beta_1=0) using (t = b_1/\text{se}(b_1)) with (n-2) df.
  • A very small p‑value (e.g., (<0.05)) indicates a statistically significant linear relationship.

SLR - Summary Statistics

The Excel regression output for Hanumantha's data provides a compact summary of all key statistics. The R square value is 0.65, meaning 65% of the variation in monthly spend is explained by annual income. The correlation between the two variables is the square root of R square: r=0.65=0.806r = \sqrt{0.65} = 0.806, indicating a strong positive linear relationship.

The standard error s=868.9s = 868.9 is the estimate of σ\sigma (the standard deviation of the error term ϵ\epsilon). Its square s2s^2 estimates σ2\sigma^2, the variance of ϵ\epsilon.

ANOVA Table

SourceSSdfMSFp-value
Regression (SSR)74,180,516174,180,51698.25≈0\approx 0
Residual (SSE)40,017,40853755,045
Total114,197,92454
  • Sum of squares regression SSR = 74,180,516.
  • Sum of squares error SSE = 40,017,408.
  • Mean square regression MSR = SSR / 1 = 74,180,516.
  • Mean square error MSE = SSE / 53 = 755,045.
  • F statistic = MSR / MSE = 98.25.
  • The p-value (Significance F) is near zero, so we reject H0:β1=0H_0: \beta_1 = 0 and conclude the linear relationship is statistically significant.

Coefficients and Estimated Regression Equation

CoefficientEstimateStandard Errort Statisticp-value
Intercept (b0b_0)10,470.37———
Annual Income (b1b_1)85.278.69.91≈0\approx 0

Estimated regression equation:
y^=10 470.37+85.27 x\hat{y} = 10\,470.37 + 85.27\,x

  • b1b_1 standard error: sb1=s∑(xi−xˉ)2=868.910 201=8.603≈8.6s_{b_1} = \frac{s}{\sqrt{\sum (x_i - \bar{x})^2}} = \frac{868.9}{\sqrt{10\,201}} = 8.603 \approx 8.6.
  • t test for β1=0\beta_1 = 0: t=b1sb1=85.278.6=9.91t = \frac{b_1}{s_{b_1}} = \frac{85.27}{8.6} = 9.91, with p-value ≈0\approx 0 → reject H0H_0.

Exam tip: Excel’s regression tool outputs all these values at once. The key interpretation steps: check R square, then overall F test, then individual coefficient t tests. The F and t tests yield the same conclusion for simple linear regression (since t2=Ft^2 = F).

Key takeaways

  • R square = 0.65, correlation = 0.806.
  • Standard error s=868.9s = 868.9 estimates σ\sigma.
  • F test (F=98.25F=98.25, p ≈ 0) confirms a significant linear relationship.
  • Estimated equation: y^=10 470.37+85.27 x\hat{y} = 10\,470.37 + 85.27\,x.
  • Slope’s t test (t=9.91t=9.91, p ≈ 0) matches F test result.

SLR - Residual Analysis

Residual for observation ii: ei=yi−y^ie_i = y_i - \hat{y}_i (observed minus predicted). Residuals are the best empirical proxy for the error term ϵ\epsilon. Three model assumptions about ϵ\epsilon must be checked:

  1. Expected value E(ϵ)=0E(\epsilon) = 0 (automatically satisfied by least squares).
  2. Constant variance — homoscedasticity — Var(ϵ)=σ2\text{Var}(\epsilon) = \sigma^2 for all xx.
  3. Independence of errors.
  4. Normality — ϵ∼N(0,σ2)\epsilon \sim N(0,\sigma^2).

If assumptions fail, hypothesis tests (t and F) may be invalid.

Residual Plots

Three diagnostic plots are commonly examined:

  1. Residual vs. independent variable xx (provided by Excel).
  2. Residual vs. predicted values y^\hat{y} (can be created from output).
  3. Standardized residual plot (also provided by Excel).

Standardized residual: ei∗=eiseie_i^* = \frac{e_i}{s_{e_i}} (mean = 0, standard deviation ≈ 1). For normally distributed errors, about 95% of standardized residuals lie between −2 and +2.

Interpreting Residual Plot Patterns

  • Panel A (horizontal band): no obvious violation → model assumptions appear satisfied.
  • Panel B (funnel): variance increases with xx → homoscedasticity assumption violated.
  • Panel C (U-shape): systematic curvature → linear model not appropriate.

For Hanumantha’s data, the residual plot shows a horizontal band pattern (like Panel A), supporting the assumptions. The same pattern appears in the residual vs. y^\hat{y} plot (for simple linear regression, both plots give identical information; for multiple regression, the y^\hat{y} plot is preferred).

Checking Normality with Standardized Residuals

Excel provides standardized residuals. In Hanumantha’s dataset, nearly all standardized residuals lie within ±2, consistent with a standard normal distribution. Thus the normality assumption is not violated.

Exam tip: Residual analysis is qualitative. Look for clear deviations (funnel, U-shape, many points outside ±2). A few points near the boundaries do not automatically invalidate the model.

Key takeaways

  • Residuals estimate ϵ\epsilon and are used to validate model assumptions.
  • Three key plots: residuals vs. xx, residuals vs. y^\hat{y}, standardized residuals.
  • Desired pattern: horizontal band (constant variance, linearity).
  • About 95% of standardized residuals should fall in [−2, +2] for normality.
  • Hanumantha’s data satisfies all assumptions (band pattern, no extreme standardized residuals).

Outliers

An outlier is a data point that does not follow the trend of the rest of the data. In simple linear regression, outliers can be detected via:

  • Scatter plot of yy vs. xx.
  • Standardized residuals — any observation with ∣ei∗∣>2|e_i^*| > 2 (roughly) is a potential outlier.

Handling Outliers

SituationAction
Erroneous data (entry/recording error)Correct the value and re-run regression; R square usually improves.
Violation of model assumptionsConsider a different model (e.g., curvilinear, transformation).
Genuine unusual valueRetain the observation — it may be a chance occurrence.

Influential Observations

An influential observation dramatically changes the estimated regression equation (slope, intercept, or both) when removed. Detection methods:

  • Scatter diagram — look for points that are:
    • An outlier in yy (far from the trend line),
    • An extreme xx value (far from xˉ\bar{x}),
    • A combination of both.

Handling Influential Observations

  1. Check for errors in data collection/recording — correct if possible.
  2. If valid, treat the observation as providing additional insight. Consider collecting more data at intermediate xx values to better understand the relationship.

Exam tip: Outliers and influential points are not automatically bad. Always investigate the reason before discarding. In small datasets, a single influential point can dominate the regression — report it transparently.

Key takeaways

  • Outliers: ∣|standardized residual∣>2| > 2; may be erroneous, model violation, or genuine.
  • Influential observation: extreme xx, outlying yy, or both; can shift the regression line substantially.
  • First step for both: check data accuracy.
  • Valid influential observations may motivate collecting additional data at intermediate xx values.

Test of Significance and Residual Analysis

A fitted regression line is not enough — we must ask whether the observed linear relationship is statistically significant or merely a fluke of sampling. Two hypothesis tests (t-test and F-test) answer this question. Both rely on assumptions about the error term ϵ\epsilon, which are checked through residual analysis.

The Regression Model and Estimation (Review)

The simple linear regression model is
y=β0+β1x+ϵy = \beta_0 + \beta_1 x + \epsilon

Using ordinary least squares (OLS), we obtain estimates b0b_0 and b1b_1, giving the estimated regression equation:

y^=b0+b1x\hat{y} = b_0 + b_1 x

OLS minimises the sum of squares error (SSE):
SSE=∑(yi−y^i)2\text{SSE} = \sum (y_i - \hat{y}_i)^2

Variation in yy is decomposed as

SST=SSR+SSE\text{SST} = \text{SSR} + \text{SSE}

where

  • SST=∑(yi−yˉ)2\text{SST} = \sum (y_i - \bar{y})^2 (total sum of squares)
  • SSR=∑(y^i−yˉ)2\text{SSR} = \sum (\hat{y}_i - \bar{y})^2 (regression sum of squares)

The coefficient of determination R2R^2 measures how well the model explains the data:

R2=SSRSST(0≤R2≤1)R^2 = \frac{\text{SSR}}{\text{SST}} \quad (0 \le R^2 \le 1)

The correlation between yy and xx is R2\sqrt{R^2} with the sign of b1b_1.

Assumptions for Hypothesis Testing

The error term ϵ\epsilon is assumed to be:

  • Normally distributed (for exact p‑values)
  • Mean zero, E(ϵ)=0E(\epsilon)=0
  • Constant variance σ2\sigma^2 (homoscedasticity)
  • Independent across observations

These assumptions are critical for the validity of the tests. The variance σ2\sigma^2 is estimated by the mean square error (MSE):

s2=MSE=SSEn−2s^2 = \text{MSE} = \frac{\text{SSE}}{n-2}

The standard error of b1b_1 is MSE/∑(xi−xˉ)2\sqrt{\text{MSE} / \sum (x_i - \bar{x})^2}.

t‑Test for β1\beta_1

  • Null hypothesis H0H_0: β1=0\beta_1 = 0 (no linear relationship)
  • Alternative HaH_a: β1≠0\beta_1 \neq 0

Test statistic:

t=b1se(b1)t = \frac{b_1}{\text{se}(b_1)}

Under H0H_0, tt follows a tt-distribution with n−2n-2 degrees of freedom. The p-value is computed; if small (e.g., <0.05< 0.05), reject H0H_0 and conclude the linear relationship is statistically significant.

F‑Test for Overall Significance

For simple linear regression, the F‑test tests the same null hypothesis (β1=0\beta_1=0).

MSR=SSR1,MSE=SSEn−2\text{MSR} = \frac{\text{SSR}}{1} \quad,\quad \text{MSE} = \frac{\text{SSE}}{n-2}

Test statistic:

F=MSRMSEF = \frac{\text{MSR}}{\text{MSE}}

If H0H_0 is true, FF is approximately 1. If β1≠0\beta_1\neq 0, F>1F > 1. The FF-statistic has 1 and n−2n-2 degrees of freedom; a small p‑value leads to rejecting H0H_0.

Exam tip: In simple linear regression, the t‑test and F‑test are equivalent — they give the same p‑value. In multiple linear regression, the F‑test tests the overall model while t‑tests test individual coefficients.

Residual Analysis

Residuals ei=yi−y^ie_i = y_i - \hat{y}_i are used to check the assumptions on ϵ\epsilon.

  • Standardised residuals =ei/s⋅1−hii= e_i / s \cdot \sqrt{1 - h_{ii}} (or simpler ei/se_i / s). Values outside ±2\pm 2 suggest possible outliers.
  • Plot residuals vs. xx — should be randomly scattered around zero with constant spread (no funnel shape, no curvilinear pattern).
  • Plot residuals vs. y^\hat{y} (or vs. yy) — same expectation. A clear trend indicates that the model is missing relevant variables or that a linear fit is inadequate.

Key takeaways

  • Both t‑test and F‑test test H0 ⁣:β1=0H_0\!:\beta_1=0; reject → linear relationship is significant.
  • MSE\text{MSE} estimates error variance σ2\sigma^2; s=MSEs = \sqrt{\text{MSE}} is the standard error of the estimate.
  • Residual plots must be examined for violations of constant variance, independence, and linearity.
  • Standardised residuals beyond ±2 may indicate outliers.

Application: Monthly Spend vs. Household Size

Using the same dataset (Hanumantha’s), we now model monthly spend (yy) as a function of household size (xx).

Estimated Regression Equation

From Excel’s regression output:

y^=2083.7+520.1×x\hat{y} = 2083.7 + 520.1 \times x

Model Fit

MeasureValue
R2R^20.396
Correlation rxyr_{xy}0.396=0.63\sqrt{0.396} = 0.63
Standard error ss1140.71

Only 39.6% of the variation in monthly spend is explained by household size — much weaker than the model using annual income (R2=0.65R^2=0.65).

ANOVA and F‑Test

SourceSSdfMSFFp‑value
Regression (SSR)45,233,346145,233,34634.7≈ 0
Error (SSE)68,964,576511,301,218
Total (SST)114,197,92252

F=34.7F = 34.7 with a very small p‑value → reject H0H_0 → linear relationship between monthly spend and household size is statistically significant.

Coefficients and t‑Test

CoefficientValueStd. Errorttp‑value
Intercept b0b_02083.7
Household size b1b_1520.188.25.89≈ 0

The t‑test confirms significance (same conclusion as F‑test).

Residual Diagnostics

  • Standardised residuals: three values exceed +2 in absolute value → potential outliers.
  • Residuals vs. xx: no funnel or curvilinear shape; band is acceptable.
  • Residuals vs. yy: shows a slight trend (not a perfect random band). This suggests that other factors (e.g., income) also influence monthly spend — a multiple regression model may improve fit.

Exam tip: Always check residual plots, not just R2R^2 and p‑values. A trend in residuals vs. fitted values is a red flag that important predictors are missing.

Key takeaways

  • R2=0.396R^2 = 0.396 means household size alone explains only 39.6% of spend variation.
  • Despite low R2R^2, the linear relationship is statistically significant (p ≈ 0 for both t and F).
  • Three observations are flagged as possible outliers (standardised residual > 2).
  • Residual pattern suggests the model is incomplete; include other variables like income in multiple regression.

Simple Linear Regression — Applications: Examples II

Two worked examples demonstrate the full SLR workflow: check scatter plot, fit model, interpret R², evaluate significance via F-test and t-test, examine residuals for violations and missing predictors.


Example 1: Buzzer Rogers Customer Satisfaction

Context — 36 survey responses. Question: Is price a significant determinant of satisfaction score?

  • Dependent variable yy = satisfaction score (0–100)
  • Independent variable xx = price paid (₹ 1000s)

Step 1 — Scatter plot

Price on x‑axis, score on y‑axis. A linear trend is visible, though with spread.

Step 2 — Regression equation & R²

From the trend line (and later from the tool pack):

y^=48.44+0.132 x\hat{y} = 48.44 + 0.132\,x
  • R² = 0.36 — only 36 % of score variation is explained by price. The plot showed a fair trend, so the low R² hints that other factors (brand, RAM, etc.) matter.
  • Correlation r=0.36=0.59r = \sqrt{0.36} = 0.59 (positive, moderate).

Step 3 — ANOVA and overall significance

SourceSSdfMSFFpp
Regression (SSR)510.81510.8= MSR/MSE = 18.9≈ 0
Error (SSE)914.83426.9
Total1425.635
  • FF = 18.9, p≈0p \approx 0 ⇒\Rightarrow reject H0:β1=0H_0:\beta_1=0. The linear relationship is statistically significant.

Step 4 — Coefficient table

PredictorCoefficientStd. Errorttpp
Intercept (b0b_0)48.44———
Price (b1b_1)0.1320.034.36≈ 0
  • t=4.36t = 4.36 matches F=t2F = t^2 (4.362≈19.04.36^2 \approx 19.0; rounding). pp small → reject H0H_0.
  • Interpretation: For every ₹ 1000 increase in price, the predicted satisfaction score rises by 0.132 points.

⚠️ Exam tip: This slope is counter‑intuitive (higher price → higher satisfaction). Regression describes the sample; the data may reflect that higher‑priced computers satisfy certain customers. Always check real‑world plausibility separately.

Step 5 — Residual diagnostics

  • Standardized residuals: Two values below –2 (–2.48, –2.12) — possible outliers (points far below the regression line).
  • Residuals vs. x: No funnel or curve, but not a perfect band. Some pattern suggests omitted variables (e.g., brand, features).
  • Residuals vs. y: Shows a trend — again hints that more predictors are needed.

Key takeaways — Example 1

  • R² = 0.36: price alone explains little of satisfaction.
  • Relationship is significant despite low R² (FF test rejects β1=0\beta_1=0).
  • Slope is positive — data says higher price was associated with higher satisfaction in this sample.
  • Residual analysis flags outliers and missing predictors → multiple regression needed.

Example 2: Manjula Nayak — Household Debt (99 Acres)

Context — ~500 Bengaluru households. Question: Is first income a significant determinant of household debt?

  • Dependent variable yy = household debt (₹)
  • Independent variable xx = first income (₹)

Step 1 — Scatter plot

First income vs. debt. A linear trend is barely visible; large variation in debt for a given income.

Step 2 — Regression equation & R²

y^=32 976+0.04 x\hat{y} = 32\,976 + 0.04\,x
  • R² = 0.31 — income only explains 31 % of debt variation.
  • Correlation r=0.31≈0.56r = \sqrt{0.31} \approx 0.56 (positive, moderate).

Step 3 — ANOVA and overall significance

SourceSSdfMSFFpp
Regression (SSR)1.86×10111.86 \times 10^{11}11.86×10111.86 \times 10^{11}= 228≈ 0
Error (SSE)4.08×10114.08 \times 10^{11}4988.19×1088.19 \times 10^8
Total5.94×10115.94 \times 10^{11}499
  • FF = 228, p≈0p \approx 0 → reject H0H_0. Linear relationship is significant.

Step 4 — Coefficient table

PredictorCoefficientStd. Errorttpp
Intercept (b0b_0)32 976———
First income (b1b_1)0.040.00215.1≈ 0
  • t=15.1t = 15.1 (very large). pp tiny → strong evidence against β1=0\beta_1=0.
  • Interpretation: For every ₹ 1 increase in first income, debt rises by ₹ 0.04.

Step 5 — Residual diagnostics

  • Standardized residuals: Few values outside ±2, but mostly within range.
  • Residuals vs. x: Oval shape — not a perfect band, no funnel or curve.
  • Residuals vs. y: Shows a trend, again hinting that other variables (ownership, monthly payments, etc.) affect debt.

Key takeaways — Example 2

  • R² = 0.31: income alone explains little of debt.
  • Relationship is highly significant (F=228F=228, t=15.1t=15.1).
  • Slope 0.04 means a 1‑rupee increase in income is associated with 0.04 rupee more debt.
  • Residual patterns signal that a multiple regression (next module) will better capture debt drivers.

Common workflow across both examples

Exam tip: When R² is low but the overall FF‑test is significant, the model is still statistically useful but practically weak — other variables are needed. Residual plots that show trends (not random scatter) are the main clue.

SLR Applications: Monthly Payment vs. Debt

Intuition: Households with higher monthly payments (loans, credit cards, utilities) are expected to carry more debt. This second example uses monthly payment as the independent variable xx to predict debt yy, comparing results with the earlier income-based model.

Estimated Regression Equation

Using the same 500-household dataset, the least‑squares regression yields:

y^=b0+b1x=28 791+2.94 x\hat{y} = b_0 + b_1 x = 28\,791 + 2.94\,x

  • b1=2.94b_1 = 2.94: for every one‑rupee increase in monthly payment, predicted debt increases by 2.94 rupees.
  • Caution: With only one predictor, this coefficient captures both the direct effect and correlations with omitted variables; it is not a pure marginal effect.

Goodness‑of‑Fit: R2R^2 and Correlation

MeasureValueInterpretation
R2R^20.37Monthly payment explains 37% of the variability in debt.
r=R2r = \sqrt{R^2}0.60Positive correlation between monthly payment and debt.

Comparison with earlier model (Debt vs. First Income):

ModelR2R^2rr
Debt vs. Income0.310.56
Debt vs. Monthly Payment0.370.60

The monthly‑payment model is marginally better, but both leave >60% of variance unexplained – strong evidence that debt depends on multiple factors.

ANOVA and Hypothesis Tests

Null hypothesis: H0:β1=0H_0: \beta_1 = 0 (no linear relationship).
Alternative: H1:β1≠0H_1: \beta_1 \neq 0.

SourceSum of SquaresdfMean SquareFFpp-value
Regression (SSR)2.18×10112.18 \times 10^{11}12.18×10112.18 \times 10^{11}288≈0\approx 0
Error (SSE)3.80×10113.80 \times 10^{11}498757 268 785757\,268\,785
Total (SST)499
  • Standard error of estimate: MSE=757 268 785≈27 518\sqrt{\text{MSE}} = \sqrt{757\,268\,785} \approx 27\,518 rupees – typical prediction error.
  • F=288F = 288, p≈0p \approx 0 → reject H0H_0; the linear relationship is statistically significant.

Coefficient details for b1b_1:

CoefficientStd. Errortt statisticpp-value
Monthly payment (b1b_1)2.940.174t=2.94/0.174=16.95t = 2.94 / 0.174 = 16.95≈0\approx 0
  • The tt‑test confirms β1≠0\beta_1 \neq 0, identical conclusion to the FF‑test.

Residual Analysis

Standardized residuals: Several observations lie below −2 or above +2, suggesting potential outliers.

Residual vs. xx (monthly payment): Forms an oval shape – not a perfect random band, but no obvious funnel or curvature. This points to heteroscedasticity or non‑linearity that may be addressed by including additional predictors.

Residual vs. y^\hat{y}: Similar pattern reinforces that the model fails to capture all systematic variation.

Exam tip: An oval‑shaped residual plot (instead of a horizontal band) is a classic sign that a simple linear regression is misspecified – often resolved by moving to multiple regression.

Key takeaways

  • Monthly payment explains 37% of debt variability (R2=0.37R^2 = 0.37, r=0.60r = 0.60), slightly better than income (31%).
  • Both slope and overall model are highly significant (F=288F=288, t=16.95t=16.95, p≈0p \approx 0).
  • Residual analysis reveals several outliers and a non‑random pattern, indicating omitted variables.
  • Coefficient b1=2.94b_1 = 2.94 cannot be interpreted as a pure marginal effect in simple regression.

Module Summary

This module introduced the fundamentals of measuring and modelling linear relationships between two variables.

Key concepts covered:

  1. Covariance and correlation – quantify the direction and strength of linear association. Cov(X,Y)=∑(xi−xˉ)(yi−yˉ)n−1,r=Covsxsy\text{Cov}(X,Y) = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{n-1}, \quad r = \frac{\text{Cov}}{s_x s_y}

    • Distinguish explainable vs. spurious correlation.
  2. Simple linear regression model Y=β0+β1X+ϵY = \beta_0 + \beta_1 X + \epsilon with E(Y)=β0+β1X\mathbb{E}(Y) = \beta_0 + \beta_1 X.

  3. Least‑squares estimation – sample statistics b0b_0 and b1b_1 estimate population parameters.

  4. Coefficient of determination R2R^2 – proportion of YY’s variability explained by XX.

  5. Assumptions about errors ϵ\epsilon – independence, constant variance, normality – required for valid inference.

  6. Hypothesis tests – tt‑test for β1\beta_1 and FF‑test for overall significance (equivalent in SLR).

  7. Residual analysis – plots and standardized residuals used to check assumptions and detect outliers/influential points.

The next module extends this framework to multiple linear regression, where YY depends on several XX variables.

Key takeaways

  • Covariance and correlation measure association; regression models the relationship.
  • R2R^2 is the key measure of fit; low R2R^2 indicates omitted variables.
  • Hypothesis tests (tt, FF) assess statistical significance of the linear relationship.
  • Residual plots are essential for model validation.