Term 2 · Module 3 of 4

Predictive Analysis

Advanced Statistics for Business

Concept of Forecasting - Introduction

Forecasting is the process of making predictions about future events based on historical data and various statistical techniques. In business analytics, forecasting enables organizations to be proactive rather than reactive: by identifying patterns in past data, companies can anticipate demand, costs, revenue, customer behavior, and other critical variables subject to fluctuation.

Forecasting supports decision-making and strategic planning across all horizons. Common methods include regression models, time series analysis, machine learning algorithms, and econometric models. This module focuses specifically on regression models and time series analysis for forecasting.

Short‑Term vs. Long‑Term Forecasting

  • Short‑term forecasting – e.g., predicting sales for the next week or month.
  • Long‑term forecasting – e.g., projecting growth over several years or a decade.

Exam tip: The same method rarely performs well for both horizons. Short‑term forecasts often emphasize recent patterns; long‑term forecasts use broader trends and structural factors.

Applications of Forecasting in Business

Application AreaPurposeExample
Demand forecastingPredict future product/service demand to align inventory with expected sales.A retail store uses prior‑year sales, current trends, and customer purchasing behavior to avoid stockouts and overstocking.
Financial forecastingPredict revenue, expenses, cash flow, and profit for budgeting and strategic planning.A tech startup forecasts cash inflows/outflows to determine when it will achieve profitability and when additional capital may be needed.
Sales forecastingEstimate future sales volumes to set revenue targets, allocate resources, and manage production.A car dealership predicts seasonal demand for specific models (e.g., higher sales in summer) to stock appropriate quantities.
Workforce forecastingAnticipate future staffing needs, especially for seasonal or growth‑driven fluctuations.A hotel chain uses past booking data to hire adequate staff for the high‑tourist season.
Risk management forecastingIdentify potential risks and take preventive measures.Banks use historical loan‑repayment data and credit scores to forecast the likelihood of loan defaults, improving lending decisions.
Fraud detection forecastingDetect abnormal transaction patterns that may indicate fraud.Credit‑card companies build models that flag transactions deviating from forecasted customer behavior.
Marketing / Customer churn forecastingPredict customer behavior (e.g., likelihood of churn or purchase) to tailor retention and promotional efforts.A streaming service (e.g., Hotstar) forecasts churn risk based on viewing habits and subscription tenure, then targets at‑risk customers with retention offers.

Why Forecasting Is Critical

  • Demand forecasting prevents both stockouts (lost sales, dissatisfied customers) and overstocking (tied‑up capital, storage costs).
  • Financial forecasting helps allocate resources, manage cash flow, and communicate growth potential to investors.
  • Sales forecasting guides production schedules, inventory levels, and sales force deployment.
  • Workforce forecasting ensures adequate staffing during peak periods without over‑hiring.
  • Risk forecasting minimizes losses (e.g., loan defaults) and maintains financial stability.
  • Fraud detection protects both the institution and its customers from unauthorized activity.
  • Customer behavior forecasting (e.g., churn, purchase likelihood) enables personalized marketing, improving conversion rates and customer experience.

Key takeaways

  • Forecasting = making data‑driven predictions about future events using historical data and statistical techniques.
  • Methods differ by horizon (short‑term vs. long‑term); one method rarely fits both.
  • Major applications include demand, financial, sales, workforce, risk, fraud, and marketing forecasting.
  • Accurate forecasts reduce uncertainty, optimize operations, and support proactive strategic planning.
  • This module will cover regression‑based and time‑series forecasting techniques in depth.

Predictive Analysis of Continuous Data

Linear regression is the most common approach for identifying relationships between variables and is widely used for forecasting continuous data. It models the relationship between one or more independent variables (predictors) and a continuous response variable (dependent variable) by fitting a linear equation to the data. Once the model is trained, new values of the independent variables can be plugged in to predict future outcomes.

Example: A company forecasts monthly sales using advertising spend, product prices, and seasonal factors. Historical data on these inputs and actual sales are used to train the regression model, then expected future inputs generate sales forecasts.

The effectiveness of linear regression for forecasting depends on how well the model generalises to unseen data—which is evaluated using cross-validation.


Cross-Validation for Evaluating Forecasting Models

Cross-validation assesses predictive performance by splitting data into multiple subsets. A typical approach:

  1. Partition the data into a training set and a testing set.
  2. Train the regression model on the training set (it never sees the testing set).
  3. Evaluate performance on the testing set (unseen data) to measure generalisation.

During training, the model learns relationships by minimising the least‑squared error; variable selection can be guided by (R^2). Performance on the testing set checks for overfitting—good training performance but poor testing indicates the model has captured noise rather than genuine patterns.

K‑Fold Cross-Validation

Instead of one fixed split, the data is divided into (k) equal subsets (folds). The model is trained (k) times, each time using (k-1) folds as training and the remaining fold as testing. Accuracy is averaged across all (k) iterations, yielding a more robust estimate—especially useful when data is limited.

Exam tip: K‑fold cross-validation ensures every observation is used for both training and testing, reducing the variance of the performance estimate. It is critical when selecting the best set of independent variables for forecasting.


Evaluating Predictive Accuracy for Continuous Data

Several metrics measure the difference between predicted ((\hat{y}_i)) and actual ((y_i)) values. Each provides different insight.

Common Metrics

MetricFormulaUnitInterpretation
R‑squared ((R^2))Proportion of variance explained by modelUnitless (0–1)Higher → more variance captured, but may overfit
Mean Absolute Error (MAE)(\displaystyle \frac{1}{n}\sum_{i=1}^{n}y_i - \hat{y}_i)
Mean Squared Error (MSE)(\displaystyle \frac{1}{n}\sum_{i=1}^{n} (y_i - \hat{y}_i)^2)Squared units of dependent variablePenalises large errors more heavily
Root Mean Squared Error (RMSE)(\displaystyle \sqrt{\frac{1}{n}\sum_{i=1}^{n} (y_i - \hat{y}_i)^2})Same as dependent variableInterpretable scale; sensitive to large errors
Mean Absolute Percentage Error (MAPE)(\displaystyle \frac{100%}{n}\sum_{i=1}^{n} \frac{y_i - \hat{y}_i}{y_i})

Detailed Notes on Each Metric

  • (R^2) (Coefficient of Determination): Represents the proportion of variance in the dependent variable explained by the model. Values range 0 to 1; 1 indicates perfect fit. High (R^2) can be misleading due to overfitting—it does not indicate absolute predictive accuracy and should be complemented by other metrics.

  • MAE: Average absolute error. For the house‑price dataset (price in lakhs of ₹), MAE is also in lakhs of ₹, giving a direct, intuitive measure of typical prediction error.

  • MSE: Squared error penalises outliers more than MAE. Because it is in squared units, it is harder to interpret directly. Useful when large errors are considered particularly harmful.

  • RMSE: The square root of MSE brings the unit back to the original scale (e.g., lakhs of ₹). It is the most interpretable error metric that still penalises large deviations.

  • MAPE: Expresses error as a percentage of actual value. For example, a 10% MAPE means predictions deviate on average by 10%. Especially valuable in finance and retail where relative accuracy drives decisions.

Exam tip: No single metric is sufficient. Use RMSE to detect large errors, MAE for average error, MAPE for relative error, and (R^2) for model comparison. Always evaluate on a held‑out test set via cross‑validation.

Key takeaways

  • Linear regression forecasts continuous outcomes by fitting a linear equation; the model must generalise to unseen data.
  • Cross‑validation (especially k‑fold) robustly evaluates generalisation by training/testing on multiple data splits.
  • Predictive accuracy is measured with (R^2), MAE, MSE, RMSE, and MAPE—each highlights different aspects of error.
  • Overfitting is detected when training performance is strong but test performance is weak.
  • Errors on the house‑price example: MAE and RMSE in lakhs of ₹, MAPE as a percentage of the actual price.

The Problem Setup

We apply linear regression to a continuous target—house prices (in lakhs of rupees). The dataset contains 4,340 apartments. For validation, the first 3,000 observations form the training set and the remaining 1,340 form the test set. The goal: use apartment features to predict the price of an unseen house.

Features considered:

  • RERA approval (binary)
  • Number of bedrooms
  • Area in sq. ft.
  • Ready to move (binary)
  • Resale (binary)
  • Posted by (categorical: builder, owner, dealer)

Because posted by is categorical, it is encoded using a baseline category. Builder is set as the baseline; two binary dummy variables are created:

  • posted_by_owner (1 = owner, 0 otherwise)
  • posted_by_dealer (1 = dealer, 0 otherwise)

The Regression Model

A linear regression is fitted on the training data:

Price=β0+β1(RERA)+β2(bedrooms)+β3(area)+β4(ready_to_move)+β5(resale)+β6(posted_owner)+β7(posted_dealer)+ε\text{Price} = \beta_0 + \beta_1(\text{RERA}) + \beta_2(\text{bedrooms}) + \beta_3(\text{area}) + \beta_4(\text{ready\_to\_move}) + \beta_5(\text{resale}) + \beta_6(\text{posted\_owner}) + \beta_7(\text{posted\_dealer}) + \varepsilon

Estimated coefficients from the output:

FeatureCoefficientInterpretation
Intercept571Baseline average price (lakhs) when all features = 0
RERA approval21Add if approved (not statistically significant)
Number of bedrooms87Increase per extra bedroom
Area (sq. ft.)—Positive and significant (value not given, but used)
Ready to move—Significant (coefficient not stated)
Resale—Significant
Posted by owner—Significant (relative to builder)
Posted by dealer—Significant

Exam tip: The intercept (571) is the predicted price when all continuous variables are zero and all dummies are zero (builder baseline, no RERA, not ready, not resale). This is often not practically meaningful but mathematically necessary.

Making Predictions

For a new apartment, the predicted price is the sum of the intercept plus each applicable coefficient. Example procedure:

  1. Start with 571 (intercept)
  2. Add 21 if RERA-approved
  3. Add 87 × number of bedrooms
  4. Add coefficient × area (sq. ft.)
  5. Add coefficient if ready-to-move
  6. Add coefficient if resale
  7. Add coefficient if posted by owner or dealer

Evaluating Forecast Accuracy

For each test observation, compute the error = actual − predicted. Four aggregated metrics:

MAE=1n∑i=1n∣yi−y^i∣MSE=1n∑i=1n(yi−y^i)2\text{MAE} = \frac{1}{n}\sum_{i=1}^{n}|y_i - \hat{y}_i| \qquad \text{MSE} = \frac{1}{n}\sum_{i=1}^{n}(y_i - \hat{y}_i)^2 RMSE=MSEMAPE=100%n∑i=1n∣yi−y^i∣yi\text{RMSE} = \sqrt{\text{MSE}} \qquad \text{MAPE} = \frac{100\%}{n}\sum_{i=1}^{n}\frac{|y_i - \hat{y}_i|}{y_i}

Initial results on test set:

  • MAE > 100 lakhs
  • RMSE high
  • MAPE > 130%

These seem poor. However, a closer look reveals the problem.

The Outlier Problem

A plot of actual vs. predicted values shows the model’s predictions max out around 1,600 lakhs, while some actual prices reach 9,000 lakhs. These extreme values (likely data entry errors or genuine outliers) inflate the error metrics.

After removing the 3–4 most extreme outliers (actual > 9,000), the actual and predicted price ranges become comparable. A scatter of actual vs. predicted clusters around the y=xy=x line, indicating reasonable performance for most observations.

Exam tip: Outliers can completely distort MAE, RMSE, and especially MAPE (since percentage error is sensitive to small denominators). Always examine the distribution before trusting global error metrics.

Improving the Model

Logarithmic transformation on the target variable can reduce skew and mitigate outlier impact:

log⁡(Price)=β0+β1(RERA)+⋯+ε\log(\text{Price}) = \beta_0 + \beta_1(\text{RERA}) + \dots + \varepsilon

After transformation, re-check for outliers and remove them only after careful analysis. This often yields a more normal-like distribution and better predictive performance.

Key takeaways

  • Linear regression is a valuable tool for forecasting continuous data, but its effectiveness must be validated with a train/test split.
  • Categorical features with >2 levels require dummy encoding with one baseline category.
  • Outliers disproportionately affect regression coefficients and error metrics; visualization (actual vs. predicted) is essential.
  • Removing outliers or applying log transformations can greatly improve model reliability.
  • Even with good overall fit, some predictions may still be biased—further feature engineering or alternative models may be needed.

Predictive Analysis of Categorical Data

Logistic regression is a method for modeling binary outcomes – events that take only two values (yes/no, success/failure, purchase/no purchase). Unlike linear regression (which works for continuous response), logistic regression outputs a probability that an event occurs, bounded between 0 and 1. It does this by passing a linear combination of predictors through the logistic function:

p=11+e−(β0+β1x1+⋯+βkxk)p = \frac{1}{1 + e^{-(\beta_0 + \beta_1 x_1 + \dots + \beta_k x_k)}}

Why it matters: businesses can use the predicted probability to classify customers as “likely to purchase” or not, then target marketing efforts accordingly.

Confusion Matrix

The confusion matrix is a 2×2 table that compares predicted classes (based on a probability threshold, typically 0.5) against actual classes.

Predicted Positive (1)Predicted Negative (0)
Actual Positive (1)True Positive (TP)False Negative (FN)
Actual Negative (0)False Positive (FP)True Negative (TN)
  • TP: correctly predicted positive cases.
  • TN: correctly predicted negative cases.
  • FP: predicted positive but actually negative (Type I error).
  • FN: predicted negative but actually positive (Type II error).

Key Metrics Derived from the Confusion Matrix

Accuracy

The fraction of all correct predictions:

Accuracy=TP+TNTP+FP+TN+FN\text{Accuracy} = \frac{TP + TN}{TP + FP + TN + FN}

Exam tip: Accuracy can be misleading when classes are imbalanced (e.g., 95% negative, 5% positive). A model that always predicts “negative” would achieve 95% accuracy but is useless for identifying positives.

Precision (Positive Predictive Value)

Out of all cases predicted as positive, how many were actually positive?

Precision=TPTP+FP\text{Precision} = \frac{TP}{TP + FP}

A high precision means few false positives.

Recall (Sensitivity, True Positive Rate)

Out of all actual positive cases, how many were correctly predicted?

Recall=TPTP+FN\text{Recall} = \frac{TP}{TP + FN}

A high recall means few false negatives.

F1 Score

The harmonic mean of precision and recall, balancing both:

F1=2⋅Precision⋅RecallPrecision+RecallF1 = 2 \cdot \frac{\text{Precision} \cdot \text{Recall}}{\text{Precision} + \text{Recall}}

  • Always between 0 and 1; >0.5 is better than random.
  • High only when both precision and recall are high.

Metrics trade-off: increasing precision often reduces recall and vice versa. The choice depends on business cost – e.g., in medical screening, missing a disease (low recall) may be worse than a false alarm (low precision).

AUC‑ROC Curve

The ROC curve plots Recall (TPR) against False Positive Rate (FPR) for all possible classification thresholds (from 0% to 100%).

  • Threshold: the cutoff probability above which we predict “positive” (e.g., 0.5, 0.7). Changing threshold shifts TP/FP counts.
  • Area Under the Curve (AUC) summarizes overall performance: AUC = 1 means perfect discrimination; AUC = 0.5 means random. AUC closer to 1 indicates better ability to separate the two classes.
  • Particularly useful when the cost of misclassification varies (e.g., credit scoring, churn prediction).

Probabilistic Accuracy: Quadratic Score

Metrics above only evaluate classification correctness (which side of a threshold the predicted probability falls), not how well the probability itself matches reality.

Quadratic score (also called Brier score) measures the accuracy of the predicted probabilities:

QS=1n∑i=1n(yi−p^i)2\text{QS} = \frac{1}{n} \sum_{i=1}^{n} (y_i - \hat{p}_i)^2

  • yiy_i = actual outcome (1 or 0)
  • p^i\hat{p}_i = predicted probability of class 1

Lower values mean better probabilistic predictions (like MSE for continuous data). A model predicting 93% for a positive case is penalised less than one predicting 52%, even if both classify correctly.

Key Takeaways

  • Logistic regression models binary outcomes via the logistic function, producing probabilities between 0 and 1.
  • The confusion matrix is the foundation for computing accuracy, precision, recall, and F1 score.
  • Precision answers “how reliable are positive predictions?”; recall answers “how many positives are caught?”; F1 balances both.
  • AUC‑ROC evaluates discrimination ability across all thresholds – high AUC → good class separation.
  • Quadratic score evaluates the accuracy of the probability estimates themselves, not just the binary classification.
  • Selecting the right evaluation metric depends on business goals and the costs of false positives vs. false negatives.

Logistic Regression for Binary Classification (EdTech Company Case)

Logistic regression extends the linear model to predict a binary outcome — e.g., whether a customer converts (1) or not (0). Instead of modeling the outcome directly, it models the log-odds (logit) of the probability of conversion:

logit(p)=ln⁡(p1−p)=β0+β1x1+⋯+βkxk\text{logit}(p) = \ln\left(\frac{p}{1-p}\right) = \beta_0 + \beta_1 x_1 + \dots + \beta_k x_k

The probability is recovered via the inverse logit (sigmoid) function:

p=elogit(p)1+elogit(p)p = \frac{e^{\text{logit}(p)}}{1 + e^{\text{logit}(p)}}

This keeps predictions in the (0,1) range, making them interpretable as probabilities.

Data and Variables

The EdTech company collected data on 9,240 potential customers. The target is Converted (1 = took a course, 0 = did not). The predictors used in the model:

VariableTypeDescription
Open to EmailBinary (0/1)Agreed to receive emails
Open to CallBinary (0/1)Agreed to receive calls
Total VisitsContinuousNumber of visits to the website
Total Time SpentContinuousMinutes spent on the site
Page Views per VisitContinuousAverage pages viewed per visit
IndianBinary (0/1)Customer from India
UnemployedBinary (0/1)Currently unemployed
StudentBinary (0/1)Currently a student
Better Career ProspectsBinary (0/1)Selected “better career” as primary reason

Training set: first 9,000 observations; test set: remaining ~240 observations.

Estimated Coefficients

The logistic regression coefficients (log-odds effects) are:

PredictorCoefficientInterpretation (on log-odds)
Intercept0.954Baseline log-odds when all predictors = 0
Open to Email1.178Positive: increased chance of conversion
Open to Call–3.981Negative: decreased chance of conversion
Total Visits0.025Small positive effect per additional visit
Total Time Spent(not given)Positive (direction stated)
Page Views per Visit–0.060Negative: more views → lower conversion (possible distraction)
Indian(large negative)Negative coefficient
Unemployed(negative)Negative (unexpected; possibly inability to pay)
Student~0 (near zero)Negligible effect
Better Career Prospects(large positive)Strong positive effect

Exam tip: A negative coefficient does not always mean the variable is “bad” — it means the log-odds of conversion decrease. Domain knowledge is needed to explain counterintuitive signs (e.g., unemployed → negative may reflect price sensitivity).

Worked Example: Predicting Probability for a New Customer

Given a customer with the following data:

  • Open to Email = 1
  • Open to Call = 1
  • Total Visits = 2
  • Total Time Spent = 60 minutes
  • Page Views per Visit = ? (coefficient –0.060 used)
  • Indian = 1
  • Unemployed = 1
  • Student = 0
  • Better Career Prospects = 1

The logit is computed by summing the intercept and each coefficient multiplied by the variable value. Using the coefficients provided (and noting that the time‑spent coefficient was used but its value not disclosed), the cumulative logit is –1.458.

Convert to probability:

p=e−1.4581+e−1.458=0.19p = \frac{e^{-1.458}}{1 + e^{-1.458}} = 0.19

The model predicts only a 19% chance of conversion for this customer.

Model Evaluation on the Test Set

The predicted probabilities are thresholded at 50% to classify. The resulting confusion matrix:

Predicted Converted (1)Predicted Not Converted (0)
Actual Converted (1)TP = 51FN = 45
Actual Not Converted (0)FP = ?TN = 118

Key metrics (as computed from the test set):

MetricFormulaValue
AccuracyTP+TNTP+FP+FN+TN\frac{TP + TN}{TP + FP + FN + TN}69.29%
PrecisionTPTP+FP\frac{TP}{TP + FP}0.638
Recall (Sensitivity)TPTP+FN\frac{TP}{TP + FN}0.531
F1 Score2×Precision×RecallPrecision+Recall2 \times \frac{\text{Precision} \times \text{Recall}}{\text{Precision} + \text{Recall}}0.580
Quadratic (Brier) Score1n∑(p^i−yi)2\frac{1}{n}\sum (\hat{p}_i - y_i)^20.213

The model has decent precision but low recall — it misses many actual converters. The Brier score (0.213, where 0 is perfect and 1 is worst) suggests predictions are reasonably confident when correct.

Feature Engineering

  • Create new features: e.g., interaction terms (Indian × Unemployed × Student), lead origin / source dummies.
  • Transform skewed continuous variables: apply log⁡(x+1)\log(x+1) or x\sqrt{x} to variables like Total Time Spent (likely right‑skewed).
  • Domain‑specific features: incorporate business knowledge (e.g., previous engagement, discount eligibility).

Robust Cross-Validation

  • The current split uses one‑fold (9,000 train / 240 test). Switch to 10‑fold cross‑validation to get more reliable performance estimates and reduce variance due to a single split.

Exam tip: When improving a logistic regression model, always start with feature engineering and transformation before tuning thresholds or trying more complex models. Cross‑validation helps confirm improvements are not coincidental.

Key takeaways

  • Logistic regression models log-odds of a binary outcome; coefficients are additive on the log‑odds scale.
  • The sigmoid transformation converts log-odds to probabilities between 0 and 1.
  • Model evaluation uses confusion‑matrix derived metrics: accuracy, precision, recall, F1, and the Brier (quadratic) score.
  • Counterintuitive coefficient signs (e.g., negative for unemployed) can arise — investigate with domain knowledge.
  • Improvement strategies: feature engineering (interactions, transformations) and robust cross‑validation (e.g., 10‑fold).

Time Series Data

Time series data is data collected at successive points in time (daily, monthly, yearly) where the chronological order carries essential information. Unlike cross‑sectional data, the sequence itself is critical for understanding how values evolve.

Why specialized methods?

Conventional regression (e.g., OLS) treats observations as independent and ignores order, so it cannot detect patterns linked to time:

  • Trend – a steady increase or decrease over time.
  • Seasonality – regular cycles (e.g., higher sales in holiday seasons).

Time‑series methods capture these patterns, yielding better forecasts and informed decisions (e.g., stocking inventory before a seasonal peak).

Worked example: mean vs. trend

Given 7 observations:

Observation12345678 (??)
Set A2.62.12.42.02.52.22.3?
Set B1.51.82.12.42.73.03.3?
  • Set A: No clear pattern → the mean (~2.3) is a reasonable prediction.
  • Set B: Clear upward trend – mean also ≈2.3, but a value around 3.0–3.3 is far more plausible. Using the mean here would be a poor forecast.

Exam tip: Always plot your time series first. A mean forecast works only when no trend or seasonality exists.

Three real‑world cases (from Forecasting: Principles and Practice)

CaseProblemFlawed methodLesson
Disposable tableware manufacturerNeed monthly forecasts for hundreds of items with trend & seasonalityAverage of last 6 or 12 months; linear regression on recent data only; linear trend from last year to this yearUsing only recent data misses long‑term patterns – models must capture trend & seasonality
Australian pharmaceutical benefit schemeUnderestimated expenditure by ~$1 billion/yearIgnored sudden jumps due to policy changes and cheaper competitor drugsExogenous regressors (external info) are essential; not all patterns are in the historical series alone
Car fleet companyForecast vehicle resale values; specialists resisted statistical modelsRelied solely on expert judgmentCombine data, exogenous variables, and expert judgment for robust forecasts

Core patterns in time series

Key takeaways

  • Time series data is ordered chronologically – the order matters.
  • Special methods are needed because conventional regression ignores time‑dependent patterns (trend, seasonality).
  • A simple average (mean) fails when a trend is present.
  • Good forecasting captures trend, seasonality, recent patterns, and exogenous regressors.
  • Expert judgment can complement, but not replace, statistical models.

The French Bakery Sales Dataset

The dataset contains 600 days of sales (quantity sold) from a French bakery, spanning 2 Jan 2021 to 30 Sep 2022. Columns:

  • Date (daily)
  • Quantity (a discrete count of products sold, unlike continuous sales revenue)
  • Day of week (for seasonality analysis)

Time‑series cross‑validation requires temporal order. When evaluating forecasting models, use the earlier part as the training set and the most recent part as the test set (never random splits). Here: train on all data up to 31 Aug 2022, test on September 2022 (last month).

Core Concepts: Trend and Seasonality

ComponentWhat it isReal‑world exampleWhy it matters
TrendLong‑term overall direction (upward, downward, flat, or piecewise changing)A growing company’s annual revenue → upward trend; a product becoming obsolete → eventual declineStrategic decisions: forecast growth, identify decline periods
SeasonalityRegular, predictable pattern repeating at fixed intervals (days, weeks, months, quarters)Bakery sales peak on weekends, drop on weekdaysAnticipate short‑term fluctuations; prepare resources

Trend can also be sinusoidal (e.g., winter wear – rises Sep–Jan, falls after, repeats yearly). Both components together let businesses separate a temporary seasonal dip from a long‑term declining trend.

Detecting Trend

1. Visual Inspection

Plot the time series as a line chart. For the bakery data:

  • Early 2021: slight increase → then decline → again increase in 2022.
  • Values fluctuate regularly (hint of weekly seasonality).

2. Moving Average

Smooths short‑term fluctuations to reveal the underlying trend.

Moving Averaget=1k∑i=t−k+1tyi\text{Moving Average}_t = \frac{1}{k} \sum_{i=t-k+1}^{t} y_i

  • Period kk = window size. In Excel: add trendline → moving average → set period.
  • k=7k=7 (weekly): reveals short‑term trend.
  • k=30k=30 (monthly): shows month‑to‑month direction.

3. Linear Trend Line

Fits a straight line through all data points – assumes constant increase/decrease. The bakery data shows a very slight upward linear trend from start to end. Displays equation on chart.

4. Piecewise Linear Trend (mentioned but not detailed)

Better for non‑constant trends: fit separate straight lines to different time segments. Will be used later when specifying trend functions.

Detecting Seasonality

Step 1 – Compute averages per seasonal period

For weekly seasonality: group data by day of week and compute average quantity sold.

DayAverage Quantity
Monday to Friday700–850
Saturday>1000
Sunday>1000

Weekends clearly above weekdays – strong weekly seasonal pattern.

Step 2 – Calculate Seasonal Index

Seasonal Index for day d=Actual value on that dayAverage for that day of week\text{Seasonal Index for day } d = \frac{\text{Actual value on that day}}{\text{Average for that day of week}} Example: a Monday with sales 900 → index =900/843.49≈1.067= 900 / 843.49 \approx 1.067 (above average).

Step 3 – Deseasonalize the Data

Remove seasonality to isolate trend: Deseasonalized value=Actual valueSeasonal Index\text{Deseasonalized value} = \frac{\text{Actual value}}{\text{Seasonal Index}} After deseasonalization, only trend and residual noise remain, enabling clearer trend analysis and better forecasting.

Exam tip: When performing time‑series cross‑validation, always use a time‑ordered train‑test split (e.g., first 80% as train, last 20% as test). Random splitting would leak future information into the training set.

Key takeaways

  • Trend captures long‑term direction; seasonality captures fixed‑interval cycles.
  • Detect trend via visual inspection, moving average, or linear/piecewise trend lines.
  • Detect seasonality by averaging values per period (e.g., day‑of‑week) and computing seasonal indices.
  • Deseasonalization separates the two components for more accurate forecasting.
  • These concepts underpin advanced models (autoregressive, etc.) discussed later in the module.

Simple Forecasting Methods

Time series forecasting starts with methods that are intuitive and easy to apply. These simple methods (naive, seasonal naive, average, weighted average, drift) are often the first step before moving to more sophisticated techniques. Each captures a different assumption about how the past relates to the future.

Naive Method

The naive method assumes the most recent observation is the best forecast for all future periods. Formally, for a time series y1,y2,…,yty_1, y_2, \dots, y_t:

y^t+h=ytfor all h≥1\hat{y}_{t+h} = y_t \quad \text{for all } h \ge 1

Intuition: If nothing is changing rapidly, the last value is a reasonable guess. Works well when data is stable or unpredictable (e.g., daily stock prices with no trend).

Worked example – Retail store daily sales (units):

DayMonTueWedThuFriSatSun
Sales100120115110130120125

Forecast for next Monday (and all future days): 125 units (Sunday’s value). Graphically, a horizontal line at the last observation.

Limitations:

  • Fails when data has strong trends or seasonality.
  • Sensitive to random fluctuations – one outlier can distort all forecasts.

Seasonal Naive Method

The seasonal naive method extends the naive idea by repeating the observation from the same period in the previous season. For a season length mm:

y^t+h=yt+h−m\hat{y}_{t+h} = y_{t + h - m}

Intuition: If a pattern repeats every week (or month, hour, etc.), the future should look like the most recent complete season.

Worked example – Same store, now with two weeks of data:

WeekMonTueWedThuFriSatSun
Week 1100120115110130120125
Week 2210240230220260300340

Forecast for Week 3: copy Week 2 values → Monday = 210, Tuesday = 240, …, Sunday = 340. Graphically, the second week’s pattern repeats indefinitely.

Limitations:

  • Assumes seasonal pattern is constant (no trend, no change in shape).
  • Ignores any overall upward or downward movement across seasons.

Exam tip: Use seasonal naive when you see clear repeating cycles (weekly, monthly) and no trend. For stock market daily data (no strong seasonality), naive often beats seasonal naive.

Average Method

The average method uses the mean of all past observations as the forecast:

y^t+1=1t∑i=1tyi\hat{y}_{t+1} = \frac{1}{t}\sum_{i=1}^{t} y_i

Intuition: When data is stable with no trend or season, the historical average is a reasonable guess.

Worked example – Coffee consumption (cups/day) over one week:

DayMonTueWedThuFriSatSun
Cups4325311

Forecast for next Monday:

y^=4+3+2+5+3+1+17=197≈2.71 cups\hat{y} = \frac{4+3+2+5+3+1+1}{7} = \frac{19}{7} \approx 2.71 \text{ cups}

Compare: naive would give 1 cup (Sunday’s value), seasonal naive would give 4 cups (last Monday’s value). The average balances extremes.

Weighted Average Method

The weighted average method generalises the average by assigning different weights to observations, typically giving more to recent ones:

y^t+1=∑i=1twiyi∑i=1twi\hat{y}_{t+1} = \frac{\sum_{i=1}^{t} w_i y_i}{\sum_{i=1}^{t} w_i}

Weights are chosen by the analyst (e.g., linear increasing, exponential decay).

Intuition: Recent data may be more relevant, especially if trends or patterns shift. Weighted average lets you emphasise the end of the series.

Worked example – Coffee data with weights 1 (Monday) to 7 (Sunday):

DayWeightSalesWeight × Sales
Mon144
Tue236
Wed326
Thu4520
Fri5315
Sat616
Sun717

Sum weights = 1+2+3+4+5+6+7 = 28. Sum weighted sales = 4+6+6+20+15+6+7 = 64.

y^=6428≈2.29 cups\hat{y} = \frac{64}{28} \approx 2.29 \text{ cups}

Because recent days (Sat, Sun) had low values, the forecast is lower than the simple average (2.71). Weighted average is useful when recent trends dominate (e.g., sales during a promotion).

Drift Method

The drift method assumes a constant average change between consecutive observations. The forecast is the last observation plus the average per‑period change:

y^t+h=yt+h⋅yt−y1t−1\hat{y}_{t+h} = y_t + h \cdot \frac{y_t - y_1}{t-1}

Intuition: When data shows a clear upward or downward trend (sales growth, population increase), drift captures that directional movement.

Worked example – Product sales over 7 days:

DayMonTueWedThuFriSatSun
Sales10121517202225

First observation y1=10y_1=10, last yt=25y_t=25, t=7t=7 days of data → t−1=6t-1=6 time steps.

Average change: 25−106=2.5\frac{25-10}{6} = 2.5 units per day.

Forecast for next Monday (h=1h=1): 25+1×2.5=27.525 + 1 \times 2.5 = 27.5 units. Forecast for Tuesday: 25+2×2.5=3025 + 2 \times 2.5 = 30 units.

The forecast line continues the slope observed from the entire history.

Limitations:

  • Assumes linear trend extends forever – unrealistic for long horizons.
  • Sensitive to the first and last observations; outliers can distort the slope.

Choosing Among Simple Methods

MethodBest when…Ignores
NaiveData is stable or random, no trend/seasonalityTrends, seasonality
Seasonal naiveStrong, repeating seasonal pattern (e.g., weekly, monthly)Long-term trends
AverageData is stationary with no trend/seasonRecent changes, trends
Weighted averageRecent observations are more important than older onesSeasonality (unless weighted accordingly)
DriftClear upward or downward linear trendSeasonality, irregular fluctuations

Exam tip: Simple methods are often used as benchmarks. Any advanced model should outperform the naive or seasonal naive forecasts. Many exam questions ask you to compute forecasts by hand using these methods – practice with small datasets.

Key Takeaways

  • Naive: forecast = last observed value. Works for flat, noisy series.
  • Seasonal naive: forecast = value from same period in previous season. Works for repeating patterns.
  • Average: forecast = historical mean. Works for stationary data.
  • Weighted average: forecast = weighted mean with higher weights on recent data. Captures recent shifts.
  • Drift: forecast = last value + linear trend from first to last point. Works for trending data.
  • All five are intuitive, easy to implement, and provide a foundation for advanced methods like exponential smoothing.

Exponential Smoothing Method

Exponential smoothing is a forecasting technique that computes a weighted average of past observations, assigning exponentially decreasing weights so that more recent data influences the forecast more heavily. It smooths out random fluctuations to reveal the underlying trend while still incorporating all historical data.

Intuition

Instead of equal weights (simple average) or arbitrary weights (weighted average), exponential smoothing uses a single parameter – the smoothing factor α\alpha – to control how quickly the forecast adapts to new information.

  • α\alpha close to 1 → forecast reacts strongly to the most recent observation (rapid adaptation).
  • α\alpha close to 0 → forecast is influenced almost equally by all past data (heavy smoothing).

The method gets its name because the weights applied to past observations decay at an exponential rate.

The Core Formula

Let yty_t be the actual observation at time tt, and FtF_t be the forecast for time tt (made at t−1t-1). The forecast for the next period t+1t+1 is:

Ft+1=α yt+(1−α) FtF_{t+1} = \alpha \, y_t + (1 - \alpha) \, F_t

Where:

  • 0≤α≤10 \leq \alpha \leq 1 – smoothing factor.
  • The initial forecast F1F_1 is usually set to the first observation y1y_1 (equivalent to the naive method for the first step).

Why “Exponential”? — Expanding the Formula

Repeatedly substituting Ft,Ft−1,…F_t, F_{t-1}, \dots reveals the weight structure:

Ft+1=αyt+(1−α)Ft=αyt+α(1−α)yt−1+(1−α)2Ft−1=αyt+α(1−α)yt−1+α(1−α)2yt−2+⋯+(1−α)tF1\begin{aligned} F_{t+1} &= \alpha y_t + (1-\alpha)F_t \\ &= \alpha y_t + \alpha(1-\alpha) y_{t-1} + (1-\alpha)^2 F_{t-1} \\ &= \alpha y_t + \alpha(1-\alpha) y_{t-1} + \alpha(1-\alpha)^2 y_{t-2} + \cdots + (1-\alpha)^t F_1 \end{aligned}

The weight on observation yt−ky_{t-k} is:

α(1−α)k\alpha (1-\alpha)^k

Because (1−α)<1(1-\alpha) < 1, the weight decreases geometrically (exponentially) as kk increases — older data matters less and less.

Worked Example: Ice Cream Shop Daily Sales

DayActual Sales yty_tForecast FtF_t (made previous day)Calculation for next forecast
Mon100–Set FTue=100F_{\text{Tue}} = 100 (initial naive)
Tue120100FWed=0.5×120+0.5×100=110F_{\text{Wed}} = 0.5 \times 120 + 0.5 \times 100 = 110
Wed115110FThu=0.5×115+0.5×110=112.5F_{\text{Thu}} = 0.5 \times 115 + 0.5 \times 110 = 112.5
Thu130112.5FFri=0.5×130+0.5×112.5=121.25F_{\text{Fri}} = 0.5 \times 130 + 0.5 \times 112.5 = 121.25
Fri(unknown)121.25FSat=0.5×(Fri actual)+0.5×121.25F_{\text{Sat}} = 0.5 \times \text{(Fri actual)} + 0.5 \times 121.25
Sat………
Sun125(computed)FMon≈0.5×125+0.5×(Sun forecast)≈127F_{\text{Mon}} \approx 0.5 \times 125 + 0.5 \times \text{(Sun forecast)} \approx 127

Result: Forecast for next Monday ≈ 127 units (with α=0.5\alpha = 0.5).

The Damping Factor

In many software implementations (e.g., Excel), the parameter requested is the damping factor 1−α1 - \alpha rather than α\alpha itself.

Exam tip: Always check which parameter a tool asks for. If it asks for “damping factor”, enter 1−α1-\alpha. A damping factor of 0.8 means α=0.2\alpha = 0.2 (very heavy smoothing).

Why Use Exponential Smoothing?

ReasonExplanation
AdaptabilityHigher α\alpha makes the forecast respond quickly to changes (e.g., a hot day spikes sales).
SimplicityOnly requires α\alpha and the last actual + forecast; can be calculated by hand or in a spreadsheet.
SmoothnessWeights down extreme fluctuations, giving a clear view of the underlying trend and seasonality.

Implementation in Excel

  1. Open Data Analysis Toolpak → select Exponential Smoothing.
  2. Input range: column of actual sales.
  3. Damping factor: enter 1−α1-\alpha (e.g., 0.2 for α=0.8\alpha=0.8).
  4. Set output range → click OK.

The resulting smoothed series tracks the data while reducing noise. A higher damping factor (lower α\alpha) produces a smoother line.

Visualising the Forecast Process


Key takeaways

  • Exponential smoothing is a weighted average with weights decaying exponentially (α(1−α)k\alpha(1-\alpha)^k).
  • Smoothing factor α\alpha (0 ≤ α ≤ 1) controls responsiveness: high α → fast adaptation, low α → heavy smoothing.
  • The damping factor = 1−α1-\alpha; Excel uses damping factor, not α.
  • Initial forecast is usually the first observation (naive).
  • Three strengths: adaptability, simplicity, smoothness.
  • The method produces forecasts that react quickly to recent changes while still using all past data.

Autocorrelation and Autoregressive Models

Autocorrelation (also called lagged correlation or serial correlation) measures how a time series is correlated with its own past values. Intuitively: does today’s number carry information about tomorrow’s? In most business data – stock prices, daily sales, temperature – the past does influence the future. Autocorrelation quantifies that influence.

Mathematical definition

For a time series y1,y2,…,yny_1, y_2, \dots, y_n, the autocorrelation at lag kk is the Pearson correlation between the series and itself shifted by kk time steps:

ρk=∑t=k+1n(yt−yˉ)(yt−k−yˉ)∑t=k+1n(yt−yˉ)2∑t=k+1n(yt−k−yˉ)2\rho_k = \frac{\sum_{t=k+1}^{n} (y_t - \bar{y})(y_{t-k} - \bar{y})}{\sqrt{\sum_{t=k+1}^{n} (y_t - \bar{y})^2 \sum_{t=k+1}^{n} (y_{t-k} - \bar{y})^2}}

In practice, to compute ρ1\rho_1:

  • Create two series: (y2,y3,…,yn)(y_2, y_3, \dots, y_n) and (y1,y2,…,yn−1)(y_1, y_2, \dots, y_{n-1}).
  • Compute their correlation (e.g., CORREL() in Excel).

Worked example – toy sales data

Given daily sales: values from 100 to 130, ending at 125 on Sunday (7 days). For lag 1, align:

OriginalLagged 1
y2=120y_2 = 120y1=100y_1 = 100
y3=115y_3 = 115y2=120y_2 = 120
……
y7=125y_7 = 125y6y_6

The correlation of these two series equals 0.196. Interpretation: a 1‑unit increase in past day’s sales is associated with a 0.2‑unit increase in the next day’s sales (linear approximation).

Extending to higher lags

  • Lag 2: correlate (y3,y4,…,yn)(y_3, y_4, \dots, y_n) with (y1,y2,…,yn−2)(y_1, y_2, \dots, y_{n-2}).
  • Lag kk: correlate (yk+1,…,yn)(y_{k+1}, \dots, y_n) with (y1,…,yn−k)(y_1, \dots, y_{n-k}).

Typically, autocorrelation decreases as the lag increases (e.g., today’s stock price influences tomorrow most, the day after a little less, etc.).

Sign of autocorrelation – what it tells you

Autocorrelation signPatternBusiness example
PositiveHigh values followed by high; low followed by lowDaily store sales (recurring customer behaviour); warm days follow warm days
NegativeHigh values followed by low; low followed by highAlternating marketing expenditure (high spend → budget cut next period)

Exam tip: Positive autocorrelation is far more common in business. Negative autocorrelation often indicates deliberate alternating policies.


Autoregressive (AR) Models

An autoregressive model exploits autocorrelation by modelling the current value as a function of its own past values. The name literally means “regressing on itself”.

AR of order 1 – AR(1)

yt=c+ϕ1yt−1+εty_t = c + \phi_1 y_{t-1} + \varepsilon_t
  • yty_t = current value
  • yt−1y_{t-1} = value at previous time point
  • ϕ1\phi_1 = coefficient (strength of influence)
  • cc = intercept
  • εt\varepsilon_t = random error

Creating the lagged predictor: shift the series by one time step. Then run ordinary linear regression with yty_t as response and yt−1y_{t-1} as predictor.

Bakery sales example

Data: daily sales from 2 Jan 2021 to 30 Sep 2022. For an AR(1) model:

  • Dependent variable: sales on day tt
  • Independent variable: sales on day t−1t-1 (lagged series)

Regression output (Excel Data Analysis ToolPak):

CoefficientEstimatepp-value
Intercept (cc)242.9≈ 0
Lagged sales (ϕ1\phi_1)0.739≈ 0
  • Adjusted R2R^2: ~54% – past day’s sales explains 54% of today’s variation.
  • Both coefficients are statistically significant (p≈0p \approx 0).

Interpretation:

  • If previous day’s sales were zero, expected sales = 242.9 units.
  • For every 1‑unit increase in past sales, today’s sales increase by 0.739 units.

Prediction: Fitted equation:

y^t+1=242.942+0.739 yt\hat{y}_{t+1} = 242.942 + 0.739 \, y_t

On 30 Sep 2022, sales = 795.95 units. Prediction for 1 Oct 2022:

y^=242.942+0.739×795.95=831.15 units\hat{y} = 242.942 + 0.739 \times 795.95 = 831.15 \text{ units}

AR of order 2 – AR(2)

yt=c+ϕ1yt−1+ϕ2yt−2+εty_t = c + \phi_1 y_{t-1} + \phi_2 y_{t-2} + \varepsilon_t

Extends to include two previous lags. The appropriate order is chosen by examining autocorrelation at multiple lags or using cross‑validation.

Exam tip: For an AR model to work, significant autocorrelation must exist in the data. Check the autocorrelation plot before fitting.


Choosing among forecasting methods – a recap

In practice, evaluate AR models and alternatives on a held‑out test set (the last few observations) using root mean squared error (RMSE).

MethodCore idea
NaïveUse last observed value for all future forecasts
Seasonal naïveUse last observed value from same season (e.g., last Monday for next Monday)
Average methodForecast = sample mean of all past values
Weighted averageWeighted mean of all past values, higher weight on recent
Exponential smoothingWeights decay exponentially: α(1−α)k\alpha (1-\alpha)^{k} for values kk periods back
Drift methodLinear trend from first to last observation, projected forward
Autoregressive (AR)Linear regression on own past values

Cross‑validation procedure for time series:

  1. Hold out the last mm observations as the test set (e.g., last month).
  2. Train each method on the remaining data.
  3. Compute RMSE on the test set.
  4. Select the method with the smallest RMSE.

Key takeaways

  • Autocorrelation measures how a time series correlates with its own past; it is the foundation of autoregressive models.
  • ρk\rho_k is computed by correlating the series with its kk-lagged version; typically decays with kk.
  • Positive autocorrelation: high → high; negative: high → low (rarer in business).
  • AR(1): yt=c+ϕ1yt−1+εty_t = c + \phi_1 y_{t-1} + \varepsilon_t; estimated via linear regression on the lagged variable.
  • AR models require significant autocorrelation; order selection can use autocorrelation analysis or cross‑validation.
  • Always compare AR with simpler benchmarks (naïve, exponential smoothing, drift) using a hold‑out RMSE.

A Case Study: Forecasting Gold Price in India (1964–2023)

This case study applies forecasting techniques to annual gold price data (1964–2023, 60 years). The dataset includes inflation as a potential predictor. The goal: predict gold price for the next year and evaluate which method works best. A cross-validation approach holds out the last 5 years (2019–2023) as a test set; the first 55 years train the models.

Questions and Data Overview

Six key questions guide the analysis:

  1. Is there a trend in gold price?
  2. Is there seasonality in this yearly series?
  3. Can past inflation (lagged) serve as a predictor in a linear regression model?
  4. Which simple methods (Naive, Seasonal Naive, Average, Weighted Average, Drift) perform well?
  5. What is the autocorrelation structure? Would an autoregressive model be suitable?
  6. Using cross-validation, which of the candidate models forecasts best?

1. Trend and Seasonality

Intuition: Visualizing the time series reveals overall behavior and whether trend or repeating cycles exist.

  • The gold price plot shows a slow linear increase until ~2003, a dip around 2000, then exponential growth, another dip around 2012, followed by renewed exponential rise. The trend is erratic – not consistently linear or exponential – making a simple linear trend unsuitable.
  • Seasonality is absent. Because the data is annual (yearly), repeating patterns within a year cannot exist. Any method relying on seasonality (e.g., Seasonal Naive, seasonal AR) is inappropriate.

Exam tip: For annual data, seasonality is impossible unless the series is recorded at sub‑annual intervals. Always check the frequency first.

2. Linear Regression with Inflation as Predictor

Intuition: Inflation is commonly believed to drive gold prices. A regression model uses lagged inflation (inflation from the previous year) to predict gold price.

  • Model: Goldt=β0+β1⋅Inflationt−1+ε\text{Gold}_t = \beta_0 + \beta_1 \cdot \text{Inflation}_{t-1} + \varepsilon
  • Fitted equation (training data 1964–2018): Predicted Gold=8191.447−(coefficient)×Inflation\text{Predicted Gold} = 8191.447 - \text{(coefficient)} \times \text{Inflation}
  • The coefficient on inflation is not statistically significant (high p‑value). Although the common belief links gold to inflation, over the 60‑year period inflation changed little while gold prices soared. Inflation alone is a weak predictor.
  • Predictions for 2019–2023 are poor. For 2019, predicted ≈ 7,775 vs. actual ≈ 35,000. RMSE is extremely high (similar to the simple average method).

Takeaway: Simple linear regression on a single, non‑trending predictor fails when the target has a strong time‑dependent growth.

3. Simple Forecasting Methods

Five simple methods were considered:

  • Naive method: Forecast next year = last observed value. (Good benchmark, always included.)
  • Seasonal naive: Requires seasonality – not applicable.
  • Average method: Forecast = mean of all training data (55 years). Unlikely to capture rising trend.
  • Weighted average method: Gives more weight to recent observations. Here, weights proportional to 1,2,3,4,5 for the last 5 years (most recent gets weight 5). Subjective – other schemes possible.
  • Drift method: Assumes constant linear change per period; not suitable because trend is not linear over the whole period.

Decision: Only Naive, Average, and Weighted Average are used. Seasonal naive and drift are dropped.

Forecasts on test set (2019–2023):

YearActual Gold PriceNaive (31,438)Average (7,223)Weighted Avg
2019~35,00031,4387,223~23,000 (est)
2020...31,4387,223increasing
...............
2023...31,4387,223...

RMSE results:

  • Average: ~43,567 (very poor)
  • Naive: 20,536
  • Weighted average: 12,869

Weighted average captures some upward trend by heavily weighting recent data, outperforming naive. Average method fails completely.

4. Autoregressive Model of Order 1 (AR1)

Intuition: Gold price is highly correlated with its own past. Autocorrelation measures this: a high, slowly decaying autocorrelation suggests an autoregressive (AR) model will work.

Autocorrelation values (training data):

LagAutocorrelation
10.989
20.964
30.935
40.908

Values are very high and decrease with lag – the farther back, the less influence on the present. This pattern supports an AR model.

AR1 model fitted:

Goldt=263.282+1.047⋅Goldt−1+ε\text{Gold}_t = 263.282 + 1.047 \cdot \text{Gold}_{t-1} + \varepsilon

  • Coefficient of lagged gold is highly significant. For every unit increase in last year’s price, this year’s price increases by ~1.047 units (plus intercept).
  • Forecasting for 2019: Use gold price in 2018 (31,438). Prediction: 263.282+1.047×31,438≈33,176263.282 + 1.047 \times 31,438 \approx 33,176. Actual: 35,000. Error ~1,824.
  • Repeating for all 5 test years yields RMSE = 6,568 – far lower than any simple method.

Why it works: The AR1 model captures the persistence and growth inherent in the series without assuming a fixed trend shape.

5. Cross-Validation Results Summary

MethodRMSE (test set)
Linear regression (inflation)~43,500 (approx)
Average method43,567
Naive method20,536
Weighted average (last 5, weights 1-5)12,869
AR1 model6,568

The AR1 model reduces RMSE by about 49% relative to the next-best weighted-average method. Further refinements include:

  • Combining regression with AR (ARIMAX).
  • Higher‑order AR (AR2, AR3).
  • Adding trend terms (linear, exponential) as additional predictors in an AR framework.

6. Beyond: Improving the Model

The case study hints at more advanced approaches. For gold price, one could add a linear trend (time index 1,2,3,…) and an exponential trend (e1,e2,e3,…e^1, e^2, e^3,\dots) as extra regressors alongside the lagged gold. Running a regression with these would likely improve predictions further.

Contrast with bakery sales data (from earlier module): That data had seasonality, so seasonal naive and seasonal AR would be appropriate. This case shows that method selection must always match the data’s characteristics.

Exam tip: Always start with a naive forecast as a baseline. Any serious model must beat its RMSE. Here, AR1 beat it by a factor of 3.

Key Takeaways

  • Annual observations cannot reveal within-year seasonality – use seasonal methods only when the sampling frequency captures a repeating seasonal cycle.
  • Visual inspection reveals trend shape (here: erratic, partly exponential) and zero seasonality.
  • Simple linear regression on a single predictor (inflation) fails when the predictor lacks a time trend of its own.
  • Among simple methods, weighted average (giving more weight to recent data) captured growth and outperformed naive.
  • Autocorrelation was very high (0.99 at lag 1), making an AR1 model the best performer (RMSE 6,568 vs. 12,869 for weighted average).
  • Cross-validation (hold‑out last 5 years) provided an honest comparison – AR1 was clearly superior.
  • Always refine models by testing additional predictors (trend terms, other economic variables) and higher AR orders.