Term 1 · Module 1 of 8

Descriptive Statistics

Business Statistics for Entrepreneurs

Descriptive Statistics: Data Types & Terminology

Descriptive statistics is the process of describing collected data through visuals and numerical summaries. It helps turn raw facts into a story for decision-making.

What is "Statistics"?

The word can mean either:

  • Numerical facts computed from data (e.g., average, median, range).
  • The process of collecting, analysing, presenting, and interpreting data — a blend of science (fixed principles) and art (subtle interpretive nuance).

Statistics is used across business functions: accounting/finance (performance analysis), economics (policy), marketing/production (grow profit, reduce cost), and information systems (system health).

Data and Data Set

  • Data: Facts and figures collected, analysed, and summarised for decision-making.
  • Data set: The entire collection of data for a specific study.

Example: Manjula Nayak's Bengaluru Household Survey An excerpt with 9 households and multiple variables:

Family SizeLocation (numeric code)Location (text)Ownership (0/1)Income 1 (₹)Income 2 (₹)Monthly Payment (₹)Utilities (₹)Debt (₹)Satisfaction (numeric)Satisfaction (Likert)
32Northwest11,700,000012,0003,00050,0004Satisfied
  • Element: The entity for which data are collected (here: each household).
  • Variable: A characteristic of interest for each element (e.g., family size, location, income).
  • Observation: The complete set of measurements for one element — one entire row.

⚠️ Quiz trap — In a different dataset (Tumkur apartment rents), a table with 7 rows and 10 columns contained 70 elements (households), each with only one variable: monthly rent. Observations were single-cell rents, not whole rows. Never assume rows = observations and columns = variables. Check what each column represents.

Types of Data

Data can be broadly qualitative or quantitative, each with finer subcategories. The type determines how you visualise and analyse it.

Qualitative Data (Categorical)

Records attributes or labels; can be numeric or non-numeric. Divides into:

ScalePropertyExample from survey
NominalLabels/names, no rankingLocation (Southwest=1, Northwest=2, Northeast=3, Southeast=4) — numeric but no order; Ownership (0=rent, 1=own)
OrdinalAll nominal properties plus inherent rankingSatisfaction (5=Extremely satisfied … 1=Extremely dissatisfied) — both non-numeric and numeric forms imply rank

Exam tip: Numbers alone don't make data quantitative. Location codes (1,2,3,4) are still qualitative because they merely label — there is no meaningful ordering.

Quantitative Data (Numerical)

Values have magnitude and order. Two sub-scales:

ScaleMeaningful ratio?Zero meaningful?Examples
Interval❌ No (10°C is not “twice as hot” as 5°C)❌ No (0°C does not mean “no temperature”)Temperature (°C), clock time (2 AM, 4 AM) — intervals matter, not ratios
Ratio✅ Yes (₹2 Lakh profit is twice ₹1 Lakh)✅ Yes (₹0 profit = break-even)Income, debt, age, profit — both ratios and zero are meaningful

Worked comparisons:

  • Ownership (0/1) → qualitative, nominal (no ranking: 1 is not “better” than 0).
  • Satisfaction numeric (1–5) → qualitative, ordinal (5 is better than 1).
  • Income (₹) → quantitative, ratio (₹0 = no income; ₹2L/₹1L = twice).
  • Temperature (Celsius) → quantitative, interval (0°C is arbitrary; 30°C/10°C ≠ "three times hotter").

Key Takeaways

  • Descriptive statistics involves summarising data via visuals and numerical summaries.
  • An observation is the full set of variables for one element, not necessarily one row — always check the data structure.
  • Qualitative data can be numeric or non-numeric; its scale is nominal (no rank) or ordinal (rank implied).
  • Quantitative data is always numeric; its scale is interval (ratio meaningless, zero arbitrary) or ratio (ratio meaningful, zero absolute).
  • The variable type (nominal, ordinal, interval, ratio) determines which visualisation and analysis techniques are appropriate.

Core Classification Framework

Every variable can be classified along three binary axes:

  1. Qualitative (categorical) vs Quantitative (numerical measurement)
  2. Numeric (values are numbers) vs Non‑numeric (labels are text)
  3. Nominal (no order) vs Ordinal (natural order) – for qualitative variables Interval (equal intervals, no true zero) vs Ratio (true zero, ratios meaningful) – for quantitative variables

Key insight: A qualitative variable can be numeric (e.g., a 1‑5 rating) – the numbers act as category labels, not measurements. Likewise, a quantitative variable is always numeric.


1. Alumni Survey Data (Jyothi Hegde — IIIT Dharwad)

100+ rows; variables describe student profiles, test scores, and outcomes.

VariableTypeNumeric?Sub‑typeRationale
Gender (M/W)QualitativeNon‑numericNominalLabels only, no order
Age (years)QuantitativeNumericRatioTrue zero (age 0), ratios meaningful
Overall (0‑100)QuantitativeNumericRatioScale includes 0, zero has meaning
Core courses (2‑10 GPA)QuantitativeNumericIntervalScale starts at 2, not 0; ratio not valid
Elective (2‑10 GPA)QuantitativeNumericIntervalSame as core
Work experience (years)QuantitativeNumericRatioTrue zero
Domicile (1=in‑state, 2=out)QualitativeNumericNominalNumbers are codes, no ordering
Rating (1‑5)QualitativeNumericOrdinalNumbers imply rank (5 > 1)
Satisfaction (labels)QualitativeNon‑numericOrdinal“Very satisfied” > “Not at all satisfied”
Salary (lakhs)QuantitativeNumericRatioTrue zero, ratios meaningful

Take‑away for exam:

  • A numeric variable is not automatically quantitative – e.g., rating and domicile are qualitative despite being numbers.
  • Test scores with a non‑zero minimum (e.g., 2‑10) are interval, not ratio, because a score of 4 is not “twice as good” as 2.

2. Credit Card Spending (Hanumanthappa – Canara Bank, Mangaluru)

55 customers; variables: annual income, household size, amount charged.

VariableTypeNumeric?Sub‑type
Annual income (₹ lakhs)QuantitativeNumericRatio
Household size (# persons)QuantitativeNumericRatio
Amount charged (₹/month)QuantitativeNumericRatio

All three are quantitative, numeric, ratio – straightforward.


3. Chroma Customer Survey (Basavaraja – Bellary)

36 responses; includes computer attributes and satisfaction score.

VariableTypeNumeric?Sub‑typeRationale
Brand (e.g., Lenovo, Dell)QualitativeNon‑numericNominalNo order
Price (₹ lakhs)QuantitativeNumericRatioTrue zero
RAM (GB)QualitativeNumericOrdinal32 GB > 16 GB in performance, but not a measurement of “RAM quantity” – it’s a category with order
Hard disk (GB)QualitativeNumericOrdinalSame logic as RAM
Operating system (Win10/11, Home/Pro)QualitativeNon‑numericNominalLater version ≠ better (e.g., Win11 not inherently superior); no clear order
Satisfaction score (0‑100)QuantitativeNumericRatioTrue zero, ratios meaningful

Exam trap: RAM and hard disk are trick variables – they look quantitative but are treated as ordinal categorical because the values are arbitrary tiers, not continuous measurements. Never assume numeric ⇒ quantitative.


4. Nandini Sweets Marketing (Umesh Kamath – Udupi)

Weekly store data; variables: date, store, sales, price, promotions (0/1).

VariableTypeNumeric?Sub‑typeRationale
Date (week)QuantitativeNumericIntervalZero date is arbitrary; ratio of dates meaningless
Store name (text)QualitativeNon‑numericNominalNo order
Units soldQuantitativeNumericRatioTrue zero
Retail price (₹)QuantitativeNumericRatioTrue zero
Coupons (0/1)QualitativeNumericNominal0/1 are codes, no order
Displays (0/1)QualitativeNumericNominalSame
Sampling (0/1)QualitativeNumericNominalSame

Key point: Binary (0/1) variables are qualitative nominal – the numbers are labels, not measurements.


5. Magic Bricks Property Data (Vijay Gowda – Mysuru)

30 apartments; variables: #bathrooms, #bedrooms, premium locality (0/1), gated complex (0/1), amenities (0/1), square footage, selling price.

VariableTypeNumeric?Sub‑type
# bathroomsQuantitativeNumericRatio
# bedroomsQuantitativeNumericRatio
Premium locality (0/1)QualitativeNumericNominal
Gated complex (0/1)QualitativeNumericNominal
Amenities (0/1)QualitativeNumericNominal
Square footageQuantitativeNumericRatio
Selling price (₹)QuantitativeNumericRatio

All quantitative variables are ratio; all qualitative variables are numeric nominal (0/1 codes).


Key Takeaways

  • Qualitative variables classify; they can be numeric (e.g., rating 1‑5) or non‑numeric. Sub‑types: nominal (no order) or ordinal (ranked).
  • Quantitative variables measure; always numeric. Sub‑types: interval (no true zero, e.g., date, test scores starting >0) or ratio (true zero, e.g., income, age, sales).
  • Binary (0/1) variables are qualitative nominal – the numbers are codes, not counts.
  • Numeric does not imply quantitative – always check whether the variable is a measurement or a label.
  • Common exam traps:
    • Test scores with non‑zero minimum → interval (not ratio).
    • RAM/HDD sizes in tiers → ordinal qualitative (not quantitative).
    • Later version of software → not automatically ordinal; judge whether order is meaningful.

Types of Data: Cross-Sectional vs. Time Series

Distinguishing between cross-sectional data and time series data affects how you analyse trends, correlations, and cause‑effect relationships.

  • Cross-sectional data – collected at the same (or approximately the same) point in time across multiple subjects (e.g., households, districts, individuals).
  • Time series data – collected over several time periods for one or more subjects (e.g., months, years).

Key intuition

  • Cross‑sectional: a snapshot across many units at one time.
  • Time series: a movie of one (or several) units across time.

Examples

Study / DatasetData collectedClassificationReason
Manjula’s household survey – rent/own, income, monthly paymentAt one point in time across many householdsCross-sectionalSnapshot across households
Jyoti Hegde – placement/satisfaction at IIIT Dharwad (gender, age, exam scores, work experience, salary)Surveyed current people at one time, though they graduated at different timesCross-sectionalSurvey conducted at a single point in time
Hanumanth Pai – credit card data (income, household size, monthly charges)One‑time data on many customersCross-sectionalSnapshot across customers
Basavaraju – customer loyalty at Chroma (computer purchase, brand, price, features)Survey done all at onceCross-sectionalSnapshot across customers
Umesh Kamath – Nandini Sweets (marketing promotions, units sold, selling price, promotion strategy over multiple periods)Data collected over several time periods for stores in Belgaum and GulbargaTime seriesRepeated measurements over time
Vijay Gowda – real estate (bathrooms, bedrooms, locality, amenities, square footage, selling price)Data collected all at onceCross-sectionalSnapshot of flats at one time

Exam tip: The “snapshot across subjects” vs. “repeated measures over time” distinction is often tested. Even if subjects graduated in different years, if you survey them now, it’s cross‑sectional. The key is the time of data collection, not the time period the data describe.


Data Collection Methods

When data are not available from existing sources, they must be collected. Three main approaches:

1. Existing (Secondary) Sources

Data already exist and can be used without separate collection.

  • Examples: Company records (operations, personnel, finances); government data (census); specialized organisations (Bloomberg, Nielsen Company).
  • Advantage: Fast, inexpensive, often large scale.

2. Observational Studies

The researcher observes and records variables of interest without intervening.

  • When to use: When existing sources are insufficient.
  • Process: Simply watch and measure what happens in a real‑world setting.
  • Examples: Observers in a store noting customer choices; HR tracking employee attrition over time.

3. Experimental Studies

Data are collected in a controlled manner – the researcher decides the target group, specific situations/conditions, and what data to collect.

  • When to use: To isolate cause‑and‑effect (often via controlled manipulation).
  • Examples: Clinical trials for pharmaceuticals (target group identified, process specified, predefined data types).

Key contrast: In observational studies you record what naturally occurs; in experimental studies you control conditions to test a hypothesis.


Key takeaways – Types of data

  • Cross‑sectional: one‑time snapshot across subjects → compare differences.
  • Time series: repeated observations over time → track trends and patterns.
  • Classification depends on when the data are collected, not what they describe.

Key takeaways – Data collection

  • Existing sources are the cheapest route, but may lack the variables you need.
  • Observational studies record natural behaviour (no intervention).
  • Experimental studies control conditions to establish causal relationships.

Data Visualization

Data visualization uses tables and graphs to summarize categorical and quantitative data. The goal is not just to create them but to choose the right format for the context and interpret it correctly.

Tabular Summaries for Categorical Data

A frequency distribution is a table that shows the count (frequency) of observations in each of several non‑overlapping categories.

Example: Satisfaction levels (5 levels: Very Satisfied, Satisfied, Neutral, Unsatisfied, Not at all Satisfied) for IIIT Dharwad respondents (103 total: 72 men, 31 women).

Satisfaction LevelMen (freq)Women (freq)Overall (freq)
Very Satisfied17926
Satisfied22426
Neutral22628
Unsatisfied9716
Not at all Satisfied257

A relative frequency is the proportion of observations in a class:

Relative frequency=Frequency of classn\text{Relative frequency} = \frac{\text{Frequency of class}}{n}

The percent frequency is the relative frequency multiplied by 100.

Worked example – Men:

  • Very Satisfied: 17/72≈0.24→24%17 / 72 \approx 0.24 \rightarrow 24\%
  • Satisfied: 22/72≈0.31→31%22 / 72 \approx 0.31 \rightarrow 31\%
  • Neutral: 22/72≈0.31→31%22 / 72 \approx 0.31 \rightarrow 31\%
  • Unsatisfied: 9/72≈0.125→12.5%9 / 72 \approx 0.125 \rightarrow 12.5\%
  • Not at all Satisfied: 2/72≈0.028→2.8%2 / 72 \approx 0.028 \rightarrow 2.8\%

Women example: Very Satisfied: 9/31≈0.29→29%9 / 31 \approx 0.29 \rightarrow 29\%.

A relative frequency distribution and percent frequency distribution are tables of these values.

Graphical Displays for Categorical Data

  • Bar chart – X‑axis: categories; Y‑axis: frequency or percent. Height of bar proportional to count/percent.

  • Pie chart – Circle divided into sectors; sector angle = relative frequency × 360°.

    Example: For men very satisfied (24%): sector angle =0.24×360∘=86.4∘= 0.24 \times 360^\circ = 86.4^\circ.

Exam tip: Pie charts are visually harder to compare than bar charts. Most data‑visualisation experts prefer bar charts for categorical data.

Comparing Two Categorical Variables

  • Side‑by‑side bar chart – Two or more bars per category (e.g., men vs. women for each satisfaction level). Orientation can be swapped (categories vs. groups on x‑axis).
  • Stacked bar chart – One bar per category, segmented by the second variable. Height of each segment = frequency.
  • Stacked percentage bar chart – Each bar scaled to 100%; segments show the percentage within each category. Use when absolute counts don’t matter, only proportions.

Choose the format based on the comparison you want to highlight.

Tabular Summaries for Quantitative Data

Frequency, relative frequency, and percent frequency distributions also work for quantitative data, but must define non‑overlapping classes with equal width and 5–10 classes. Each data point falls into exactly one class.

Example – Salary (in ₹ Lakhs) using 7 classes of width 5:

Salary ClassMen (freq)Women (freq)Overall (freq)
15–19.........
20–24.........
25–29.........
30–34.........
35–39.........
40–44.........
45–50.........

Guidelines for class intervals:

  • Same width for all classes.
  • Typically 5–10 classes.
  • Limits chosen so each observation belongs to exactly one class.

Relative frequency and percent frequency are computed the same way as for categorical data.

Graphical Displays for Quantitative Data

  • Histogram – Like a bar chart but rectangles touch (no gap). X‑axis: classes; Y‑axis: frequency. Reveals the shape of the distribution (e.g., skewness).

  • Scatter plot – Shows relationship between two quantitative variables. Each point = one observation (x‑coordinate for one variable, y‑coordinate for the other). A trend line may be added to suggest a pattern.

    Example: Plot salary vs. age for men and women from the survey.

Choosing the Right Display

PurposeRecommended display
Show distribution of a single variable (categorical or quantitative)Bar chart, pie chart (categorical); histogram (quantitative)
Compare two categorical variablesSide‑by‑side bar chart, stacked bar chart, stacked percentage bar chart
Show relationship between two quantitative variablesScatter plot (optionally with trend line)

Data Dashboards

In practice, businesses use data dashboards – a collection of visual displays that organise key metrics for monitoring performance and supporting quick decisions.

Key takeaways

  • Frequency distribution: count per non‑overlapping class; relative frequency = frequency / n; percent = ×100.
  • Bar charts preferred over pie charts for comparing categories.
  • Side‑by‑side and stacked bar charts compare two variables.
  • Histograms (touching rectangles) reveal distribution shape.
  • Scatter plots show relationships between quantitative variables.
  • Class intervals for quantitative data should be equal‑width, 5–10 classes, non‑overlapping.

Measures of Central Tendency

Numerical summaries for data fall into two categories: measures of central tendency (where the data cluster) and measures of dispersion (how spread out they are). Central tendency includes the mean, median, mode, percentiles, and quartiles.

Mean (Arithmetic Average)

The mean is the average value of a variable – the most common measure of central location. For a sample it is denoted xˉ\bar{x} (x‑bar); for a population it is denoted μ\mu. The sample mean is a point estimator of the population mean.

xˉ=1n∑i=1nxi\bar{x} = \frac{1}{n} \sum_{i=1}^{n} x_i

Excel: =AVERAGE(range)

Property: The mean is sensitive to extreme values (outliers). A single very large or very small observation can pull the mean away from the centre.


Median

The median is the middle value when the data are sorted in ascending order.

  • If nn is odd: the median is the (n+1)/2(n+1)/2th observation.
  • If nn is even: the median is the average of the two middle observations.

Robustness: The median is not affected by extreme values.

Example Dataset: 1, 1, 2, 2, 4 → mean = 2, median = 2. Change the last value to 394: mean = 80, median = 2. The median remains unchanged.

Because of this, the median is the preferred measure for skewed distributions such as annual income or property prices.

Excel: =MEDIAN(range)


Mode

The mode is the value that occurs with the greatest frequency.

  • Unimodal – exactly one mode.
  • Bimodal – two values tied for highest frequency (e.g., 1,1,2,2,4 has modes 1 and 2).
  • Multimodal – more than two modes; rarely reported because listing many values is unhelpful.

If every value is unique, no mode exists.

Excel:

  • =MODE.SNGL(range) – returns the first most frequent value.
  • =MODE.MULT(range) – returns an array of all modes.
  • If no mode, Excel returns #N/A.

Percentiles

The ppth percentile divides the data so that approximately p%p\% of observations are less than or equal to that value, and (100−p)%(100-p)\% are greater.

Computation (the method used by PERCENTILE.EXC):

  1. Sort data in ascending order.
  2. Compute index Lp=p100(n+1)L_p = \frac{p}{100}(n+1)
  3. If LpL_p is an integer, take the LpL_pth observation.
  4. If LpL_p is not an integer, interpolate between the floor and ceiling observations.

Example: 80th percentile with n=200n=200 → L80=0.8×201=160.8L_{80}=0.8\times201=160.8 → interpolate between the 160th and 161st values.

Excel: =PERCENTILE.EXC(array, k) where k=p/100k = p/100.

Interpretation: A test score at the 80th percentile means about 80% of test‑takers scored lower (and 20% higher).


Quartiles

Quartiles divide the data into four parts, each containing roughly 25% of observations. They are just specific percentiles:

QuartilePercentileAlso known as
Q1 (first quartile)25thLower quartile
Q2 (second quartile)50thMedian
Q3 (third quartile)75thUpper quartile

Excel: Use =PERCENTILE.EXC with k=0.25,0.5,0.75k = 0.25, 0.5, 0.75.


Applied Examples

The following five datasets illustrate how measures of central tendency behave in practice.

1. IIIT Dharwad – salaries of respondents (n=103n=103)

MeasureValue (₹ lakhs)
Mean25.9
Median25
Mode25
Q124
Q2 (median)25
Q327
80th percentile27
90th percentile29

Interpretation: Mean and median are close – no extreme salaries pulling the average.

2. Canara Bank – monthly credit card spend (n=55n=55)

MeasureValue (₹)
Mean3,852
Median3,864
Mode3,657
Q12,890
Q34,531

Mean and median are similar (no extreme spenders), but mode differs – the most common spend amount is lower.

3. Chroma – customer satisfaction score (n=36n=36)

MeasureValue
Mean57.31
Median59
Q153
Q361.75

Median slightly higher than mean; moderate spread.

4. Nandini sweets – weekly sales (n=132n=132)

MeasureValue (₹)
Mean135.98
Median105
Mode91
Q182.25
Q3136.5

Mean (≈136\approx 136) is much higher than the median (105105), suggesting that a few weeks with very high sales pull the average up.

5. Jayalakshmipuram – property prices (nn not stated, but data given)

MeasureValue (₹ lakhs)
Mean202.57
Median195
Q1149.5
Q3243.75

Mean > median, indicating some high‑priced properties inflating the average – a common pattern in real estate.

Exam tip: The median is the better measure of “typical” value when data are skewed or contain outliers (e.g., income, house prices). Always check whether mean and median are close – if not, the mean is likely distorted.

Key takeaways

  • Mean = sum ÷ nn, sensitive to outliers; median = middle value, robust; mode = most frequent.
  • Percentiles and quartiles describe the distribution’s spread. Q2 = median.
  • Excel commands: AVERAGE, MEDIAN, MODE.SNGL/MODE.MULT, PERCENTILE.EXC.
  • When mean and median differ substantially, suspect extreme values in the tail.

Measures of Dispersion

Measures of dispersion (or variability) quantify how spread out a dataset is. Intuition: two datasets can have the same mean but very different spreads — think of salaries clustered around ₹25 L vs. salaries scattered from ₹10 L to ₹50 L.

Range

The range is the simplest measure: the difference between the largest and smallest values.

Range=max⁡(xi)−min⁡(xi)\text{Range} = \max(x_i) - \min(x_i)

  • In Excel: =MAX(data) - MIN(data).
  • Limitation: based on only two observations → highly sensitive to extreme values. Rarely used alone.

Interquartile Range (IQR)

The interquartile range (IQR) overcomes the range’s dependency on extremes. It is the range of the middle 50% of the data:

IQR=Q3−Q1\text{IQR} = Q_3 - Q_1

  • Q1Q_1 = 25th percentile, Q3Q_3 = 75th percentile.
  • Not affected by outliers.

Variance

The sample variance uses all data points. It is the average of the squared deviations from the mean.

For nn observations x1,x2,…,xnx_1, x_2, \dots, x_n with sample mean xˉ\bar{x}:

  1. Compute each deviation: xi−xˉx_i - \bar{x}
  2. Square each deviation: (xi−xˉ)2(x_i - \bar{x})^2
  3. Sum all squared deviations: ∑(xi−xˉ)2\sum (x_i - \bar{x})^2
  4. Divide by n−1n-1 (not nn) — a correction that makes the sample variance an unbiased estimator of the population variance.

s2=∑i=1n(xi−xˉ)2n−1s^2 = \frac{\sum_{i=1}^n (x_i - \bar{x})^2}{n-1}

The sum of deviations ∑(xi−xˉ)\sum (x_i - \bar{x}) always equals zero — a direct consequence of the definition of the mean.

Exam tip: Dividing by n−1n-1 (instead of nn) makes little difference when nn is large, but it is essential for correct statistical inference.

Worked example — dataset: 1, 1, 2, 2, 4

  • xˉ=2\bar{x} = 2
  • Deviations: −1,−1,0,0,2-1, -1, 0, 0, 2 → squared: 1,1,0,0,41, 1, 0, 0, 4 → sum = 6
  • s2=65−1=1.5s^2 = \frac{6}{5-1} = 1.5

In Excel: =VAR.S(data).

Standard Deviation

The standard deviation is the square root of the variance. It is measured in the same units as the original data, making it more interpretable.

s=s2s = \sqrt{s^2}

For the above example: s=1.5≈1.22s = \sqrt{1.5} \approx 1.22.

In Excel: =STDEV.S(data).

Coefficient of Variation (CV)

The coefficient of variation is a relative measure of dispersion: standard deviation divided by the mean, expressed as a percentage.

CV=sxˉ×100%\text{CV} = \frac{s}{\bar{x}} \times 100\%

  • Useful for comparing variability across datasets with different units or means.
  • Larger CV → greater relative spread.

Summary of Measures: Five Example Datasets

Dataset (Variable)MeanMedianMinMaxRangessCVIQR
Jyothi – Salary (₹ L)25.92251850323.6614%3
Hanumantha – Monthly spend (₹)3,8523,864———1,45438%1,641
Basavaraja – Satisfaction score57.31—4266246.3811%8.75
Umesh – Weekly sales (units)135.981053443239888.3665%54.25
Vijay – Property price (₹ L)202195554113569346%94.25

Key takeaways

  • Range is simple but fragile; IQR is robust to outliers.
  • Variance and standard deviation use all data; standard deviation is in original units.
  • Coefficient of variation normalises spread by the mean — ideal for comparisons.
  • All these measures are point estimates of the corresponding population parameters.

Z‑Scores (Standardized Scores)

A Z‑score tells how many standard deviations a data point lies above or below the mean. It standardises any variable to a common scale.

zi=xi−xˉsz_i = \frac{x_i - \bar{x}}{s}

  • Positive zz: observation above the mean; negative zz: below the mean.
  • Dimensionless → can compare variables with different units (e.g., household size vs. property price).

Properties of Z‑scores

  • The set of ziz_i for any dataset always has mean = 0 and standard deviation = 1.
  • Why? Subtracting the mean shifts the centre to zero; dividing by the standard deviation scales the spread to 1.

In Excel: =STANDARDIZE(x, mean, sd).

Detecting Outliers with Z‑Scores

A common rule: data points with |zz| > 3 are considered outliers (more than 3 standard deviations from the mean).

Example — Jyothi’s salary data (xˉ=25.92\bar{x}=25.92, s=3.66s=3.66):

  • Minimum z=−2.17z = -2.17 (salary ₹18 L)
  • Maximum z=+6.58z = +6.58 (salary ₹50 L)
  • Two points with z>3z > 3: salaries ₹38 L (z=3.3z=3.3) and ₹50 L (z=6.58z=6.58) → flagged as outliers by the zz-score method.

Key takeaways

  • Z‑score standardises any variable to mean 0, SD 1.
  • ∣z∣>3|z| > 3 is a common outlier threshold.
  • Z‑scores enable cross-variable comparisons.

Five‑Number Summary and Box Plot

The five‑number summary captures the distribution’s centre, spread, and extremes:

  1. Smallest value
  2. First quartile Q1Q_1 (25th percentile)
  3. Median Q2Q_2
  4. Third quartile Q3Q_3 (75th percentile)
  5. Largest value

A box‑and‑whisker plot (box plot) visualises this summary.

Construction Rules

  • Box: ends at Q1Q_1 and Q3Q_3; a line inside marks the median.
  • Interquartile range: IQR=Q3−Q1\text{IQR} = Q_3 - Q_1
  • Fences:
    • Lower fence: Q1−1.5×IQRQ_1 - 1.5 \times \text{IQR}
    • Upper fence: Q3+1.5×IQRQ_3 + 1.5 \times \text{IQR}
  • Whiskers: extend from the box to the smallest and largest data points that lie inside the fences.
  • Outliers: any data point outside the fences is plotted as a small circle.

Exam tip: The 1.5×IQR rule and the ∣z∣>3|z|>3 rule can identify different outliers. No single method is definitive; use both when exploring data.

Example — Jyothi’s Salary Data

  • Q1=24Q_1 = 24, Q3=27Q_3 = 27, median = 25, IQR=3\text{IQR}=3
  • Lower fence: 24−1.5×3=19.524 - 1.5\times3 = 19.5
  • Upper fence: 27+1.5×3=31.527 + 1.5\times3 = 31.5
  • Whiskers: lower whisker ends at 21 (the smallest salary ≥ 19.5); upper whisker ends at 31 (the largest salary ≤ 31.5).
  • Outliers (outside fences): ₹18 L (below lower fence), ₹32 L, ₹35 L, ₹38 L, ₹50 L (above upper fence).

Note: The zz-score method flagged only ₹38 L and ₹50 L; the box plot additionally flags ₹18 L, ₹32 L, ₹35 L.

In Excel: box plots can be created from the charting options (histogram group).

Key takeaways

  • Five‑number summary gives a compact description of shape, spread, and outliers.
  • Box plot uses the 1.5×IQR rule to identify outliers.
  • Compare box‑plot outliers with zz-score outliers for a fuller picture.
  • “A picture is worth a thousand words” — box plots are excellent exploratory tools.

Using the Analysis ToolPak for Descriptive Statistics

Excel’s built-in statistical functions compute one statistic at a time (e.g., =AVERAGE, =VAR.S). The Analysis ToolPak provides a Descriptive Statistics tool that calculates multiple summary measures simultaneously for one or more variables.

Activation and Use

  1. Enable ToolPak: File → Options → Add-Ins → Manage Excel Add-Ins → Analysis Toolpak → OK.
  2. Run tool: Data → Data Analysis → Descriptive Statistics.
  3. Specify inputs: select input range (e.g., C1:I104), check Summary statistics, choose output range.
  4. Output includes: mean, median, mode, variance, standard deviation, range, minimum, maximum, count, kurtosis, skewness, and more — all in one table.

Recap of Descriptive Statistics Covered

The module previously covered:

  • Measures of central tendency: mean, median, mode, percentiles, quartiles.
  • Measures of dispersion: range, interquartile range (IQR), variance, standard deviation, coefficient of variation.
  • Five-number summary and box‑and‑whisker plot.
  • Outlier detection: using z‑score (e.g., |z| > 3) and the box plot (points beyond 1.5×IQR).

Key takeaways – Analysis ToolPak

  • The ToolPak’s Descriptive Statistics tool computes all common summary statistics at once.
  • Activate via File → Options → Add-Ins → Analysis Toolpak.
  • Use the tool from Data → Data Analysis for any selected data range.
  • Output includes both central tendency and dispersion measures automatically.

Basic Concepts of Probability

Probability is a numerical measure of the likelihood of an event, on a scale from 0 (impossible) to 1 (certain). A probability of 0.5 means the event is as likely as not.

Definitions and Notation

  • Random experiment: a process with more than one possible outcome.
  • Sample space (SS): the set of all possible outcomes (each outcome is a sample point).
  • Event: any subset of the sample space.
  • Complement of event AA (AcA^c): all sample points not in AA. P(Ac)=1−P(A)P(A^c) = 1 - P(A)

Example: Movie Revenue

Sanchita Shenoy runs PVR Cinemas’ special re-runs of KGF Chapter 2 and Kantara.

  • KGF revenue can be {3, 4, 5, 6} lakhs.
  • Kantara revenue can be {4, 5, 6} lakhs.
  • A movie is a hit if revenue ≥ 5 lakhs.

The sample space contains 12 outcomes (3×4 combinations). Probabilities are obtained from historical data:

OutcomeKGF (₹ lakhs)Kantara (₹ lakhs)Probability (from data)
134(not given)
2350.10
3360.10
444(not given)
545(not given)
646(not given)
7540.10
8550.08
9560.02
1064(not given)
11650.06
12660.02

(The probabilities shown sum to 1.)

Define events:

  • AA: KGF is a hit (revenue 5 or 6 lakhs) → outcomes 7–12. P(A)=0.10+0.08+0.02+(others)=0.34P(A) = 0.10 + 0.08 + 0.02 + (\text{others}) = 0.34
  • BB: Kantara is a hit (revenue 5 or 6 lakhs) → outcomes 2,3,5,6,8,9,11,12. P(B)=0.10+0.10+⋯=0.60P(B) = 0.10 + 0.10 + \dots = 0.60

Complements:

  • P(Ac)=1−0.34=0.66P(A^c) = 1 - 0.34 = 0.66 (KGF not a hit).
  • P(Bc)=1−0.60=0.40P(B^c) = 1 - 0.60 = 0.40 (Kantara not a hit).

Union and Intersection

  • Union (A∪BA \cup B): event containing all sample points in AA, BB, or both. Intuition: at least one movie is a hit.
  • Intersection (A∩BA \cap B): event containing sample points in both AA and BB. Intuition: both movies are hits.

For the example:

  • Outcomes in A∪BA \cup B = {2,3,5,6,7,8,9,10,11,12} (all outcomes except those where neither is a hit). P(A∪B)=0.76P(A \cup B) = 0.76
  • Outcomes in A∩BA \cap B = {8,9,11,12} (both hits). P(A∩B)=0.08+0.02+0.06+0.02=0.18P(A \cap B) = 0.08 + 0.02 + 0.06 + 0.02 = 0.18

Addition Law

The addition law gives the probability of the union of two events:

P(A∪B)=P(A)+P(B)−P(A∩B)P(A \cup B) = P(A) + P(B) - P(A \cap B)

Verification with the movie example: 0.34+0.60−0.18=0.76(matches)0.34 + 0.60 - 0.18 = 0.76 \quad \text{(matches)}

Exam tip: The subtraction of P(A∩B)P(A \cap B) corrects for double‑counting the overlap. Forgetting this is a common mistake.

Mutually Exclusive Events

Two events are mutually exclusive (disjoint) if they have no sample points in common — A∩B=∅A \cap B = \varnothing. Then P(A∩B)=0P(A \cap B) = 0, and the addition law simplifies to:

P(A∪B)=P(A)+P(B)P(A \cup B) = P(A) + P(B)

Example: A single coin toss — events “heads” and “tails” are mutually exclusive.

In Venn diagram terms, mutually exclusive events are non‑overlapping circles.

Key takeaways – Probability Basics

  • Probability ranges from 0 to 1; P(S)=1P(S) = 1.
  • Complement: P(Ac)=1−P(A)P(A^c) = 1 - P(A).
  • Union probability includes all outcomes in either event; intersection probability includes only common outcomes.
  • Addition law: P(A∪B)=P(A)+P(B)−P(A∩B)P(A \cup B) = P(A) + P(B) - P(A \cap B).
  • For mutually exclusive events, P(A∩B)=0P(A \cap B) = 0 so P(A∪B)=P(A)+P(B)P(A \cup B) = P(A) + P(B).

Bayes’ Theorem

Bayes’ Theorem links prior probability (initial belief about an event) to posterior probability (updated belief after observing new evidence). Intuitively: when you learn that event BB has occurred, your estimate of the probability of event AA may change. Bayes’ Theorem quantifies exactly how.

Conditional Probability

The probability of AA given that BB has occurred is conditional probability:

P(A∣B)=P(A∩B)P(B)P(A \mid B) = \frac{P(A \cap B)}{P(B)}

Similarly,

P(B∣A)=P(A∩B)P(A)P(B \mid A) = \frac{P(A \cap B)}{P(A)}

Multiplication Law

The multiplication law expresses the intersection probability in two equivalent ways:

P(A∩B)=P(A∣B) P(B)=P(B∣A) P(A)P(A \cap B) = P(A \mid B) \, P(B) = P(B \mid A) \, P(A)

These relations are the building blocks of Bayes’ Theorem.

Relating Prior and Posterior

Given a prior P(A)P(A) and new information BB, the posterior P(A∣B)P(A \mid B) is:

P(A∣B)=P(B∣A) P(A)P(B)P(A \mid B) = \frac{P(B \mid A) \, P(A)}{P(B)}

Interpretation: The posterior is proportional to the likelihood of the evidence given AA times the prior, normalized by the total probability of BB.

Worked Example (Movie Hits)

  • AA = movie KGF is a hit, prior P(A)=0.34P(A)=0.34
  • BB = movie Kantara is a hit, P(B)=0.60P(B)=0.60
  • Joint P(A∩B)=0.18P(A \cap B)=0.18

Compute posterior P(A∣B)P(A \mid B):

P(A∣B)=0.180.60=0.30P(A \mid B) = \frac{0.18}{0.60} = 0.30

The posterior (0.30) is lower than the prior (0.34) – knowing Kantara was a hit reduced the chance of KGF being a hit.

Compute P(B∣A)P(B \mid A):

P(B∣A)=0.180.34≈0.53P(B \mid A) = \frac{0.18}{0.34} \approx 0.53

Again, the posterior (0.53) is lower than the prior of BB (0.60).

Joint Probability Table

A joint probability table arranges all intersection probabilities for quick computation of conditionals and marginals.

BB (Kantara hit)BcB^c (not hit)Marginal
AA (KGF hit)0.180.160.34
AcA^c (not hit)0.420.240.66
Marginal0.600.401.00
  • Marginal probabilities are in the last row/column (e.g., P(A)=0.34P(A)=0.34, P(B)=0.60P(B)=0.60).
  • To find P(B∣A)P(B \mid A), divide the joint in row AA (0.18) by the row marginal (0.34).

Exam tip: Always check whether you are conditioning on a row or a column. The denominator is always the marginal of the conditioning event.

Independence of Events

Two events AA and BB are independent if the occurrence of one does not affect the probability of the other:

P(A∣B)=P(A)orP(B∣A)=P(B)P(A \mid B) = P(A) \quad \text{or} \quad P(B \mid A) = P(B)

Equivalently:

P(A∩B)=P(A) P(B)P(A \cap B) = P(A) \, P(B)

Mutually exclusive events with non-zero probability cannot be independent (if one happens, the other cannot, so P(A∣B)=0≠P(A)P(A \mid B)=0 \neq P(A)).

  • In the movie example: P(A∩B)=0.18P(A \cap B)=0.18, P(A)P(B)=0.34×0.60=0.204P(A)P(B)=0.34 \times 0.60=0.204. Since 0.18≠0.2040.18 \neq 0.204, the events are not independent.

Law of Total Probability

The law of total probability expresses the total probability of AA as a sum over mutually exclusive and collectively exhaustive events:

P(A)=P(A∩B)+P(A∩Bc)P(A) = P(A \cap B) + P(A \cap B^c)

Generalized to nn events B1,B2,…,BnB_1, B_2, \dots, B_n that partition the sample space:

P(A)=∑i=1nP(A∩Bi)=∑i=1nP(A∣Bi) P(Bi)P(A) = \sum_{i=1}^{n} P(A \cap B_i) = \sum_{i=1}^{n} P(A \mid B_i) \, P(B_i)

Example: For movie KGF being a hit, consider three revenue outcomes for Kantara (B1B_1=₹4 lakhs, B2B_2=₹5 lakhs, B3B_3=₹6 lakhs). Then P(A)=P(A∩B1)+P(A∩B2)+P(A∩B3)=0.34P(A) = P(A \cap B_1) + P(A \cap B_2) + P(A \cap B_3) = 0.34.

Worked Example: Jyothi Hegde’s Placement Data

Categorize respondents by work experience (WW: ≤2 years or >2 years) and salary (SS: ≤₹25 lakhs or >₹25 lakhs). Total n=103n=103.

S≤25S \leq 25S>25S > 25Row total
W≤2W \leq 2341347
W>2W > 2223456
Column total5647103

Convert to probabilities (divide by 103):

S≤25S \leq 25S>25S > 25Marginal
W≤2W \leq 20.3300.1260.456
W>2W > 20.2140.3300.544
Marginal0.5440.4561.000

Compute probabilities:

  • P(W>2)=56/103≈0.544P(W > 2) = 56/103 \approx 0.544
  • P(S>25)=47/103≈0.456P(S > 25) = 47/103 \approx 0.456
  • P(W≤2∩S>25)=13/103≈0.126P(W \leq 2 \cap S > 25) = 13/103 \approx 0.126
  • P(W>2∣S>25)=34/47≈0.723P(W > 2 \mid S > 25) = 34/47 \approx 0.723
  • P(S>25∣W>2)=34/56≈0.607P(S > 25 \mid W > 2) = 34/56 \approx 0.607

Key takeaways

  • Prior → posterior: update probability when new information arrives.
  • Conditional probability: P(A∣B)=P(A∩B)/P(B)P(A \mid B) = P(A \cap B)/P(B).
  • Bayes’ Theorem: P(A∣B)=P(B∣A)P(A)/P(B)P(A \mid B) = P(B \mid A)P(A)/P(B).
  • Use a joint probability table to organize and compute conditionals and marginals.
  • Independence is tested via P(A∩B)=P(A)P(B)P(A \cap B)=P(A)P(B); mutually exclusive events are dependent.
  • Law of total probability sums over partitions to find P(A)P(A).