Descriptive Statistics: Data Types & Terminology
Descriptive statistics is the process of describing collected data through visuals and numerical summaries. It helps turn raw facts into a story for decision-making.
What is "Statistics"?
The word can mean either:
- Numerical facts computed from data (e.g., average, median, range).
- The process of collecting, analysing, presenting, and interpreting data — a blend of science (fixed principles) and art (subtle interpretive nuance).
Statistics is used across business functions: accounting/finance (performance analysis), economics (policy), marketing/production (grow profit, reduce cost), and information systems (system health).
Data and Data Set
- Data: Facts and figures collected, analysed, and summarised for decision-making.
- Data set: The entire collection of data for a specific study.
Example: Manjula Nayak's Bengaluru Household Survey An excerpt with 9 households and multiple variables:
| Family Size | Location (numeric code) | Location (text) | Ownership (0/1) | Income 1 (₹) | Income 2 (₹) | Monthly Payment (₹) | Utilities (₹) | Debt (₹) | Satisfaction (numeric) | Satisfaction (Likert) |
|---|---|---|---|---|---|---|---|---|---|---|
| 3 | 2 | Northwest | 1 | 1,700,000 | 0 | 12,000 | 3,000 | 50,000 | 4 | Satisfied |
- Element: The entity for which data are collected (here: each household).
- Variable: A characteristic of interest for each element (e.g., family size, location, income).
- Observation: The complete set of measurements for one element — one entire row.
⚠️ Quiz trap — In a different dataset (Tumkur apartment rents), a table with 7 rows and 10 columns contained 70 elements (households), each with only one variable: monthly rent. Observations were single-cell rents, not whole rows. Never assume rows = observations and columns = variables. Check what each column represents.
Types of Data
Data can be broadly qualitative or quantitative, each with finer subcategories. The type determines how you visualise and analyse it.
Qualitative Data (Categorical)
Records attributes or labels; can be numeric or non-numeric. Divides into:
| Scale | Property | Example from survey |
|---|---|---|
| Nominal | Labels/names, no ranking | Location (Southwest=1, Northwest=2, Northeast=3, Southeast=4) — numeric but no order; Ownership (0=rent, 1=own) |
| Ordinal | All nominal properties plus inherent ranking | Satisfaction (5=Extremely satisfied … 1=Extremely dissatisfied) — both non-numeric and numeric forms imply rank |
Exam tip: Numbers alone don't make data quantitative. Location codes (1,2,3,4) are still qualitative because they merely label — there is no meaningful ordering.
Quantitative Data (Numerical)
Values have magnitude and order. Two sub-scales:
| Scale | Meaningful ratio? | Zero meaningful? | Examples |
|---|---|---|---|
| Interval | ❌ No (10°C is not “twice as hot” as 5°C) | ❌ No (0°C does not mean “no temperature”) | Temperature (°C), clock time (2 AM, 4 AM) — intervals matter, not ratios |
| Ratio | ✅ Yes (₹2 Lakh profit is twice ₹1 Lakh) | ✅ Yes (₹0 profit = break-even) | Income, debt, age, profit — both ratios and zero are meaningful |
Worked comparisons:
- Ownership (0/1) → qualitative, nominal (no ranking: 1 is not “better” than 0).
- Satisfaction numeric (1–5) → qualitative, ordinal (5 is better than 1).
- Income (₹) → quantitative, ratio (₹0 = no income; ₹2L/₹1L = twice).
- Temperature (Celsius) → quantitative, interval (0°C is arbitrary; 30°C/10°C ≠ "three times hotter").
Key Takeaways
- Descriptive statistics involves summarising data via visuals and numerical summaries.
- An observation is the full set of variables for one element, not necessarily one row — always check the data structure.
- Qualitative data can be numeric or non-numeric; its scale is nominal (no rank) or ordinal (rank implied).
- Quantitative data is always numeric; its scale is interval (ratio meaningless, zero arbitrary) or ratio (ratio meaningful, zero absolute).
- The variable type (nominal, ordinal, interval, ratio) determines which visualisation and analysis techniques are appropriate.
Core Classification Framework
Every variable can be classified along three binary axes:
- Qualitative (categorical) vs Quantitative (numerical measurement)
- Numeric (values are numbers) vs Non‑numeric (labels are text)
- Nominal (no order) vs Ordinal (natural order) – for qualitative variables Interval (equal intervals, no true zero) vs Ratio (true zero, ratios meaningful) – for quantitative variables
Key insight: A qualitative variable can be numeric (e.g., a 1‑5 rating) – the numbers act as category labels, not measurements. Likewise, a quantitative variable is always numeric.
1. Alumni Survey Data (Jyothi Hegde — IIIT Dharwad)
100+ rows; variables describe student profiles, test scores, and outcomes.
| Variable | Type | Numeric? | Sub‑type | Rationale |
|---|---|---|---|---|
| Gender (M/W) | Qualitative | Non‑numeric | Nominal | Labels only, no order |
| Age (years) | Quantitative | Numeric | Ratio | True zero (age 0), ratios meaningful |
| Overall (0‑100) | Quantitative | Numeric | Ratio | Scale includes 0, zero has meaning |
| Core courses (2‑10 GPA) | Quantitative | Numeric | Interval | Scale starts at 2, not 0; ratio not valid |
| Elective (2‑10 GPA) | Quantitative | Numeric | Interval | Same as core |
| Work experience (years) | Quantitative | Numeric | Ratio | True zero |
| Domicile (1=in‑state, 2=out) | Qualitative | Numeric | Nominal | Numbers are codes, no ordering |
| Rating (1‑5) | Qualitative | Numeric | Ordinal | Numbers imply rank (5 > 1) |
| Satisfaction (labels) | Qualitative | Non‑numeric | Ordinal | “Very satisfied” > “Not at all satisfied” |
| Salary (lakhs) | Quantitative | Numeric | Ratio | True zero, ratios meaningful |
Take‑away for exam:
- A numeric variable is not automatically quantitative – e.g., rating and domicile are qualitative despite being numbers.
- Test scores with a non‑zero minimum (e.g., 2‑10) are interval, not ratio, because a score of 4 is not “twice as good” as 2.
2. Credit Card Spending (Hanumanthappa – Canara Bank, Mangaluru)
55 customers; variables: annual income, household size, amount charged.
| Variable | Type | Numeric? | Sub‑type |
|---|---|---|---|
| Annual income (₹ lakhs) | Quantitative | Numeric | Ratio |
| Household size (# persons) | Quantitative | Numeric | Ratio |
| Amount charged (₹/month) | Quantitative | Numeric | Ratio |
All three are quantitative, numeric, ratio – straightforward.
3. Chroma Customer Survey (Basavaraja – Bellary)
36 responses; includes computer attributes and satisfaction score.
| Variable | Type | Numeric? | Sub‑type | Rationale |
|---|---|---|---|---|
| Brand (e.g., Lenovo, Dell) | Qualitative | Non‑numeric | Nominal | No order |
| Price (₹ lakhs) | Quantitative | Numeric | Ratio | True zero |
| RAM (GB) | Qualitative | Numeric | Ordinal | 32 GB > 16 GB in performance, but not a measurement of “RAM quantity” – it’s a category with order |
| Hard disk (GB) | Qualitative | Numeric | Ordinal | Same logic as RAM |
| Operating system (Win10/11, Home/Pro) | Qualitative | Non‑numeric | Nominal | Later version ≠ better (e.g., Win11 not inherently superior); no clear order |
| Satisfaction score (0‑100) | Quantitative | Numeric | Ratio | True zero, ratios meaningful |
Exam trap: RAM and hard disk are trick variables – they look quantitative but are treated as ordinal categorical because the values are arbitrary tiers, not continuous measurements. Never assume numeric ⇒ quantitative.
4. Nandini Sweets Marketing (Umesh Kamath – Udupi)
Weekly store data; variables: date, store, sales, price, promotions (0/1).
| Variable | Type | Numeric? | Sub‑type | Rationale |
|---|---|---|---|---|
| Date (week) | Quantitative | Numeric | Interval | Zero date is arbitrary; ratio of dates meaningless |
| Store name (text) | Qualitative | Non‑numeric | Nominal | No order |
| Units sold | Quantitative | Numeric | Ratio | True zero |
| Retail price (₹) | Quantitative | Numeric | Ratio | True zero |
| Coupons (0/1) | Qualitative | Numeric | Nominal | 0/1 are codes, no order |
| Displays (0/1) | Qualitative | Numeric | Nominal | Same |
| Sampling (0/1) | Qualitative | Numeric | Nominal | Same |
Key point: Binary (0/1) variables are qualitative nominal – the numbers are labels, not measurements.
5. Magic Bricks Property Data (Vijay Gowda – Mysuru)
30 apartments; variables: #bathrooms, #bedrooms, premium locality (0/1), gated complex (0/1), amenities (0/1), square footage, selling price.
| Variable | Type | Numeric? | Sub‑type |
|---|---|---|---|
| # bathrooms | Quantitative | Numeric | Ratio |
| # bedrooms | Quantitative | Numeric | Ratio |
| Premium locality (0/1) | Qualitative | Numeric | Nominal |
| Gated complex (0/1) | Qualitative | Numeric | Nominal |
| Amenities (0/1) | Qualitative | Numeric | Nominal |
| Square footage | Quantitative | Numeric | Ratio |
| Selling price (₹) | Quantitative | Numeric | Ratio |
All quantitative variables are ratio; all qualitative variables are numeric nominal (0/1 codes).
Key Takeaways
- Qualitative variables classify; they can be numeric (e.g., rating 1‑5) or non‑numeric. Sub‑types: nominal (no order) or ordinal (ranked).
- Quantitative variables measure; always numeric. Sub‑types: interval (no true zero, e.g., date, test scores starting >0) or ratio (true zero, e.g., income, age, sales).
- Binary (0/1) variables are qualitative nominal – the numbers are codes, not counts.
- Numeric does not imply quantitative – always check whether the variable is a measurement or a label.
- Common exam traps:
- Test scores with non‑zero minimum → interval (not ratio).
- RAM/HDD sizes in tiers → ordinal qualitative (not quantitative).
- Later version of software → not automatically ordinal; judge whether order is meaningful.
Types of Data: Cross-Sectional vs. Time Series
Distinguishing between cross-sectional data and time series data affects how you analyse trends, correlations, and cause‑effect relationships.
- Cross-sectional data – collected at the same (or approximately the same) point in time across multiple subjects (e.g., households, districts, individuals).
- Time series data – collected over several time periods for one or more subjects (e.g., months, years).
Key intuition
- Cross‑sectional: a snapshot across many units at one time.
- Time series: a movie of one (or several) units across time.
Examples
| Study / Dataset | Data collected | Classification | Reason |
|---|---|---|---|
| Manjula’s household survey – rent/own, income, monthly payment | At one point in time across many households | Cross-sectional | Snapshot across households |
| Jyoti Hegde – placement/satisfaction at IIIT Dharwad (gender, age, exam scores, work experience, salary) | Surveyed current people at one time, though they graduated at different times | Cross-sectional | Survey conducted at a single point in time |
| Hanumanth Pai – credit card data (income, household size, monthly charges) | One‑time data on many customers | Cross-sectional | Snapshot across customers |
| Basavaraju – customer loyalty at Chroma (computer purchase, brand, price, features) | Survey done all at once | Cross-sectional | Snapshot across customers |
| Umesh Kamath – Nandini Sweets (marketing promotions, units sold, selling price, promotion strategy over multiple periods) | Data collected over several time periods for stores in Belgaum and Gulbarga | Time series | Repeated measurements over time |
| Vijay Gowda – real estate (bathrooms, bedrooms, locality, amenities, square footage, selling price) | Data collected all at once | Cross-sectional | Snapshot of flats at one time |
Exam tip: The “snapshot across subjects” vs. “repeated measures over time” distinction is often tested. Even if subjects graduated in different years, if you survey them now, it’s cross‑sectional. The key is the time of data collection, not the time period the data describe.
Data Collection Methods
When data are not available from existing sources, they must be collected. Three main approaches:
1. Existing (Secondary) Sources
Data already exist and can be used without separate collection.
- Examples: Company records (operations, personnel, finances); government data (census); specialized organisations (Bloomberg, Nielsen Company).
- Advantage: Fast, inexpensive, often large scale.
2. Observational Studies
The researcher observes and records variables of interest without intervening.
- When to use: When existing sources are insufficient.
- Process: Simply watch and measure what happens in a real‑world setting.
- Examples: Observers in a store noting customer choices; HR tracking employee attrition over time.
3. Experimental Studies
Data are collected in a controlled manner – the researcher decides the target group, specific situations/conditions, and what data to collect.
- When to use: To isolate cause‑and‑effect (often via controlled manipulation).
- Examples: Clinical trials for pharmaceuticals (target group identified, process specified, predefined data types).
Key contrast: In observational studies you record what naturally occurs; in experimental studies you control conditions to test a hypothesis.
Key takeaways – Types of data
- Cross‑sectional: one‑time snapshot across subjects → compare differences.
- Time series: repeated observations over time → track trends and patterns.
- Classification depends on when the data are collected, not what they describe.
Key takeaways – Data collection
- Existing sources are the cheapest route, but may lack the variables you need.
- Observational studies record natural behaviour (no intervention).
- Experimental studies control conditions to establish causal relationships.
Data Visualization
Data visualization uses tables and graphs to summarize categorical and quantitative data. The goal is not just to create them but to choose the right format for the context and interpret it correctly.
Tabular Summaries for Categorical Data
A frequency distribution is a table that shows the count (frequency) of observations in each of several non‑overlapping categories.
Example: Satisfaction levels (5 levels: Very Satisfied, Satisfied, Neutral, Unsatisfied, Not at all Satisfied) for IIIT Dharwad respondents (103 total: 72 men, 31 women).
| Satisfaction Level | Men (freq) | Women (freq) | Overall (freq) |
|---|---|---|---|
| Very Satisfied | 17 | 9 | 26 |
| Satisfied | 22 | 4 | 26 |
| Neutral | 22 | 6 | 28 |
| Unsatisfied | 9 | 7 | 16 |
| Not at all Satisfied | 2 | 5 | 7 |
A relative frequency is the proportion of observations in a class:
The percent frequency is the relative frequency multiplied by 100.
Worked example – Men:
- Very Satisfied:
- Satisfied:
- Neutral:
- Unsatisfied:
- Not at all Satisfied:
Women example: Very Satisfied: .
A relative frequency distribution and percent frequency distribution are tables of these values.
Graphical Displays for Categorical Data
-
Bar chart – X‑axis: categories; Y‑axis: frequency or percent. Height of bar proportional to count/percent.
-
Pie chart – Circle divided into sectors; sector angle = relative frequency × 360°.
Example: For men very satisfied (24%): sector angle .
Exam tip: Pie charts are visually harder to compare than bar charts. Most data‑visualisation experts prefer bar charts for categorical data.
Comparing Two Categorical Variables
- Side‑by‑side bar chart – Two or more bars per category (e.g., men vs. women for each satisfaction level). Orientation can be swapped (categories vs. groups on x‑axis).
- Stacked bar chart – One bar per category, segmented by the second variable. Height of each segment = frequency.
- Stacked percentage bar chart – Each bar scaled to 100%; segments show the percentage within each category. Use when absolute counts don’t matter, only proportions.
Choose the format based on the comparison you want to highlight.
Tabular Summaries for Quantitative Data
Frequency, relative frequency, and percent frequency distributions also work for quantitative data, but must define non‑overlapping classes with equal width and 5–10 classes. Each data point falls into exactly one class.
Example – Salary (in ₹ Lakhs) using 7 classes of width 5:
| Salary Class | Men (freq) | Women (freq) | Overall (freq) |
|---|---|---|---|
| 15–19 | ... | ... | ... |
| 20–24 | ... | ... | ... |
| 25–29 | ... | ... | ... |
| 30–34 | ... | ... | ... |
| 35–39 | ... | ... | ... |
| 40–44 | ... | ... | ... |
| 45–50 | ... | ... | ... |
Guidelines for class intervals:
- Same width for all classes.
- Typically 5–10 classes.
- Limits chosen so each observation belongs to exactly one class.
Relative frequency and percent frequency are computed the same way as for categorical data.
Graphical Displays for Quantitative Data
-
Histogram – Like a bar chart but rectangles touch (no gap). X‑axis: classes; Y‑axis: frequency. Reveals the shape of the distribution (e.g., skewness).
-
Scatter plot – Shows relationship between two quantitative variables. Each point = one observation (x‑coordinate for one variable, y‑coordinate for the other). A trend line may be added to suggest a pattern.
Example: Plot salary vs. age for men and women from the survey.
Choosing the Right Display
| Purpose | Recommended display |
|---|---|
| Show distribution of a single variable (categorical or quantitative) | Bar chart, pie chart (categorical); histogram (quantitative) |
| Compare two categorical variables | Side‑by‑side bar chart, stacked bar chart, stacked percentage bar chart |
| Show relationship between two quantitative variables | Scatter plot (optionally with trend line) |
Data Dashboards
In practice, businesses use data dashboards – a collection of visual displays that organise key metrics for monitoring performance and supporting quick decisions.
Key takeaways
- Frequency distribution: count per non‑overlapping class; relative frequency = frequency / n; percent = ×100.
- Bar charts preferred over pie charts for comparing categories.
- Side‑by‑side and stacked bar charts compare two variables.
- Histograms (touching rectangles) reveal distribution shape.
- Scatter plots show relationships between quantitative variables.
- Class intervals for quantitative data should be equal‑width, 5–10 classes, non‑overlapping.
Measures of Central Tendency
Numerical summaries for data fall into two categories: measures of central tendency (where the data cluster) and measures of dispersion (how spread out they are). Central tendency includes the mean, median, mode, percentiles, and quartiles.
Mean (Arithmetic Average)
The mean is the average value of a variable – the most common measure of central location. For a sample it is denoted (x‑bar); for a population it is denoted . The sample mean is a point estimator of the population mean.
Excel: =AVERAGE(range)
Property: The mean is sensitive to extreme values (outliers). A single very large or very small observation can pull the mean away from the centre.
Median
The median is the middle value when the data are sorted in ascending order.
- If is odd: the median is the th observation.
- If is even: the median is the average of the two middle observations.
Robustness: The median is not affected by extreme values.
Example Dataset: 1, 1, 2, 2, 4 → mean = 2, median = 2. Change the last value to 394: mean = 80, median = 2. The median remains unchanged.
Because of this, the median is the preferred measure for skewed distributions such as annual income or property prices.
Excel: =MEDIAN(range)
Mode
The mode is the value that occurs with the greatest frequency.
- Unimodal – exactly one mode.
- Bimodal – two values tied for highest frequency (e.g., 1,1,2,2,4 has modes 1 and 2).
- Multimodal – more than two modes; rarely reported because listing many values is unhelpful.
If every value is unique, no mode exists.
Excel:
=MODE.SNGL(range)– returns the first most frequent value.=MODE.MULT(range)– returns an array of all modes.- If no mode, Excel returns
#N/A.
Percentiles
The th percentile divides the data so that approximately of observations are less than or equal to that value, and are greater.
Computation (the method used by PERCENTILE.EXC):
- Sort data in ascending order.
- Compute index
- If is an integer, take the th observation.
- If is not an integer, interpolate between the floor and ceiling observations.
Example: 80th percentile with → → interpolate between the 160th and 161st values.
Excel: =PERCENTILE.EXC(array, k) where .
Interpretation: A test score at the 80th percentile means about 80% of test‑takers scored lower (and 20% higher).
Quartiles
Quartiles divide the data into four parts, each containing roughly 25% of observations. They are just specific percentiles:
| Quartile | Percentile | Also known as |
|---|---|---|
| Q1 (first quartile) | 25th | Lower quartile |
| Q2 (second quartile) | 50th | Median |
| Q3 (third quartile) | 75th | Upper quartile |
Excel: Use =PERCENTILE.EXC with .
Applied Examples
The following five datasets illustrate how measures of central tendency behave in practice.
1. IIIT Dharwad – salaries of respondents ()
| Measure | Value (₹ lakhs) |
|---|---|
| Mean | 25.9 |
| Median | 25 |
| Mode | 25 |
| Q1 | 24 |
| Q2 (median) | 25 |
| Q3 | 27 |
| 80th percentile | 27 |
| 90th percentile | 29 |
Interpretation: Mean and median are close – no extreme salaries pulling the average.
2. Canara Bank – monthly credit card spend ()
| Measure | Value (₹) |
|---|---|
| Mean | 3,852 |
| Median | 3,864 |
| Mode | 3,657 |
| Q1 | 2,890 |
| Q3 | 4,531 |
Mean and median are similar (no extreme spenders), but mode differs – the most common spend amount is lower.
3. Chroma – customer satisfaction score ()
| Measure | Value |
|---|---|
| Mean | 57.31 |
| Median | 59 |
| Q1 | 53 |
| Q3 | 61.75 |
Median slightly higher than mean; moderate spread.
4. Nandini sweets – weekly sales ()
| Measure | Value (₹) |
|---|---|
| Mean | 135.98 |
| Median | 105 |
| Mode | 91 |
| Q1 | 82.25 |
| Q3 | 136.5 |
Mean () is much higher than the median (), suggesting that a few weeks with very high sales pull the average up.
5. Jayalakshmipuram – property prices ( not stated, but data given)
| Measure | Value (₹ lakhs) |
|---|---|
| Mean | 202.57 |
| Median | 195 |
| Q1 | 149.5 |
| Q3 | 243.75 |
Mean > median, indicating some high‑priced properties inflating the average – a common pattern in real estate.
Exam tip: The median is the better measure of “typical” value when data are skewed or contain outliers (e.g., income, house prices). Always check whether mean and median are close – if not, the mean is likely distorted.
Key takeaways
- Mean = sum ÷ , sensitive to outliers; median = middle value, robust; mode = most frequent.
- Percentiles and quartiles describe the distribution’s spread. Q2 = median.
- Excel commands:
AVERAGE,MEDIAN,MODE.SNGL/MODE.MULT,PERCENTILE.EXC. - When mean and median differ substantially, suspect extreme values in the tail.
Measures of Dispersion
Measures of dispersion (or variability) quantify how spread out a dataset is. Intuition: two datasets can have the same mean but very different spreads — think of salaries clustered around ₹25 L vs. salaries scattered from ₹10 L to ₹50 L.
Range
The range is the simplest measure: the difference between the largest and smallest values.
- In Excel:
=MAX(data) - MIN(data). - Limitation: based on only two observations → highly sensitive to extreme values. Rarely used alone.
Interquartile Range (IQR)
The interquartile range (IQR) overcomes the range’s dependency on extremes. It is the range of the middle 50% of the data:
- = 25th percentile, = 75th percentile.
- Not affected by outliers.
Variance
The sample variance uses all data points. It is the average of the squared deviations from the mean.
For observations with sample mean :
- Compute each deviation:
- Square each deviation:
- Sum all squared deviations:
- Divide by (not ) — a correction that makes the sample variance an unbiased estimator of the population variance.
The sum of deviations always equals zero — a direct consequence of the definition of the mean.
Exam tip: Dividing by (instead of ) makes little difference when is large, but it is essential for correct statistical inference.
Worked example — dataset: 1, 1, 2, 2, 4
- Deviations: → squared: → sum = 6
In Excel: =VAR.S(data).
Standard Deviation
The standard deviation is the square root of the variance. It is measured in the same units as the original data, making it more interpretable.
For the above example: .
In Excel: =STDEV.S(data).
Coefficient of Variation (CV)
The coefficient of variation is a relative measure of dispersion: standard deviation divided by the mean, expressed as a percentage.
- Useful for comparing variability across datasets with different units or means.
- Larger CV → greater relative spread.
Summary of Measures: Five Example Datasets
| Dataset (Variable) | Mean | Median | Min | Max | Range | CV | IQR | |
|---|---|---|---|---|---|---|---|---|
| Jyothi – Salary (₹ L) | 25.92 | 25 | 18 | 50 | 32 | 3.66 | 14% | 3 |
| Hanumantha – Monthly spend (₹) | 3,852 | 3,864 | — | — | — | 1,454 | 38% | 1,641 |
| Basavaraja – Satisfaction score | 57.31 | — | 42 | 66 | 24 | 6.38 | 11% | 8.75 |
| Umesh – Weekly sales (units) | 135.98 | 105 | 34 | 432 | 398 | 88.36 | 65% | 54.25 |
| Vijay – Property price (₹ L) | 202 | 195 | 55 | 411 | 356 | 93 | 46% | 94.25 |
Key takeaways
- Range is simple but fragile; IQR is robust to outliers.
- Variance and standard deviation use all data; standard deviation is in original units.
- Coefficient of variation normalises spread by the mean — ideal for comparisons.
- All these measures are point estimates of the corresponding population parameters.
Z‑Scores (Standardized Scores)
A Z‑score tells how many standard deviations a data point lies above or below the mean. It standardises any variable to a common scale.
- Positive : observation above the mean; negative : below the mean.
- Dimensionless → can compare variables with different units (e.g., household size vs. property price).
Properties of Z‑scores
- The set of for any dataset always has mean = 0 and standard deviation = 1.
- Why? Subtracting the mean shifts the centre to zero; dividing by the standard deviation scales the spread to 1.
In Excel: =STANDARDIZE(x, mean, sd).
Detecting Outliers with Z‑Scores
A common rule: data points with || > 3 are considered outliers (more than 3 standard deviations from the mean).
Example — Jyothi’s salary data (, ):
- Minimum (salary ₹18 L)
- Maximum (salary ₹50 L)
- Two points with : salaries ₹38 L () and ₹50 L () → flagged as outliers by the -score method.
Key takeaways
- Z‑score standardises any variable to mean 0, SD 1.
- is a common outlier threshold.
- Z‑scores enable cross-variable comparisons.
Five‑Number Summary and Box Plot
The five‑number summary captures the distribution’s centre, spread, and extremes:
- Smallest value
- First quartile (25th percentile)
- Median
- Third quartile (75th percentile)
- Largest value
A box‑and‑whisker plot (box plot) visualises this summary.
Construction Rules
- Box: ends at and ; a line inside marks the median.
- Interquartile range:
- Fences:
- Lower fence:
- Upper fence:
- Whiskers: extend from the box to the smallest and largest data points that lie inside the fences.
- Outliers: any data point outside the fences is plotted as a small circle.
Exam tip: The 1.5×IQR rule and the rule can identify different outliers. No single method is definitive; use both when exploring data.
Example — Jyothi’s Salary Data
- , , median = 25,
- Lower fence:
- Upper fence:
- Whiskers: lower whisker ends at 21 (the smallest salary ≥ 19.5); upper whisker ends at 31 (the largest salary ≤ 31.5).
- Outliers (outside fences): ₹18 L (below lower fence), ₹32 L, ₹35 L, ₹38 L, ₹50 L (above upper fence).
Note: The -score method flagged only ₹38 L and ₹50 L; the box plot additionally flags ₹18 L, ₹32 L, ₹35 L.
In Excel: box plots can be created from the charting options (histogram group).
Key takeaways
- Five‑number summary gives a compact description of shape, spread, and outliers.
- Box plot uses the 1.5×IQR rule to identify outliers.
- Compare box‑plot outliers with -score outliers for a fuller picture.
- “A picture is worth a thousand words” — box plots are excellent exploratory tools.
Using the Analysis ToolPak for Descriptive Statistics
Excel’s built-in statistical functions compute one statistic at a time (e.g., =AVERAGE, =VAR.S). The Analysis ToolPak provides a Descriptive Statistics tool that calculates multiple summary measures simultaneously for one or more variables.
Activation and Use
- Enable ToolPak:
File → Options → Add-Ins → Manage Excel Add-Ins → Analysis Toolpak → OK. - Run tool:
Data → Data Analysis → Descriptive Statistics. - Specify inputs: select input range (e.g.,
C1:I104), check Summary statistics, choose output range. - Output includes: mean, median, mode, variance, standard deviation, range, minimum, maximum, count, kurtosis, skewness, and more — all in one table.
Recap of Descriptive Statistics Covered
The module previously covered:
- Measures of central tendency: mean, median, mode, percentiles, quartiles.
- Measures of dispersion: range, interquartile range (IQR), variance, standard deviation, coefficient of variation.
- Five-number summary and box‑and‑whisker plot.
- Outlier detection: using z‑score (e.g., |z| > 3) and the box plot (points beyond 1.5×IQR).
Key takeaways – Analysis ToolPak
- The ToolPak’s Descriptive Statistics tool computes all common summary statistics at once.
- Activate via File → Options → Add-Ins → Analysis Toolpak.
- Use the tool from Data → Data Analysis for any selected data range.
- Output includes both central tendency and dispersion measures automatically.
Basic Concepts of Probability
Probability is a numerical measure of the likelihood of an event, on a scale from 0 (impossible) to 1 (certain). A probability of 0.5 means the event is as likely as not.
Definitions and Notation
- Random experiment: a process with more than one possible outcome.
- Sample space (): the set of all possible outcomes (each outcome is a sample point).
- Event: any subset of the sample space.
- Complement of event (): all sample points not in .
Example: Movie Revenue
Sanchita Shenoy runs PVR Cinemas’ special re-runs of KGF Chapter 2 and Kantara.
- KGF revenue can be {3, 4, 5, 6} lakhs.
- Kantara revenue can be {4, 5, 6} lakhs.
- A movie is a hit if revenue ≥ 5 lakhs.
The sample space contains 12 outcomes (3×4 combinations). Probabilities are obtained from historical data:
| Outcome | KGF (₹ lakhs) | Kantara (₹ lakhs) | Probability (from data) |
|---|---|---|---|
| 1 | 3 | 4 | (not given) |
| 2 | 3 | 5 | 0.10 |
| 3 | 3 | 6 | 0.10 |
| 4 | 4 | 4 | (not given) |
| 5 | 4 | 5 | (not given) |
| 6 | 4 | 6 | (not given) |
| 7 | 5 | 4 | 0.10 |
| 8 | 5 | 5 | 0.08 |
| 9 | 5 | 6 | 0.02 |
| 10 | 6 | 4 | (not given) |
| 11 | 6 | 5 | 0.06 |
| 12 | 6 | 6 | 0.02 |
(The probabilities shown sum to 1.)
Define events:
- : KGF is a hit (revenue 5 or 6 lakhs) → outcomes 7–12.
- : Kantara is a hit (revenue 5 or 6 lakhs) → outcomes 2,3,5,6,8,9,11,12.
Complements:
- (KGF not a hit).
- (Kantara not a hit).
Union and Intersection
- Union (): event containing all sample points in , , or both. Intuition: at least one movie is a hit.
- Intersection (): event containing sample points in both and . Intuition: both movies are hits.
For the example:
- Outcomes in = {2,3,5,6,7,8,9,10,11,12} (all outcomes except those where neither is a hit).
- Outcomes in = {8,9,11,12} (both hits).
Addition Law
The addition law gives the probability of the union of two events:
Verification with the movie example:
Exam tip: The subtraction of corrects for double‑counting the overlap. Forgetting this is a common mistake.
Mutually Exclusive Events
Two events are mutually exclusive (disjoint) if they have no sample points in common — . Then , and the addition law simplifies to:
Example: A single coin toss — events “heads” and “tails” are mutually exclusive.
In Venn diagram terms, mutually exclusive events are non‑overlapping circles.
Key takeaways – Probability Basics
- Probability ranges from 0 to 1; .
- Complement: .
- Union probability includes all outcomes in either event; intersection probability includes only common outcomes.
- Addition law: .
- For mutually exclusive events, so .
Bayes’ Theorem
Bayes’ Theorem links prior probability (initial belief about an event) to posterior probability (updated belief after observing new evidence). Intuitively: when you learn that event has occurred, your estimate of the probability of event may change. Bayes’ Theorem quantifies exactly how.
Conditional Probability
The probability of given that has occurred is conditional probability:
Similarly,
Multiplication Law
The multiplication law expresses the intersection probability in two equivalent ways:
These relations are the building blocks of Bayes’ Theorem.
Relating Prior and Posterior
Given a prior and new information , the posterior is:
Interpretation: The posterior is proportional to the likelihood of the evidence given times the prior, normalized by the total probability of .
Worked Example (Movie Hits)
- = movie KGF is a hit, prior
- = movie Kantara is a hit,
- Joint
Compute posterior :
The posterior (0.30) is lower than the prior (0.34) – knowing Kantara was a hit reduced the chance of KGF being a hit.
Compute :
Again, the posterior (0.53) is lower than the prior of (0.60).
Joint Probability Table
A joint probability table arranges all intersection probabilities for quick computation of conditionals and marginals.
| (Kantara hit) | (not hit) | Marginal | |
|---|---|---|---|
| (KGF hit) | 0.18 | 0.16 | 0.34 |
| (not hit) | 0.42 | 0.24 | 0.66 |
| Marginal | 0.60 | 0.40 | 1.00 |
- Marginal probabilities are in the last row/column (e.g., , ).
- To find , divide the joint in row (0.18) by the row marginal (0.34).
Exam tip: Always check whether you are conditioning on a row or a column. The denominator is always the marginal of the conditioning event.
Independence of Events
Two events and are independent if the occurrence of one does not affect the probability of the other:
Equivalently:
Mutually exclusive events with non-zero probability cannot be independent (if one happens, the other cannot, so ).
- In the movie example: , . Since , the events are not independent.
Law of Total Probability
The law of total probability expresses the total probability of as a sum over mutually exclusive and collectively exhaustive events:
Generalized to events that partition the sample space:
Example: For movie KGF being a hit, consider three revenue outcomes for Kantara (=₹4 lakhs, =₹5 lakhs, =₹6 lakhs). Then .
Worked Example: Jyothi Hegde’s Placement Data
Categorize respondents by work experience (: ≤2 years or >2 years) and salary (: ≤₹25 lakhs or >₹25 lakhs). Total .
| Row total | |||
|---|---|---|---|
| 34 | 13 | 47 | |
| 22 | 34 | 56 | |
| Column total | 56 | 47 | 103 |
Convert to probabilities (divide by 103):
| Marginal | |||
|---|---|---|---|
| 0.330 | 0.126 | 0.456 | |
| 0.214 | 0.330 | 0.544 | |
| Marginal | 0.544 | 0.456 | 1.000 |
Compute probabilities:
Key takeaways
- Prior → posterior: update probability when new information arrives.
- Conditional probability: .
- Bayes’ Theorem: .
- Use a joint probability table to organize and compute conditionals and marginals.
- Independence is tested via ; mutually exclusive events are dependent.
- Law of total probability sums over partitions to find .