Key Concepts: Continuous vs. Discrete
A continuous random variable can take any value within an interval (finite or infinite), including fractional values. Examples: time between customer arrivals, fluid volume in a bottle, oven temperature. Unlike discrete variables (e.g., number of defects, number of customers), continuous variables have a probability density function (pdf) , not a probability mass function.
| Feature | Discrete | Continuous |
|---|---|---|
| Possible values | Countable (integers) | Uncountable (any real in an interval) |
| Probability at a point | (the mass) | (area under a point is zero) |
| Probability for an interval | Sum of over values in | Integral of from to |
| Total probability |
Probability Density Function (PDF) and Cumulative Distribution Function (CDF)
For a continuous random variable :
- pdf: for all , and .
- CDF: .
- Tail probability: .
The probability that lies in is the area under between and :
Exam tip: For a continuous random variable, . Always think in terms of intervals.
Expectation and Variance
- Expectation (mean): (replace sum with integral).
- Variance: .
Formulas mirror the discrete case, substituting summation with integration.
Uniform Distribution
A random variable is uniformly distributed over if every value in that interval is equally likely. The pdf is constant:
Key formulas
| Quantity | Expression |
|---|---|
| CDF | |
| Tail probability | |
| Probability in | |
| Expectation | |
| Variance |
Worked Example: Driving Time
Nilesh Shah’s driving time from Vadodara to Ahmedabad is uniformly distributed between minutes and minutes.
- Expected time: minutes.
- : (75% chance of 130 minutes or less).
- : (87.5% chance of more than 105 minutes).
- : (any single point has zero probability for a continuous variable).
Note: Probabilities for intervals (e.g., ) can be computed similarly.
Key takeaways
- Continuous random variables have a pdf; probabilities are areas under the curve.
- The CDF gives ; tail probability is .
- For continuous variables, for any specific .
- Expectation and variance use integrals instead of sums.
- Uniform distribution: constant pdf over ; , .
- In a uniform model, probability is proportional to interval length.
Uniform Distribution Example
The uniform distribution models a random variable that is equally likely to take any value within a given interval . Intuitively: every value in that range has the same chance—there is no “peaked” area.
Worked Example: Battery Life of a Phone
Meena’s phone battery lasts between 8 and 12 hours before needing a charge. Let = time (in hours) until recharge. Model: .
PDF (not needed for calculations here):
CDF for any in :
1. Probability that recharge needed within 9 hours ()
25% chance the phone needs charging in the first 9 hours.
2. Probability that battery lasts at least 11 hours ()
25% chance the battery lasts 11 hours or more.
3. 80th percentile – the value such that .
There is an 80% chance the phone will need charging within 11.2 hours.
Exam tip: For uniform distributions, percentiles are a straight linear interpolation: .
Discrete Uniform – Reminder
The discrete uniform distribution assigns equal probability to a finite set of values (e.g., a fair die: for ). The continuous version is its natural extension to an interval.
Key takeaways (Uniform)
- : all values equally likely in .
- CDF: for .
- Percentile : .
- Used when only the minimum and maximum are known and nothing else.
Exponential Distribution
The exponential distribution models the time between random events that occur at a constant average rate (events per unit time). Intuitively: if events happen “on average every 1/λ time units”, the waiting time until the next event follows an exponential distribution.
Formal Definition
Let = time between events. .
Probability density function (PDF):
Cumulative distribution function (CDF):
Tail probability (survival function):
Probability between and ():
Mean and Variance
| Property | Formula | Note |
|---|---|---|
| Mean | Average waiting time | |
| Variance | ||
| Standard deviation | Equal to the mean – a unique property |
The mean equals the standard deviation for the exponential distribution. This is a quick check: if the sample mean and sample SD are very different, the data may not be exponential.
Key Properties
- Skewed right – the PDF peaks at and decays; most probability lies near zero.
- Memoryless property – . The future waiting time does not depend on how long you have already waited.
- Connection to Poisson: If events occur at rate per unit time (Poisson process), the waiting time between consecutive events is exponential with mean , and the number of events in a fixed interval is Poisson with mean .
Applications
- Time to complete a manual task (assembly, phone call)
- Lifetime of electronic components (bulbs, circuits)
- Time until equipment requires repair
- Waiting time in a queue (call center, customer service)
Worked Example: Call Center Waiting Time
Megha’s waiting time (in minutes) for a service desk agent is exponential with mean 2 minutes. So per minute.
1. Probability she waits less than 1 minute ()
2. Probability she waits more than 3 minutes ()
3. Probability she waits between 1 and 3 minutes ()
4. 90th percentile – find such that :
90% of calls are answered within about 4.6 minutes.
Computation in Excel
Use the function =EXPON.DIST(x, λ, cumulative):
cumulative = FALSEreturns the PDF .cumulative = TRUEreturns the CDF .
For inverse calculations (given probability, find ), Excel provides =EXPON.INV(probability, λ) or use the formula for the -th percentile.
Key takeaways (Exponential)
- Models waiting time between random events at constant rate λ.
- PDF: , CDF: .
- Mean = Standard deviation = .
- Memoryless: .
- Connected to Poisson: number of events per interval ~ Poisson(λt), waiting time ~ Exponential(λ).
- Use
EXPON.DISTin Excel for probabilities;EXPON.INVfor percentiles.
Memoryless Property of the Exponential Distribution
The memoryless property is a defining characteristic of the exponential distribution. Intuitively: if you are waiting for an event that follows an exponential distribution, the probability of waiting an additional units does not depend on how long you have already waited — the process "forgets" the elapsed time.
Formal definition
Let , with mean . For any and ,
Equivalently, the remaining time conditional on has the same exponential distribution as :
Derivation
The tail probability of an exponential is .
Thus has the same distribution as .
Exam tip: The memoryless property is unique to the exponential distribution (and its discrete analogue, the geometric). It is the reason exponential waiting times are used in queueing models: the future waiting time is independent of the past.
Worked example — Shanti's bus wait
Shanti waits for a bus. Waiting time (minutes) is exponential with mean minutes.
| Question | Computation | Result |
|---|---|---|
| Probability she waits min | ||
| Probability she waits min | ||
| Probability she waits between and min | ||
| Given she already waited min, probability she waits another min | ||
| Given she already waited min, probability she waits another min |
The last two results are identical — the waiting time resets regardless of elapsed time, illustrating the memoryless property.
Key takeaways
- Memoryless property: .
- Equivalent to the remaining time having the same exponential distribution.
- Only the exponential (continuous) and geometric (discrete) have this property.
- Conditional waiting times are computed using — no dependence on .
- Always check that the random variable is exponential before applying memorylessness.
Relation Between Exponential and Poisson Distributions
The exponential distribution is the waiting time distribution that underlies a Poisson process. If events occur according to a Poisson process with rate (average number of events per unit time or space), then:
- The number of events in a fixed interval of length is .
- The time between consecutive events (interarrival times) is .
This duality allows us to switch between counting problems and waiting‑time problems.
| Process | Interarrival times | Counting (number of events in ) |
|---|---|---|
| Deterministic | Fixed, constant | Fixed number |
| Poisson | Exponential() | Poisson() |
Example — highway defects
Recall the Solapur highway example: defects occur as a Poisson process with average defects per kilometer.
- Distance between defects is with mean km.
- Number of defects in km is .
| Question | Computation | Result |
|---|---|---|
| Probability of at least km without a defect | ||
| Probability of at least km without a defect |
The same result can be obtained from the Poisson: .
Exam tip: Whenever you see a problem about “time between events” or “distance between occurrences”, check whether the process is Poisson. If yes, the interarrival time is exponential — use the exponential tail formula directly.
Key takeaways
- In a Poisson process, interarrival times are i.i.d. exponential with the same rate .
- Average interarrival time = ; average number of events per unit time = .
- Switching between Poisson and exponential: .
- Memoryless property of the exponential explains the “lack of clustering” in a Poisson process — the hazard rate is constant.
Worked Example: Oil Change Service Time — Exponential & Poisson
Scenario: Rohit’s garage offers a 50% discount if an oil change takes more than 20 minutes. The average oil change time is 20 minutes. Assuming the time (in minutes) for one oil change follows an exponential distribution, we evaluate the risk of this promise. Later we shift to counting the number of oil changes in a fixed time window, which follows a Poisson distribution.
1. Time for a single oil change: Exponential model
If the mean time is minutes, the rate parameter is per minute.
The exponential CDF: .
Key relationship: Exponential distribution models the time between events (here, the time to complete one oil change).
Probability that time < 20 minutes (promise kept)
There is a 63.2% chance the oil change finishes within 20 minutes. Rohit will likely give many discounts.
Probability that time is between 15 and 30 minutes
There is a 24.9% chance the oil change takes between 15 and 30 minutes.
| Event | Probability |
|---|---|
| 0.632 | |
| 0.249 |
2. Number of oil changes in a time window: Poisson model
Number of oil changes in a fixed time period follows a Poisson distribution with mean , where is the exponential rate (0.05 per minute). This connection holds because the exponential distribution models inter-arrival times of a Poisson process.
Four‑hour morning shift (240 minutes)
Probability of more than 15 oil changes in 4 hours:
. Using Poisson tables or software:
Eight‑hour day (480 minutes)
Probability of more than 30 oil changes in 8 hours:
| Time window | Average oil changes () | Threshold | |
|---|---|---|---|
| 4 hours (morning) | 12 | 0.228 | 15 |
| 8 hours (full day) | 24 | 0.132 | 30 |
Relationship between Exponential and Poisson
The Poisson distribution counts events when the inter‑event times are i.i.d. exponential with the same rate .
Key Takeaways
- Exponential models the time until one event (oil change). Use .
- Poisson models the number of events in a fixed interval when events occur at a constant average rate.
- The two are linked: .
- Rohit’s offer is risky: 63% chance of finishing on time (but 37% chance of discount). The probabilities of completing many oil changes per shift are modest (22.8% for >15 in 4 hours, 13.2% for >30 in 8 hours).
-
Exam tip: When a problem gives an average time and asks for counts over a time period, switch from exponential to Poisson using . Always check if the exponential inter‑event time assumption is justified (often stated as “random arrivals” or “memoryless” process).
The Normal Random Variable and its Distribution
The normal distribution is the most widely used continuous probability distribution. It models many naturally occurring phenomena—heights, weights, test scores, rainfall—and is central to statistical inference. Its probability density function (PDF) produces the familiar bell-shaped curve, symmetric about the mean.
The entire family of normal distributions is differentiated by two parameters:
- Mean – determines the centre; can be any real number.
- Standard deviation – determines the spread; .
Properties of the Normal Curve
- Highest point at , which is also the median and mode.
- Symmetric about : left half is a mirror image of the right half. Skewness = 0.
- Tails extend to infinity in both directions, never touching the horizontal axis.
- Total area under curve = 1 (like all PDFs). Because of symmetry, area to the left of = 0.5; area to the right = 0.5.
- Spread determined by : larger → wider, flatter curve; smaller → taller, narrower curve.
The Empirical Rule (68–95–99.7 Rule)
For any normal random variable , the percentage of values within standard deviations of the mean is fixed:
| Interval | Percentage of values | Probability |
|---|---|---|
| 68.3% | ||
| 95.4% | ||
| 99.7% |
Exam tip: The empirical rule provides quick probability approximations and is often tested directly.
The Standard Normal Distribution
A normal distribution with and is called the standard normal distribution. Its random variable is denoted by instead of .
The cumulative distribution function (CDF) of is often written as (capital phi), and the PDF as . Probability tables for eliminate the need for integration.
Computing Normal Probabilities with Excel
Two built-in functions handle any normal distribution directly (no need to standardise):
-
NORM.DIST
cumulative = TRUE→ returns (CDF).cumulative = FALSE→ returns the PDF value at .
-
NORM.INV
- Returns such that (inverse CDF).
Template example (general normal: , )
| Input | Computed Probability | Result |
|---|---|---|
| 0.841 | ||
| 0.159 | ||
| 0.683 |
For inverse:
- Enter cumulative probability 0.841 → get .
- Enter interval probability 0.683 → get , .
Worked Examples: Standard Normal Probabilities
Set , in the template.
Example 1
| Probability | Result |
|---|---|
| 0.159 | |
| 0.841 | |
| 0.933 | |
| 0.994 | |
| 0.499 |
Note: for any continuous random variable, so and .
Example 2
| Probability | Result |
|---|---|
| 0.288 | |
| 0.433 | |
| 0.345 | |
| 0.579 | |
| 0.885 | |
| 0.242 |
Example 3: Inverse Calculations
| Condition | Range of |
|---|---|
| Area to the left of is 0.21 | (i.e., ) |
| Area between and is 0.9 | |
| Area between and is 0.2 | |
| Area to the left of is 0.955 | (i.e., ) |
| Area to the right of is 0.692 | (i.e., ) |
Key Takeaways
- The normal distribution is bell-shaped, symmetric about , with median mode.
- It is parameterised by (centre) and (spread).
- The empirical rule: ≈68% within 1σ, ≈95% within 2σ, ≈99.7% within 3σ.
- The standard normal has , ; its CDF is .
- Excel:
NORM.DISTfor cumulative/PDF,NORM.INVfor inverse; both work for any normal. - For continuous variables, point probabilities are zero; probabilities always refer to intervals.
Standard Normal Distribution
The standard normal distribution is the symmetric, bell-shaped distribution centered at 0 with a standard deviation of 1. Its random variable is denoted . This distribution serves as the universal reference for all normal distributions — any normal variable can be transformed into a standard normal.
An important concept for statistical inference is the critical value. For a standard normal, a critical value is defined such that the tail probability to the right of it equals :
By symmetry, has a left‑tail probability of . Consequently, the central probability between and is :
The familiar 68‑95‑99.7 rule is a special case: for , , and .
Common critical values
| (confidence level) | ||
|---|---|---|
| 0.80 | 0.10 | 1.28 |
| 0.90 | 0.05 | 1.64 |
| 0.95 | 0.025 | 1.96 |
| 0.99 | 0.005 | 2.58 |
Exam tip: Memorise the three critical values 1.64, 1.96, and 2.58 — they appear repeatedly in confidence intervals and hypothesis tests.
Two standard types of questions
- Given , find probability — use the standard normal table or template.
- Given probability, find — inverse lookup. Sketching the bell curve and shading the relevant area always helps avoid sign errors.
Converting any normal to the standard normal
If , then the transformed variable
follows a standard normal distribution. measures how many standard deviations is away from the mean — positive above, negative below.
Worked example
Let . Find .
- Convert boundaries:
- Then .
- From the standard normal table, .
The same result could be obtained by directly using a normal template with , ; the conversion simply standardises the problem.
Key takeaways
- Critical value : right‑tail area = ; symmetric about 0.
- Central probability lies between and .
- Common values: 1.28, 1.64, 1.96, 2.58 for confidence levels 80%, 90%, 95%, 99%.
- Any normal can be transformed via to a standard normal.
- tells the distance from the mean in standard deviation units.
- Always sketch the curve and shade the area when solving probability problems.
Examples of Normal Distribution Applications
The normal distribution is a versatile model for many real-world variables. These worked examples illustrate how to compute probabilities for intervals, tails, and cumulative regions, and how to perform inverse calculations (percentiles) using given mean and standard deviation . The key tool is the z‑score: , which maps any to a standard normal .
Exam tip: Always check units (e.g., minutes vs. hours). A small mistake in input changes the entire answer.
1. Stock Returns (Munaf)
A stock fund has a mean annual return of and a standard deviation . Returns are normally distributed.
- Unsatisfactory (return ):
→ (1% chance). - Excellent (return ):
→ (74.8% chance). - Moderate (return between 5% and 10%):
(24.3% chance). - 90th percentile (inverse):
- Cumulative: 90% of returns are less than 15.8%.
- Tail: 90% of returns are greater than 8.1%.
- Interval: 90% of returns lie between 7% and 17%.
Key mechanism: The normal distribution template directly uses and to output these probabilities without manual z‑score calculation.
2. Screen Time (Bhavani)
Daily screen time of school kids: hours, hours.
- More than 4 hours:
→ (96% of kids exceed 4 hours). - Between 6 and 12 hours:
, → (75.7% of kids). - Top 20% (inverse):
Need → hours. A kid with >10.5 hours is in the top 20%. - Bottom 20% (inverse):
Need → hours. A kid with <6.3 hours is in the bottom 20%. - (Bonus) 80% interval: 80% of kids have screen time between 5.2 and 11.6 hours.
3. Exam Duration (Hirav)
Time to complete an exam: minutes, minutes. 60 students.
- 2 hours or less (≤120 min):
→ → students. - Between 2 and 3 hours (120–180 min):
, → → students. - 3 hours or more (≥180 min):
→ students.
Exam tip: The symmetry of the normal distribution means probabilities at equal distances above and below the mean are identical. Here 120 and 180 min are both 1.5σ from the mean, giving the same tail probability (0.067).
4. Sugar Packet Weight (Naman)
Two filling machines; weights are normally distributed. Naman evaluates which machine is better.
| Machine | (kg) | (kg) | kg) | 90% interval | |
|---|---|---|---|---|---|
| 1 | 1.01 | 0.02 | 0.309 (30.9%) | 0.669 (66.9%) | [0.97, 1.04] |
| 2 | 1.03 | 0.04 | 0.227 (22.7%) | 0.465 (46.5%) | [0.96, 1.09] |
- Machine 1 has fewer underweight? No, 30.9% > 22.7%, so Machine 2 produces fewer underweight packets.
- Machine 2 has a narrower range for 90% of packets? No, its interval is wider (0.96–1.09 vs. 0.97–1.04).
Naman’s claim that 90% of packets lie between 0.99 and 1.01 is false for both machines.
Trade-off: Machine 2 reduces underweight risk but increases variability, making the middle‑90% range much broader.
5. Glass Thickness (Vibhor)
Glass sheet thickness: mm, mm.
- Thickness < 2.9 mm (cumulative):
→ (20.2% of sheets too thin). - Thickness > 3.2 mm (tail):
→ (4.8% too thick). - 95% thickness range (inverse):
For a central probability of 0.95, the endpoints are → [2.76 mm, 3.24 mm].
Key takeaways
- All calculations rely on the z‑score transformation , linking any normal variable to the standard normal.
- Three standard probability queries: cumulative (less than), tail (greater than), and interval (between). Use the CDF differences.
- Inverse problems (percentiles) find the value for a given cumulative probability (e.g., top 20% → ).
- Always check units (hours vs. minutes, percentage vs. decimal) before plugging numbers into a template or formula.
- Symmetry simplifies: .
Normal Approximation to the Binomial Distribution
When n is large, computing binomial probabilities directly (via factorials in ) becomes unwieldy. The normal distribution provides a remarkably accurate approximation because the binomial’s probability mass function is approximately bell‑shaped — especially when p is not too close to 0 or 1.
Conditions for a Good Approximation
The approximation works well when both:
If these hold, match the mean and variance of the normal to those of the binomial:
So we use to approximate .
The Continuity Correction
Because the binomial is discrete and the normal is continuous, a continuity correction of ±0.5 is applied to improve accuracy.
| Desired binomial probability | Continuity‑corrected normal probability |
|---|---|
Exam tip: The continuity correction is the most common mistake. Always add/subtract 0.5; never approximate a discrete probability using a single point from a continuous distribution (that probability would be 0).
Worked Examples
Example 1: Unbiased coin, 16 tosses
→ .
-
Exact binomial:
Normal approx: -
Exact binomial:
Normal approx:
Example 2: Invoices with errors (Kunal Poonawala, Junagadh)
→ .
-
Exact binomial:
Normal approx: (close, though slightly off; approximation improves with larger ). -
Exact binomial:
Normal approx:
The approximation is reliable, especially when and are well above 5.
Connection to the Poisson Distribution
The Poisson distribution can also be approximated by the normal when its mean is large. Moreover, for a binomial with small (failure probability), the variance , making the binomial similar to a Poisson. Hence, in certain cases the Poisson can be approximated via the binomial → normal chain.
Key Takeaways
- Normal approximation to the binomial is valid when and .
- Match mean and variance .
- Always apply a continuity correction (±0.5) because the binomial is discrete.
- The approximation is remarkably accurate and is often used by software internally.
Linear Combinations of Random Variables (Normal Case)
Any linear combination of normally distributed random variables is itself normally distributed – not just its mean and variance, but the entire distribution.
General Results (any distribution)
Given random variables and constants , define
Expectation (always linear):
Variance:
-
If the are independent:
-
If the are correlated (not independent), covariance terms appear:
These formulas hold for any distributions of the .
Special Case: Normally Distributed
If each and they are jointly normal (or independent normal), then is also normally distributed:
This property is unique to the normal family and is crucial for statistical inference (e.g., sums of normal data remain normal).
Exam tip: When a problem states “ are independent normal random variables,” any linear combination (like or ) is also normal. You only need to compute its mean and variance.
Key Takeaways
- For any random variables, is linear; variance adds with squares plus covariance terms if not independent.
- For normal random variables, the linear combination is also normal — a key result for later use (e.g., Central Limit Theorem).
- This property does not hold for most other distributions; it is special to the normal.
Linear Combinations of Normal Random Variables
A key property: any linear combination of independent normal random variables is itself normally distributed. This makes it possible to model sums, differences, and averages of normal data with a single normal distribution.
The Basic Result: (Y = aX + b)
If (X \sim N(\mu,\ \sigma^2)) and (Y = aX + b) (with constants (a) and (b)), then
[
Y \sim N\big(a\mu + b,\ a^2\sigma^2\big).
]
Intuition – Shifting and scaling a normal distribution preserves its bell shape; only the mean and variance change linearly.
Standardization as a Special Case
Set (a = \frac{1}{\sigma}) and (b = -\frac{\mu}{\sigma}):
[ Y = \frac{X - \mu}{\sigma} \sim N(0,1) ]
This is the standard normal random variable (Z).
Sum of Two Independent Normals
Let (X_1 \sim N(\mu_1, \sigma_1^2)) and (X_2 \sim N(\mu_2, \sigma_2^2)) be independent. Then
[
Y = X_1 + X_2 \sim N(\mu_1 + \mu_2,\ \sigma_1^2 + \sigma_2^2).
]
If the variables are not independent, the sum is still normal, but the variance includes covariance terms.
General Linear Combination
For (n) independent normal variables (X_i \sim N(\mu_i, \sigma_i^2)) and constants (a_1, \dots, a_n, b):
[ Y = a_1 X_1 + a_2 X_2 + \cdots + a_n X_n + b ]
has distribution
[ Y \sim N!\left( \sum_{i=1}^n a_i \mu_i + b,\ \sum_{i=1}^n a_i^2 \sigma_i^2 \right). ]
Important Special Cases: Sum and Sample Mean
Assume (X_1, X_2, \dots, X_n) are iid (independent and identically distributed) as (N(\mu, \sigma^2)).
| Quantity | Definition | Distribution | Mean | Variance | Standard Deviation |
|---|---|---|---|---|---|
| Sum | (S_n = \sum_{i=1}^n X_i) | (N(n\mu,\ n\sigma^2)) | (n\mu) | (n\sigma^2) | (\sqrt{n},\sigma) |
| Sample mean | (\bar{X} = \frac{1}{n}\sum_{i=1}^n X_i) | (N(\mu,\ \sigma^2/n)) | (\mu) | (\frac{\sigma^2}{n}) | (\frac{\sigma}{\sqrt{n}}) |
Exam tip: Standard deviations do not add up directly. For the sum, (\text{SD}(S_n) = \sqrt{n},\sigma); for the average, (\text{SD}(\bar{X}) = \sigma/\sqrt{n}). The variance reduces by a factor of (n) when averaging – this is the foundation of sampling precision.
Worked Examples
1. Piston–Cylinder Gap
Problem – Piston head radius (X_1 \sim N(30\ \text{mm},\ 0.0025\ \text{mm}^2)), cylinder inside radius (X_2 \sim N(30.25\ \text{mm},\ 0.0036\ \text{mm}^2)). The gap is (Y = X_2 - X_1). Find
(a) (P(Y \le 0)) (piston does not fit)
(b) (P(0.1 \le Y \le 0.35)) (optimal performance)
Solution
(Y) is normal because it is a linear combination of independent normals.
[ \begin{aligned} \mu_Y &= 30.25 - 30 = 0.25 \ \sigma_Y^2 &= 0.0025 + 0.0036 = 0.0061 \ \sigma_Y &= \sqrt{0.0061} \approx 0.078 \end{aligned} ]
(a)
[
P(Y \le 0) = \Phi!\left(\frac{0 - 0.25}{0.078}\right) \approx \Phi(-3.205) \approx 0.0007
]
(b)
[
P(0.1 \le Y \le 0.35) = \Phi!\left(\frac{0.35-0.25}{0.078}\right) - \Phi!\left(\frac{0.1-0.25}{0.078}\right) \approx \Phi(1.282) - \Phi(-1.923) \approx 0.872
]
Interpretation – Extremely low chance of misfit; ~87 % chance of optimal gap.
2. Average Height of Cotton Plants
Problem – Individual plant height after two weeks: (X_i \sim N(29.4\ \text{cm},\ 4.41\ \text{cm}^2)). For 20 plants, the sample mean (\bar{X}) is normal. Find the 95th percentile of (\bar{X}).
Solution
[
\bar{X} \sim N!\left(29.4,\ \frac{4.41}{20} = 0.2205\right),\quad \sigma_{\bar{X}} = \sqrt{0.2205} \approx 0.469
]
The 95th percentile is
[
\mu + z_{0.05},\sigma_{\bar{X}} = 29.4 + 1.645 \times 0.469 \approx 30.17\ \text{cm}.
]
Notice that the standard deviation of the average (0.47 cm) is much smaller than that of an individual plant (2.1 cm) – averaging greatly reduces variability.
3. Stock Returns Comparison
Problem – Stock A: (X_A \sim N(8%,\ 2.25%^2)), Stock B: (X_B \sim N(9.5%,\ 4%^2)), independent.
Define (Y = X_B - X_A). Find:
(a) (P(X_A \text{ moderate})) meaning (5 \le X_A \le 10)
(b) (P(X_B \text{ excellent})) meaning (X_B \ge 10)
(c) (P(X_B > X_A)) i.e., (P(Y > 0))
(d) (P(Y \ge 2))
Solution
(a) (X_A \sim N(8, 2.25)).
[
P(5 \le X_A \le 10) = \Phi!\left(\frac{10-8}{1.5}\right) - \Phi!\left(\frac{5-8}{1.5}\right) \approx \Phi(1.333) - \Phi(-2) \approx 0.886
]
(b) (X_B \sim N(9.5, 4)).
[
P(X_B \ge 10) = 1 - \Phi!\left(\frac{10-9.5}{2}\right) \approx 1 - \Phi(0.25) \approx 0.4013
]
(c) (Y = X_B - X_A \sim N(1.5,\ 2.25+4=6.25)), so (\sigma_Y = 2.5).
[
P(Y > 0) = 1 - \Phi!\left(\frac{0-1.5}{2.5}\right) = 1 - \Phi(-0.6) \approx 0.726
]
(d)
[
P(Y \ge 2) = 1 - \Phi!\left(\frac{2-1.5}{2.5}\right) = 1 - \Phi(0.2) \approx 0.421
]
Key Takeaways
- Any linear combination of independent normal random variables is normally distributed.
- For (Y = \sum a_i X_i + b): mean = (\sum a_i \mu_i + b), variance = (\sum a_i^2 \sigma_i^2) (independence assumed).
- Sum of iid normals: mean (n\mu), variance (n\sigma^2).
- Sample mean of iid normals: mean (\mu), variance (\sigma^2/n) – precision improves with (n).
- Difference of two independent normals: variance adds; subtract the means.
- Use standardization to compute probabilities for any linear combination.
Chi-Square Distribution
The chi-square distribution arises naturally from squared standard normal variables. If , then follows a chi-square distribution with 1 degree of freedom. More generally, if are independent standard normal random variables, then
follows a chi-square distribution with degrees of freedom (denoted ).
Intuitively, it measures the sum of squared deviations; because squares are always non‑negative, the distribution is defined only for positive values.
Properties
- Single parameter: the degrees of freedom ().
- Mean: .
- Variance: .
- Shape: right‑skewed for small ; as increases, the skewness decreases and the density approaches a normal distribution (by the Central Limit Theorem).
- Support: .
Applications
- Non‑parametric hypothesis tests (e.g., goodness‑of‑fit).
- Feature selection and classification.
- Testing independence in contingency tables (logistic regression).
Critical Values
The chi‑square critical value is the value such that the right‑tail probability equals :
Unlike the standard normal, critical values depend on the degrees of freedom. The table below gives critical values for several and (computed using Microsoft Excel’s CHISQ.DIST function).
| 5 | 7.3 | 9.2 | 11.1 | 15.1 |
| 10 | 13.4 | 16.0 | 18.3 | 23.2 |
| 15 | 19.3 | 22.3 | 25.0 | 37.6 |
| 20 | 25.0 | 28.4 | 31.4 | 37.6 |
Exam tip: As increases, critical values increase because the distribution spreads out (variance ). The values in the last column are the largest because the tail probability is smallest.
Computing in Excel
Use CHISQ.DIST(x, k, cumulative):
x: value of the chi‑square random variable.k: degrees of freedom.cumulative:TRUEreturns the cumulative distribution function;FALSEreturns the probability density function.
A template can automate mean, variance, probability calculations, and inverse (critical value) look‑up.
Key takeaways
- Chi‑square is the sum of independent squared standard normals.
- Mean = , variance = .
- Always positive, right‑skewed, converges to normal as .
- Critical values increase with ; used in goodness‑of‑fit and independence tests.
Student’s t‑Distribution
The Student’s t‑distribution (or simply t‑distribution) is defined when a standard normal variable is divided by the square root of an independent chi‑square variable scaled by its degrees of freedom:
where , , and and are independent. follows a t‑distribution with degrees of freedom.
Intuitively, the t‑distribution is used in place of the standard normal when the population variance is unknown and the sample size is small. It has fatter tails (more probability in the extremes) than the normal.
History
Discovered by William Sealy Gosset while working at the Guinness Brewery. He published under the pseudonym “Student” because his employer prohibited employees from publishing research. Hence the name Student’s t‑distribution.
Properties
- Symmetric and bell‑shaped, centered at 0.
- Mean: (for ).
- Variance: (for ; undefined for ).
- Shape: flatter (more spread) than the standard normal; as , the t‑distribution converges to .
- Support: .
Applications
- Hypothesis tests about population means (especially small samples).
- Comparing means of two populations.
- Diagnostics in linear regression (e.g., t‑tests for coefficients).
Critical Values
Because the t‑distribution is symmetric, critical values are often defined for two‑tailed regions. The value is the positive number such that:
By symmetry, . Therefore .
The table below gives for common and degrees of freedom (computed using Excel’s T.DIST function).
| 5 | 1.48 | 2.02 | 2.57 | 4.03 |
| 10 | 1.37 | 1.81 | 2.23 | 3.17 |
| 15 | 1.34 | 1.75 | 2.13 | 2.95 |
| 20 | 1.33 | 1.73 | 2.09 | 2.85 |
Notice that as increases, the critical values decrease (the distribution becomes less flat) and approach the corresponding standard normal critical values (e.g., for , ).
Computing in Excel
Use T.DIST(x, k, cumulative):
x: value of the t random variable.k: degrees of freedom.cumulative:TRUEfor cumulative probability,FALSEfor PDF.
For critical values, the inverse function T.INV or a template can be used.
Exam tip: The t‑distribution is flatter than the normal, so its critical values are larger than the corresponding normal critical values for small . As grows, they converge.
Key takeaways
- t‑distribution: .
- Symmetric, zero‑mean, fatter tails than normal.
- Variance ; converges to standard normal as .
- Critical values decrease with increasing ; used in mean hypothesis tests and regression diagnostics.
The F Distribution
The F distribution (Fisher–Snedecor distribution, after Ronald Fisher) models the ratio of two independent chi-square random variables, each divided by its own degrees of freedom.
Here = numerator degrees of freedom and = denominator degrees of freedom. The order matters: .
Properties
- Defined only for positive values of the random variable.
- Unimodal and right-skewed (long right tail).
- Mean: for (when , the mean is undefined).
- Standard deviation decreases as increases.
- As both and become large, the PDF becomes sharply spiked around .
Where it is used
- Analysis of variance (ANOVA)
- Feature selection in machine learning
- Testing overall fit of multiple linear regression models
Excel Function: F.DIST
In Excel, the function F.DIST(x, k1, k2, cumulative) computes probabilities.
| Parameter | Meaning |
|---|---|
x | Value of the F random variable |
k1 | Numerator degrees of freedom |
k2 | Denominator degrees of freedom |
cumulative | TRUE → cumulative probability; FALSE → probability density |
Inverse (critical values) can be obtained via F.INV(alpha, k1, k2) for right‑tail probability .
Worked Example: Critical Values (Right‑Tail)
For an F‑distribution with given , the critical value is the value such that the area to its right equals .
Using an F‑distribution template (or Excel), with :
| (5, 5) | 2.23 | 3.45 | 5.05 | 10.97 |
| (5, 10) | 1.80 | 2.52 | 3.33 | 5.64 |
| (10, 5) | 2.19 | 3.30 | 4.74 | 10.05 |
| (20, 20) | 1.47 | 1.79 | 2.21 | 2.94 |
Exam tip: Swapping and changes critical values — always double‑check which is numerator and which is denominator. When is large, critical values approach 1 because the distribution becomes concentrated near 1.
Key takeaways
- — ratio of independent scaled chi‑squares.
- Positive, unimodal, right‑skewed; mean for .
- Order of degrees of freedom matters: .
- As increases, variance decreases and distribution peaks nearer 1.
- Widely used in ANOVA, regression, and machine learning feature selection.
- Excel:
F.DIST(x, k1, k2, TRUE/FALSE)for cumulative/density;F.INVfor critical values.