Term 1 · Module 3 of 8

Continuous Probability Distributions

Business Statistics for Entrepreneurs

Key Concepts: Continuous vs. Discrete

A continuous random variable can take any value within an interval (finite or infinite), including fractional values. Examples: time between customer arrivals, fluid volume in a bottle, oven temperature. Unlike discrete variables (e.g., number of defects, number of customers), continuous variables have a probability density function (pdf) f(x)f(x), not a probability mass function.

FeatureDiscreteContinuous
Possible valuesCountable (integers)Uncountable (any real in an interval)
Probability at a pointP(X=x)=f(x)P(X=x)=f(x) (the mass)P(X=x)=0P(X=x)=0 (area under a point is zero)
Probability for an intervalSum of f(x)f(x) over values in [a,b][a,b]Integral of f(x) dxf(x)\,dx from aa to bb
Total probability∑f(x)=1\sum f(x)=1∫f(x) dx=1\int f(x)\,dx=1

Probability Density Function (PDF) and Cumulative Distribution Function (CDF)

For a continuous random variable XX:

  • pdf: f(x)≥0f(x) \ge 0 for all xx, and ∫−∞∞f(x) dx=1\int_{-\infty}^{\infty} f(x)\,dx = 1.
  • CDF: F(x)=P(X≤x)=∫−∞xf(t) dtF(x) = P(X \le x) = \int_{-\infty}^{x} f(t)\,dt.
  • Tail probability: Fˉ(x)=P(X>x)=1−F(x)\bar{F}(x) = P(X > x) = 1 - F(x).

The probability that XX lies in [a,b][a,b] is the area under f(x)f(x) between aa and bb:

P(a≤X≤b)=∫abf(x) dx=F(b)−F(a).P(a \le X \le b) = \int_{a}^{b} f(x)\,dx = F(b) - F(a).

Exam tip: For a continuous random variable, P(X=any single value)=0P(X = \text{any single value}) = 0. Always think in terms of intervals.

Expectation and Variance

  • Expectation (mean): E[X]=∫−∞∞x f(x) dxE[X] = \int_{-\infty}^{\infty} x\,f(x)\,dx (replace sum with integral).
  • Variance: Var(X)=∫−∞∞(x−E[X])2 f(x) dx\text{Var}(X) = \int_{-\infty}^{\infty} (x - E[X])^2\,f(x)\,dx.

Formulas mirror the discrete case, substituting summation with integration.


Uniform Distribution

A random variable is uniformly distributed over [L,U][L, U] if every value in that interval is equally likely. The pdf is constant:

f(x)=1U−L,L≤x≤U,and f(x)=0 elsewhere.f(x) = \frac{1}{U - L}, \quad L \le x \le U, \quad \text{and } f(x)=0 \text{ elsewhere.}

Key formulas

QuantityExpression
CDF F(x)F(x)x−LU−L\displaystyle \frac{x - L}{U - L}
Tail probability Fˉ(x)\bar{F}(x)U−xU−L\displaystyle \frac{U - x}{U - L}
Probability in [a,b][a,b]b−aU−L\displaystyle \frac{b - a}{U - L}
Expectation E[X]E[X]L+U2\displaystyle \frac{L + U}{2}
Variance Var(X)\text{Var}(X)(U−L)212\displaystyle \frac{(U - L)^2}{12}

Worked Example: Driving Time

Nilesh Shah’s driving time from Vadodara to Ahmedabad is uniformly distributed between L=100L=100 minutes and U=140U=140 minutes.

  • Expected time: E[X]=100+1402=120E[X] = \frac{100+140}{2} = 120 minutes.
  • P(X≤130)P(X \le 130): F(130)=130−100140−100=3040=0.75F(130) = \frac{130-100}{140-100} = \frac{30}{40} = 0.75 (75% chance of 130 minutes or less).
  • P(X>105)P(X > 105): Fˉ(105)=140−10540=3540=0.875\bar{F}(105) = \frac{140-105}{40} = \frac{35}{40} = 0.875 (87.5% chance of more than 105 minutes).
  • P(X=120)P(X = 120): =0=0 (any single point has zero probability for a continuous variable).

Note: Probabilities for intervals (e.g., P(105<X<130)=F(130)−F(105)=0.75−0.125=0.625P(105 < X < 130) = F(130)-F(105) = 0.75 - 0.125 = 0.625) can be computed similarly.


Key takeaways

  • Continuous random variables have a pdf; probabilities are areas under the curve.
  • The CDF gives P(X≤x)P(X \le x); tail probability is 1−CDF1-\text{CDF}.
  • For continuous variables, P(X=a)=0P(X=a)=0 for any specific aa.
  • Expectation and variance use integrals instead of sums.
  • Uniform distribution: constant pdf over [L,U][L,U]; E[X]=L+U2E[X]=\frac{L+U}{2}, Var(X)=(U−L)212\text{Var}(X)=\frac{(U-L)^2}{12}.
  • In a uniform model, probability is proportional to interval length.

Uniform Distribution Example

The uniform distribution models a random variable that is equally likely to take any value within a given interval [L,U][L, U]. Intuitively: every value in that range has the same chance—there is no “peaked” area.

Worked Example: Battery Life of a Phone

Meena’s phone battery lasts between 8 and 12 hours before needing a charge. Let XX = time (in hours) until recharge. Model: X∼Uniform(L=8,U=12)X \sim \text{Uniform}(L=8, U=12).

PDF (not needed for calculations here):

f(x)=1U−L=14,8≤x≤12f(x) = \frac{1}{U-L} = \frac{1}{4}, \quad 8 \le x \le 12

CDF for any xx in [L,U][L,U]:

F(x)=P(X≤x)=x−LU−LF(x) = P(X \le x) = \frac{x - L}{U - L}

1. Probability that recharge needed within 9 hours (P(X≤9)P(X \le 9))

F(9)=9−812−8=14=0.25F(9) = \frac{9-8}{12-8} = \frac{1}{4} = 0.25

25% chance the phone needs charging in the first 9 hours.

2. Probability that battery lasts at least 11 hours (P(X≥11)P(X \ge 11))

P(X≥11)=1−F(11)=12−1112−8=0.25P(X \ge 11) = 1 - F(11) = \frac{12-11}{12-8} = 0.25

25% chance the battery lasts 11 hours or more.

3. 80th percentile – the value xx such that P(X≤x)=0.8P(X \le x) = 0.8.

x−812−8=0.8⇒x=8+(0.8)(4)=11.2 hours\frac{x - 8}{12 - 8} = 0.8 \quad \Rightarrow \quad x = 8 + (0.8)(4) = 11.2 \text{ hours}

There is an 80% chance the phone will need charging within 11.2 hours.


Exam tip: For uniform distributions, percentiles are a straight linear interpolation: x=L+p⋅(U−L)x = L + p \cdot (U-L).

Discrete Uniform – Reminder

The discrete uniform distribution assigns equal probability to a finite set of values (e.g., a fair die: P(X=k)=1/6P(X=k)=1/6 for k=1,…,6k=1,\dots,6). The continuous version is its natural extension to an interval.

Key takeaways (Uniform)

  • X∼Uniform(L,U)X \sim \text{Uniform}(L,U): all values equally likely in [L,U][L,U].
  • CDF: F(x)=x−LU−LF(x) = \frac{x-L}{U-L} for L≤x≤UL \le x \le U.
  • Percentile pp: x=L+p⋅(U−L)x = L + p \cdot (U-L).
  • Used when only the minimum and maximum are known and nothing else.

Exponential Distribution

The exponential distribution models the time between random events that occur at a constant average rate λ\lambda (events per unit time). Intuitively: if events happen “on average every 1/λ time units”, the waiting time until the next event follows an exponential distribution.

Formal Definition

Let XX = time between events. X≥0X \ge 0.

Probability density function (PDF):

f(x)=λe−λx,x≥0f(x) = \lambda e^{-\lambda x}, \quad x \ge 0

Cumulative distribution function (CDF):

F(x)=P(X≤x)=1−e−λxF(x) = P(X \le x) = 1 - e^{-\lambda x}

Tail probability (survival function):

Fˉ(x)=P(X>x)=e−λx\bar{F}(x) = P(X > x) = e^{-\lambda x}

Probability between aa and bb (0≤a<b0 \le a < b):

P(a<X≤b)=Fˉ(a)−Fˉ(b)=e−λa−e−λbP(a < X \le b) = \bar{F}(a) - \bar{F}(b) = e^{-\lambda a} - e^{-\lambda b}

Mean and Variance

PropertyFormulaNote
Mean E[X]E[X]1λ\frac{1}{\lambda}Average waiting time
Variance Var(X)\text{Var}(X)1λ2\frac{1}{\lambda^2}
Standard deviation1λ\frac{1}{\lambda}Equal to the mean – a unique property

The mean equals the standard deviation for the exponential distribution. This is a quick check: if the sample mean and sample SD are very different, the data may not be exponential.

Key Properties

  • Skewed right – the PDF peaks at x=0x=0 and decays; most probability lies near zero.
  • Memoryless property – P(X>s+t∣X>s)=P(X>t)P(X > s+t \mid X > s) = P(X > t). The future waiting time does not depend on how long you have already waited.
  • Connection to Poisson: If events occur at rate λ\lambda per unit time (Poisson process), the waiting time between consecutive events is exponential with mean 1/λ1/\lambda, and the number of events in a fixed interval is Poisson with mean λt\lambda t.

Applications

  • Time to complete a manual task (assembly, phone call)
  • Lifetime of electronic components (bulbs, circuits)
  • Time until equipment requires repair
  • Waiting time in a queue (call center, customer service)

Worked Example: Call Center Waiting Time

Megha’s waiting time (in minutes) for a service desk agent is exponential with mean 2 minutes. So λ=1/2=0.5\lambda = 1/2 = 0.5 per minute.

1. Probability she waits less than 1 minute (P(X<1)P(X < 1))

F(1)=1−e−0.5×1=1−e−0.5≈0.393F(1) = 1 - e^{-0.5 \times 1} = 1 - e^{-0.5} \approx 0.393

2. Probability she waits more than 3 minutes (P(X>3)P(X > 3))

Fˉ(3)=e−0.5×3=e−1.5≈0.223\bar{F}(3) = e^{-0.5 \times 3} = e^{-1.5} \approx 0.223

3. Probability she waits between 1 and 3 minutes (P(1<X≤3)P(1 < X \le 3))

e−0.5×1−e−0.5×3=e−0.5−e−1.5≈0.6065−0.2231=0.383e^{-0.5 \times 1} - e^{-0.5 \times 3} = e^{-0.5} - e^{-1.5} \approx 0.6065 - 0.2231 = 0.383

4. 90th percentile – find xx such that F(x)=0.9F(x) = 0.9:

1−e−0.5x=0.9⇒e−0.5x=0.1⇒−0.5x=ln⁡(0.1)⇒x=−2ln⁡(0.1)≈4.605 minutes1 - e^{-0.5 x} = 0.9 \quad \Rightarrow \quad e^{-0.5 x} = 0.1 \quad \Rightarrow \quad -0.5 x = \ln(0.1) \quad \Rightarrow \quad x = -2 \ln(0.1) \approx 4.605 \text{ minutes}

90% of calls are answered within about 4.6 minutes.

Computation in Excel

Use the function =EXPON.DIST(x, λ, cumulative):

  • cumulative = FALSE returns the PDF f(x)f(x).
  • cumulative = TRUE returns the CDF F(x)F(x).

For inverse calculations (given probability, find xx), Excel provides =EXPON.INV(probability, λ) or use the formula x=−ln⁡(1−p)λx = -\frac{\ln(1-p)}{\lambda} for the pp-th percentile.

Key takeaways (Exponential)

  • Models waiting time between random events at constant rate λ.
  • PDF: f(x)=λe−λxf(x) = \lambda e^{-\lambda x}, CDF: F(x)=1−e−λxF(x) = 1 - e^{-\lambda x}.
  • Mean = Standard deviation = 1/λ1/\lambda.
  • Memoryless: P(X>s+t∣X>s)=P(X>t)P(X > s+t \mid X > s) = P(X > t).
  • Connected to Poisson: number of events per interval ~ Poisson(λt), waiting time ~ Exponential(λ).
  • Use EXPON.DIST in Excel for probabilities; EXPON.INV for percentiles.

Memoryless Property of the Exponential Distribution

The memoryless property is a defining characteristic of the exponential distribution. Intuitively: if you are waiting for an event that follows an exponential distribution, the probability of waiting an additional yy units does not depend on how long you have already waited — the process "forgets" the elapsed time.

Formal definition

Let X∼Exponential(λ)X \sim \text{Exponential}(\lambda), with mean 1/λ1/\lambda. For any t0>0t_0 > 0 and y>0y > 0,

P(X>t0+y∣X>t0)=P(X>y).P(X > t_0 + y \mid X > t_0) = P(X > y).

Equivalently, the remaining time Y=X−t0Y = X - t_0 conditional on X>t0X > t_0 has the same exponential distribution as XX:

P(Y>y)=e−λy.P(Y > y) = e^{-\lambda y}.

Derivation

The tail probability of an exponential is P(X>x)=e−λxP(X > x) = e^{-\lambda x}.

P(Y>y)=P(X>t0+y∣X>t0)=P(X>t0+y)P(X>t0)=e−λ(t0+y)e−λt0=e−λy.\begin{aligned} P(Y > y) &= P(X > t_0 + y \mid X > t_0) \\ &= \frac{P(X > t_0 + y)}{P(X > t_0)} \\ &= \frac{e^{-\lambda (t_0 + y)}}{e^{-\lambda t_0}} \\ &= e^{-\lambda y}. \end{aligned}

Thus YY has the same distribution as XX.

Exam tip: The memoryless property is unique to the exponential distribution (and its discrete analogue, the geometric). It is the reason exponential waiting times are used in queueing models: the future waiting time is independent of the past.

Worked example — Shanti's bus wait

Shanti waits for a bus. Waiting time XX (minutes) is exponential with mean 2020 minutes.

λ=120=0.05,tails: P(X>x)=e−0.05x.\lambda = \frac{1}{20} = 0.05 ,\quad \text{tails: } P(X > x) = e^{-0.05 x}.
QuestionComputationResult
Probability she waits ≤15\le 15 minP(X≤15)=1−e−0.05⋅15P(X \le 15) = 1 - e^{-0.05 \cdot 15}0.5280.528
Probability she waits >30> 30 minP(X>30)=e−0.05⋅30P(X > 30) = e^{-0.05 \cdot 30}0.2230.223
Probability she waits between 1515 and 3030 minP(15<X≤30)=e−0.05⋅15−e−0.05⋅30P(15 < X \le 30) = e^{-0.05\cdot 15} - e^{-0.05 \cdot 30}0.2490.249
Given she already waited 1010 min, probability she waits another ≥10\ge 10 minP(X>20∣X>10)=e−0.05⋅10P(X > 20 \mid X > 10) = e^{-0.05 \cdot 10}0.6070.607
Given she already waited 2020 min, probability she waits another ≥10\ge 10 minP(X>30∣X>20)=e−0.05⋅10P(X > 30 \mid X > 20) = e^{-0.05 \cdot 10}0.6070.607

The last two results are identical — the waiting time resets regardless of elapsed time, illustrating the memoryless property.


Key takeaways

  • Memoryless property: P(X>t0+y∣X>t0)=P(X>y)P(X > t_0 + y \mid X > t_0) = P(X > y).
  • Equivalent to the remaining time having the same exponential distribution.
  • Only the exponential (continuous) and geometric (discrete) have this property.
  • Conditional waiting times are computed using e−λye^{-\lambda y} — no dependence on t0t_0.
  • Always check that the random variable is exponential before applying memorylessness.

Relation Between Exponential and Poisson Distributions

The exponential distribution is the waiting time distribution that underlies a Poisson process. If events occur according to a Poisson process with rate λ\lambda (average number of events per unit time or space), then:

  • The number of events in a fixed interval of length tt is Poisson(λt)\text{Poisson}(\lambda t).
  • The time between consecutive events (interarrival times) is Exponential(λ)\text{Exponential}(\lambda).

This duality allows us to switch between counting problems and waiting‑time problems.

ProcessInterarrival timesCounting (number of events in tt)
DeterministicFixed, constantFixed number λt\lambda t
PoissonExponential(λ\lambda)Poisson(λt\lambda t)

Example — highway defects

Recall the Solapur highway example: defects occur as a Poisson process with average 33 defects per kilometer.

λ=3 defects/km.\lambda = 3 \text{ defects/km}.
  • Distance between defects XX is Exponential(λ=3)\text{Exponential}(\lambda = 3) with mean 1/3≈0.3331/3 \approx 0.333 km.
  • Number of defects in LL km is Poisson(3L)\text{Poisson}(3L).
QuestionComputationResult
Probability of at least 0.50.5 km without a defectP(X>0.5)=e−3⋅0.5=e−1.5P(X > 0.5) = e^{-3 \cdot 0.5} = e^{-1.5}0.2230.223
Probability of at least 22 km without a defectP(X>2)=e−3⋅2=e−6P(X > 2) = e^{-3 \cdot 2} = e^{-6}0.0020.002

The same result can be obtained from the Poisson: P(zero defects in L km)=e−λLP(\text{zero defects in }L\text{ km}) = e^{-\lambda L}.

Exam tip: Whenever you see a problem about “time between events” or “distance between occurrences”, check whether the process is Poisson. If yes, the interarrival time is exponential — use the exponential tail formula e−λxe^{-\lambda x} directly.

Key takeaways

  • In a Poisson process, interarrival times are i.i.d. exponential with the same rate λ\lambda.
  • Average interarrival time = 1/λ1/\lambda; average number of events per unit time = λ\lambda.
  • Switching between Poisson and exponential: P(zero events in t)=e−λt=P(interarrival time>t)P(\text{zero events in }t) = e^{-\lambda t} = P(\text{interarrival time} > t).
  • Memoryless property of the exponential explains the “lack of clustering” in a Poisson process — the hazard rate is constant.

Worked Example: Oil Change Service Time — Exponential & Poisson

Scenario: Rohit’s garage offers a 50% discount if an oil change takes more than 20 minutes. The average oil change time is 20 minutes. Assuming the time XX (in minutes) for one oil change follows an exponential distribution, we evaluate the risk of this promise. Later we shift to counting the number of oil changes YY in a fixed time window, which follows a Poisson distribution.

1. Time for a single oil change: Exponential model

If the mean time is μ=20\mu = 20 minutes, the rate parameter is λ=1μ=0.05\lambda = \frac{1}{\mu} = 0.05 per minute.
The exponential CDF: F(x)=P(X≤x)=1−e−λxF(x) = P(X \le x) = 1 - e^{-\lambda x}.

Key relationship: Exponential distribution models the time between events (here, the time to complete one oil change).

Probability that time < 20 minutes (promise kept)

P(X<20)=F(20)=1−e−0.05×20=1−e−1≈0.632P(X < 20) = F(20) = 1 - e^{-0.05 \times 20} = 1 - e^{-1} \approx 0.632

There is a 63.2% chance the oil change finishes within 20 minutes. Rohit will likely give many discounts.

Probability that time is between 15 and 30 minutes

P(15<X<30)=F(30)−F(15)=(1−e−0.05×30)−(1−e−0.05×15)=e−0.75−e−1.5≈0.472−0.223=0.249P(15 < X < 30) = F(30) - F(15) = (1 - e^{-0.05 \times 30}) - (1 - e^{-0.05 \times 15}) = e^{-0.75} - e^{-1.5} \approx 0.472 - 0.223 = 0.249

There is a 24.9% chance the oil change takes between 15 and 30 minutes.

EventProbability
X<20X < 200.632
15<X<3015 < X < 300.249

2. Number of oil changes in a time window: Poisson model

Number of oil changes YY in a fixed time period tt follows a Poisson distribution with mean Λ=λt\Lambda = \lambda t, where λ\lambda is the exponential rate (0.05 per minute). This connection holds because the exponential distribution models inter-arrival times of a Poisson process.

Four‑hour morning shift (240 minutes)

Λ=0.05×240=12⇒Y∼Poisson(12)\Lambda = 0.05 \times 240 = 12 \quad \Rightarrow \quad Y \sim \text{Poisson}(12)

Probability of more than 15 oil changes in 4 hours:
P(Y≥15)=1−P(Y≤14)P(Y \ge 15) = 1 - P(Y \le 14). Using Poisson tables or software: P(Y≥15)≈0.228P(Y \ge 15) \approx 0.228

Eight‑hour day (480 minutes)

Λ=0.05×480=24⇒Y∼Poisson(24)\Lambda = 0.05 \times 480 = 24 \quad \Rightarrow \quad Y \sim \text{Poisson}(24)

Probability of more than 30 oil changes in 8 hours:
P(Y≥30)≈0.132P(Y \ge 30) \approx 0.132

Time windowAverage oil changes (Λ\Lambda)P(Y≥threshold)P(Y \ge \text{threshold})Threshold
4 hours (morning)120.22815
8 hours (full day)240.13230

Relationship between Exponential and Poisson

The Poisson distribution counts events when the inter‑event times are i.i.d. exponential with the same rate λ\lambda.

Key Takeaways

  • Exponential models the time until one event (oil change). Use λ=1/mean\lambda = 1/\text{mean}.
  • Poisson models the number of events in a fixed interval when events occur at a constant average rate.
  • The two are linked: Λ=λ×t\Lambda = \lambda \times t.
  • Rohit’s offer is risky: 63% chance of finishing on time (but 37% chance of discount). The probabilities of completing many oil changes per shift are modest (22.8% for >15 in 4 hours, 13.2% for >30 in 8 hours).
  • Exam tip: When a problem gives an average time and asks for counts over a time period, switch from exponential to Poisson using Λ=rate×time\Lambda = \text{rate} \times \text{time}. Always check if the exponential inter‑event time assumption is justified (often stated as “random arrivals” or “memoryless” process).

The Normal Random Variable and its Distribution

The normal distribution is the most widely used continuous probability distribution. It models many naturally occurring phenomena—heights, weights, test scores, rainfall—and is central to statistical inference. Its probability density function (PDF) produces the familiar bell-shaped curve, symmetric about the mean.

f(x)=1σ2π e−(x−μ)22σ2,−∞<x<∞f(x) = \frac{1}{\sigma\sqrt{2\pi}} \, e^{-\frac{(x-\mu)^2}{2\sigma^2}}, \quad -\infty < x < \infty

The entire family of normal distributions is differentiated by two parameters:

  • Mean μ\mu – determines the centre; can be any real number.
  • Standard deviation σ\sigma – determines the spread; σ>0\sigma > 0.

Properties of the Normal Curve

  • Highest point at x=μx = \mu, which is also the median and mode.
  • Symmetric about μ\mu: left half is a mirror image of the right half. Skewness = 0.
  • Tails extend to infinity in both directions, never touching the horizontal axis.
  • Total area under curve = 1 (like all PDFs). Because of symmetry, area to the left of μ\mu = 0.5; area to the right = 0.5.
  • Spread determined by σ\sigma: larger σ\sigma → wider, flatter curve; smaller σ\sigma → taller, narrower curve.

The Empirical Rule (68–95–99.7 Rule)

For any normal random variable XX, the percentage of values within kk standard deviations of the mean is fixed:

IntervalPercentage of valuesProbability
μ±1σ\mu \pm 1\sigma68.3%P(μ−σ≤X≤μ+σ)=0.683P(\mu - \sigma \leq X \leq \mu + \sigma) = 0.683
μ±2σ\mu \pm 2\sigma95.4%P(μ−2σ≤X≤μ+2σ)=0.954P(\mu - 2\sigma \leq X \leq \mu + 2\sigma) = 0.954
μ±3σ\mu \pm 3\sigma99.7%P(μ−3σ≤X≤μ+3σ)=0.997P(\mu - 3\sigma \leq X \leq \mu + 3\sigma) = 0.997

Exam tip: The empirical rule provides quick probability approximations and is often tested directly.

The Standard Normal Distribution

A normal distribution with μ=0\mu = 0 and σ=1\sigma = 1 is called the standard normal distribution. Its random variable is denoted by ZZ instead of XX.

f(z)=12π e−z2/2,−∞<z<∞f(z) = \frac{1}{\sqrt{2\pi}} \, e^{-z^2/2}, \quad -\infty < z < \infty

The cumulative distribution function (CDF) of ZZ is often written as Φ(z)\Phi(z) (capital phi), and the PDF as ϕ(z)\phi(z). Probability tables for Φ(z)\Phi(z) eliminate the need for integration.

Computing Normal Probabilities with Excel

Two built-in functions handle any normal distribution directly (no need to standardise):

  • NORM.DIST(x,mean,standard_dev,cumulative)(x, \text{mean}, \text{standard\_dev}, \text{cumulative})

    • cumulative = TRUE → returns Φ(x)=P(X≤x)\Phi(x) = P(X \leq x) (CDF).
    • cumulative = FALSE → returns the PDF value at xx.
  • NORM.INV(probability,mean,standard_dev)(\text{probability}, \text{mean}, \text{standard\_dev})

    • Returns xx such that P(X≤x)=probabilityP(X \leq x) = \text{probability} (inverse CDF).

Template example (general normal: μ=10\mu = 10, σ=3\sigma = 3)

InputComputed ProbabilityResult
x=13x = 13P(X≤13)P(X \leq 13)0.841
x=13x = 13P(X≥13)P(X \geq 13)0.159
a=7,b=13a = 7, b = 13P(7≤X≤13)P(7 \leq X \leq 13)0.683

For inverse:

  • Enter cumulative probability 0.841 → get x=13x = 13.
  • Enter interval probability 0.683 → get a=7a = 7, b=13b = 13.

Worked Examples: Standard Normal Probabilities

Set μ=0\mu = 0, σ=1\sigma = 1 in the template.

Example 1

ProbabilityResult
P(Z≤−1)P(Z \leq -1)0.159
P(Z≥−1)P(Z \geq -1)0.841
P(Z≥−1.5)P(Z \geq -1.5)0.933
P(Z≥−2.5)P(Z \geq -2.5)0.994
P(−3<Z≤0)P(-3 < Z \leq 0)0.499

Note: P(Z=a)=0P(Z = a) = 0 for any continuous random variable, so P(Z≤a)=P(Z<a)P(Z \leq a) = P(Z < a) and P(Z≥a)=P(Z>a)P(Z \geq a) = P(Z > a).

Example 2

ProbabilityResult
P(0≤Z≤0.8)P(0 \leq Z \leq 0.8)0.288
P(−1.5≤Z≤0)P(-1.5 \leq Z \leq 0)0.433
P(Z≥0.4)P(Z \geq 0.4)0.345
P(Z≥−0.2)P(Z \geq -0.2)0.579
P(Z≤1.2)P(Z \leq 1.2)0.885
P(Z≤−0.7)P(Z \leq -0.7)0.242

Example 3: Inverse Calculations

ConditionRange of zz
Area to the left of zz is 0.21z≤−0.806z \leq -0.806 (i.e., z=−0.806z = -0.806)
Area between −z-z and +z+z is 0.9−1.64≤z≤+1.64-1.64 \leq z \leq +1.64
Area between −z-z and +z+z is 0.2−0.25≤z≤+0.25-0.25 \leq z \leq +0.25
Area to the left of zz is 0.955z≤2.58z \leq 2.58 (i.e., z=2.58z = 2.58)
Area to the right of zz is 0.692z≥−0.502z \geq -0.502 (i.e., z=−0.502z = -0.502)

Key Takeaways

  • The normal distribution is bell-shaped, symmetric about μ\mu, with μ=\mu = median == mode.
  • It is parameterised by μ\mu (centre) and σ\sigma (spread).
  • The empirical rule: ≈68% within 1σ, ≈95% within 2σ, ≈99.7% within 3σ.
  • The standard normal ZZ has μ=0\mu = 0, σ=1\sigma = 1; its CDF is Φ(z)\Phi(z).
  • Excel: NORM.DIST for cumulative/PDF, NORM.INV for inverse; both work for any normal.
  • For continuous variables, point probabilities are zero; probabilities always refer to intervals.

Standard Normal Distribution

The standard normal distribution is the symmetric, bell-shaped distribution centered at 0 with a standard deviation of 1. Its random variable is denoted ZZ. This distribution serves as the universal reference for all normal distributions — any normal variable can be transformed into a standard normal.

An important concept for statistical inference is the critical value. For a standard normal, a critical value zα/2z_{\alpha/2} is defined such that the tail probability to the right of it equals α/2\alpha/2:

P(Z>zα/2)=α/2P(Z > z_{\alpha/2}) = \alpha/2

By symmetry, −zα/2-z_{\alpha/2} has a left‑tail probability of α/2\alpha/2. Consequently, the central probability between −zα/2-z_{\alpha/2} and +zα/2+z_{\alpha/2} is 1−α1-\alpha:

P(−zα/2≤Z≤zα/2)=1−αP(-z_{\alpha/2} \leq Z \leq z_{\alpha/2}) = 1-\alpha

The familiar 68‑95‑99.7 rule is a special case: for 1−α=0.681-\alpha = 0.68, α/2=0.16\alpha/2 = 0.16, and z0.16=1z_{0.16} = 1.

Common critical values

1−α1-\alpha (confidence level)α/2\alpha/2zα/2z_{\alpha/2}
0.800.101.28
0.900.051.64
0.950.0251.96
0.990.0052.58

Exam tip: Memorise the three critical values 1.64, 1.96, and 2.58 — they appear repeatedly in confidence intervals and hypothesis tests.

Two standard types of questions

  1. Given ZZ, find probability — use the standard normal table or template.
  2. Given probability, find ZZ — inverse lookup. Sketching the bell curve and shading the relevant area always helps avoid sign errors.

Converting any normal to the standard normal

If X∼N(μ,σ)X \sim N(\mu, \sigma), then the transformed variable

Z=X−μσZ = \frac{X - \mu}{\sigma}

follows a standard normal distribution. ZZ measures how many standard deviations XX is away from the mean — positive above, negative below.

Worked example
Let X∼N(μ=10,  σ=2)X \sim N(\mu=10,\;\sigma=2). Find P(10<X<14)P(10 < X < 14).

  1. Convert boundaries:
    Z1=10−102=0,Z2=14−102=2Z_1 = \frac{10-10}{2} = 0, \quad Z_2 = \frac{14-10}{2} = 2
  2. Then P(10<X<14)=P(0<Z<2)P(10 < X < 14) = P(0 < Z < 2).
  3. From the standard normal table, P(0<Z<2)=0.4772P(0 < Z < 2) = 0.4772.

The same result could be obtained by directly using a normal template with μ=10\mu=10, σ=2\sigma=2; the conversion simply standardises the problem.

Key takeaways

  • Critical value zα/2z_{\alpha/2}: right‑tail area = α/2\alpha/2; symmetric about 0.
  • Central probability 1−α1-\alpha lies between −zα/2-z_{\alpha/2} and +zα/2+z_{\alpha/2}.
  • Common zα/2z_{\alpha/2} values: 1.28, 1.64, 1.96, 2.58 for confidence levels 80%, 90%, 95%, 99%.
  • Any normal XX can be transformed via Z=(X−μ)/σZ = (X-\mu)/\sigma to a standard normal.
  • ZZ tells the distance from the mean in standard deviation units.
  • Always sketch the curve and shade the area when solving probability problems.

Examples of Normal Distribution Applications

The normal distribution is a versatile model for many real-world variables. These worked examples illustrate how to compute probabilities for intervals, tails, and cumulative regions, and how to perform inverse calculations (percentiles) using given mean μ\mu and standard deviation σ\sigma. The key tool is the z‑score: z=(x−μ)/σz = (x - \mu)/\sigma, which maps any xx to a standard normal N(0,1)N(0,1).

Exam tip: Always check units (e.g., minutes vs. hours). A small mistake in input changes the entire answer.

1. Stock Returns (Munaf)

A stock fund has a mean annual return of μ=12%\mu = 12\% and a standard deviation σ=3%\sigma = 3\%. Returns are normally distributed.

  • Unsatisfactory (return <5%< 5\%):
    z=5−123=−2.33z = \frac{5-12}{3} = -2.33 → P(X<5%)=0.01P(X < 5\%) = 0.01 (1% chance).
  • Excellent (return >10%> 10\%):
    z=10−123=−0.667z = \frac{10-12}{3} = -0.667 → P(X>10%)=0.748P(X > 10\%) = 0.748 (74.8% chance).
  • Moderate (return between 5% and 10%):
    P(5%<X<10%)=P(X<10%)−P(X<5%)=0.748−0.01=0.243P(5\% < X < 10\%) = P(X < 10\%) - P(X < 5\%) = 0.748 - 0.01 = 0.243 (24.3% chance).
  • 90th percentile (inverse):
    • Cumulative: 90% of returns are less than 15.8%.
    • Tail: 90% of returns are greater than 8.1%.
    • Interval: 90% of returns lie between 7% and 17%.

Key mechanism: The normal distribution template directly uses μ\mu and σ\sigma to output these probabilities without manual z‑score calculation.

2. Screen Time (Bhavani)

Daily screen time of school kids: μ=8.4\mu = 8.4 hours, σ=2.5\sigma = 2.5 hours.

  • More than 4 hours:
    z=4−8.42.5=−1.76z = \frac{4-8.4}{2.5} = -1.76 → P(X>4)=0.96P(X > 4) = 0.96 (96% of kids exceed 4 hours).
  • Between 6 and 12 hours:
    z6=6−8.42.5=−0.96z_6 = \frac{6-8.4}{2.5} = -0.96, z12=12−8.42.5=1.44z_{12} = \frac{12-8.4}{2.5} = 1.44 → P(6<X<12)=0.757P(6 < X < 12) = 0.757 (75.7% of kids).
  • Top 20% (inverse):
    Need P(X<x)=0.8P(X < x) = 0.8 → x=10.5x = 10.5 hours. A kid with >10.5 hours is in the top 20%.
  • Bottom 20% (inverse):
    Need P(X<x)=0.2P(X < x) = 0.2 → x=6.3x = 6.3 hours. A kid with <6.3 hours is in the bottom 20%.
  • (Bonus) 80% interval: 80% of kids have screen time between 5.2 and 11.6 hours.

3. Exam Duration (Hirav)

Time to complete an exam: μ=150\mu = 150 minutes, σ=20\sigma = 20 minutes. 60 students.

  • 2 hours or less (≤120 min):
    z=120−15020=−1.5z = \frac{120-150}{20} = -1.5 → P(X≤120)=0.067P(X \leq 120) = 0.067 → 0.067×60≈40.067 \times 60 \approx 4 students.
  • Between 2 and 3 hours (120–180 min):
    z120=−1.5z_{120} = -1.5, z180=+1.5z_{180} = +1.5 → P(120<X<180)=0.866P(120 < X < 180) = 0.866 → 0.866×60≈520.866 \times 60 \approx 52 students.
  • 3 hours or more (≥180 min):
    P(X≥180)=0.067P(X \geq 180) = 0.067 → 0.067×60≈40.067 \times 60 \approx 4 students.

Exam tip: The symmetry of the normal distribution means probabilities at equal distances above and below the mean are identical. Here 120 and 180 min are both 1.5σ from the mean, giving the same tail probability (0.067).

4. Sugar Packet Weight (Naman)

Two filling machines; weights are normally distributed. Naman evaluates which machine is better.

Machineμ\mu (kg)σ\sigma (kg)P(underweight<1P(\text{underweight} < 1 kg)P(1≤X≤1.05)P(1 \leq X \leq 1.05)90% interval
11.010.020.309 (30.9%)0.669 (66.9%)[0.97, 1.04]
21.030.040.227 (22.7%)0.465 (46.5%)[0.96, 1.09]
  • Machine 1 has fewer underweight? No, 30.9% > 22.7%, so Machine 2 produces fewer underweight packets.
  • Machine 2 has a narrower range for 90% of packets? No, its interval is wider (0.96–1.09 vs. 0.97–1.04).
    Naman’s claim that 90% of packets lie between 0.99 and 1.01 is false for both machines.

Trade-off: Machine 2 reduces underweight risk but increases variability, making the middle‑90% range much broader.

5. Glass Thickness (Vibhor)

Glass sheet thickness: μ=3\mu = 3 mm, σ=0.12\sigma = 0.12 mm.

  • Thickness < 2.9 mm (cumulative):
    z=2.9−30.12≈−0.833z = \frac{2.9-3}{0.12} \approx -0.833 → P(X<2.9)=0.202P(X < 2.9) = 0.202 (20.2% of sheets too thin).
  • Thickness > 3.2 mm (tail):
    z=3.2−30.12≈1.667z = \frac{3.2-3}{0.12} \approx 1.667 → P(X>3.2)=0.048P(X > 3.2) = 0.048 (4.8% too thick).
  • 95% thickness range (inverse):
    For a central probability of 0.95, the endpoints are 3±1.96×0.123 \pm 1.96 \times 0.12 → [2.76 mm, 3.24 mm].

Key takeaways

  • All calculations rely on the z‑score transformation z=(x−μ)/σz = (x - \mu)/\sigma, linking any normal variable to the standard normal.
  • Three standard probability queries: cumulative (less than), tail (greater than), and interval (between). Use the CDF differences.
  • Inverse problems (percentiles) find the value xx for a given cumulative probability (e.g., top 20% → p=0.8p=0.8).
  • Always check units (hours vs. minutes, percentage vs. decimal) before plugging numbers into a template or formula.
  • Symmetry simplifies: P(X<μ−kσ)=P(X>μ+kσ)P(X < \mu - k\sigma) = P(X > \mu + k\sigma).

Normal Approximation to the Binomial Distribution

When n is large, computing binomial probabilities directly (via factorials in (nx)pxqn−x\binom{n}{x} p^x q^{n-x}) becomes unwieldy. The normal distribution provides a remarkably accurate approximation because the binomial’s probability mass function is approximately bell‑shaped — especially when p is not too close to 0 or 1.

Conditions for a Good Approximation

The approximation works well when both:

np≥5andnq≥5(where q=1−p)np \ge 5 \quad \text{and} \quad nq \ge 5 \quad (\text{where } q = 1-p)

If these hold, match the mean and variance of the normal to those of the binomial:

μ=npσ2=npq\mu = np \qquad \sigma^2 = npq

So we use Y∼N(μ=np,  σ2=npq)Y \sim N(\mu = np,\; \sigma^2 = npq) to approximate X∼Bin(n,p)X \sim \text{Bin}(n,p).

The Continuity Correction

Because the binomial is discrete and the normal is continuous, a continuity correction of ±0.5 is applied to improve accuracy.

Desired binomial probabilityContinuity‑corrected normal probability
P(X≤a)P(X \le a)P(Y≤a+0.5)P(Y \le a + 0.5)
P(X≥b)P(X \ge b)P(Y≥b−0.5)P(Y \ge b - 0.5)
P(a≤X≤b)P(a \le X \le b)P(a−0.5≤Y≤b+0.5)P(a - 0.5 \le Y \le b + 0.5)
P(X=k)P(X = k)P(k−0.5≤Y≤k+0.5)P(k - 0.5 \le Y \le k + 0.5)

Exam tip: The continuity correction is the most common mistake. Always add/subtract 0.5; never approximate a discrete probability using a single point from a continuous distribution (that probability would be 0).

Worked Examples

Example 1: Unbiased coin, 16 tosses
n=16,  p=0.5n=16,\; p=0.5 → μ=8,  σ2=4  (σ=2)\mu=8,\; \sigma^2=4\;( \sigma=2).

  • Exact binomial: P(X≤5)=0.105P(X \le 5) = 0.105
    Normal approx: P(Y≤5.5)=0.1056(≈0.105)P(Y \le 5.5) = 0.1056 \quad (\approx 0.105)

  • Exact binomial: P(8≤X≤11)=0.5598P(8 \le X \le 11) = 0.5598
    Normal approx: P(7.5≤Y≤11.5)=0.5586(≈0.5598)P(7.5 \le Y \le 11.5) = 0.5586 \quad (\approx 0.5598)

Example 2: Invoices with errors (Kunal Poonawala, Junagadh)
n=100,  p=0.1n=100,\; p=0.1 → μ=10,  σ2=9  (σ=3)\mu=10,\; \sigma^2=9\;(\sigma=3).

  • Exact binomial: P(X=12)=0.09P(X = 12) = 0.09
    Normal approx: P(11.5≤Y≤12.5)=0.1052P(11.5 \le Y \le 12.5) = 0.1052 (close, though slightly off; approximation improves with larger npnp).

  • Exact binomial: P(X≤13)=0.876P(X \le 13) = 0.876
    Normal approx: P(Y≤13.5)=0.879(≈0.876)P(Y \le 13.5) = 0.879 \quad (\approx 0.876)

The approximation is reliable, especially when npnp and nqnq are well above 5.

Connection to the Poisson Distribution

The Poisson distribution can also be approximated by the normal when its mean λ\lambda is large. Moreover, for a binomial with small qq (failure probability), the variance npq≈npnpq \approx np, making the binomial similar to a Poisson. Hence, in certain cases the Poisson can be approximated via the binomial → normal chain.


Key Takeaways

  • Normal approximation to the binomial is valid when np≥5np \ge 5 and nq≥5nq \ge 5.
  • Match mean μ=np\mu = np and variance σ2=npq\sigma^2 = npq.
  • Always apply a continuity correction (±0.5) because the binomial is discrete.
  • The approximation is remarkably accurate and is often used by software internally.

Linear Combinations of Random Variables (Normal Case)

Any linear combination of normally distributed random variables is itself normally distributed – not just its mean and variance, but the entire distribution.

General Results (any distribution)

Given random variables X1,X2,…,XnX_1, X_2, \dots, X_n and constants a1,a2,…,ana_1, a_2, \dots, a_n, define

Y=a1X1+a2X2+⋯+anXnY = a_1 X_1 + a_2 X_2 + \cdots + a_n X_n

Expectation (always linear):

E[Y]=a1E[X1]+a2E[X2]+⋯+anE[Xn]E[Y] = a_1 E[X_1] + a_2 E[X_2] + \cdots + a_n E[X_n]

Variance:

  • If the XiX_i are independent:

    Var(Y)=a12Var(X1)+a22Var(X2)+⋯+an2Var(Xn)\text{Var}(Y) = a_1^2 \text{Var}(X_1) + a_2^2 \text{Var}(X_2) + \cdots + a_n^2 \text{Var}(X_n)

  • If the XiX_i are correlated (not independent), covariance terms appear:

    Var(Y)=∑i=1nai2Var(Xi)+2∑i<jaiajCov(Xi,Xj)\text{Var}(Y) = \sum_{i=1}^n a_i^2 \text{Var}(X_i) + 2\sum_{i<j} a_i a_j \text{Cov}(X_i, X_j)

These formulas hold for any distributions of the XiX_i.

Special Case: Normally Distributed XiX_i

If each Xi∼N(μi,σi2)X_i \sim N(\mu_i, \sigma_i^2) and they are jointly normal (or independent normal), then YY is also normally distributed:

Y∼N ⁣(∑aiμi,  ∑ai2σi2+2∑i<jaiajCov(Xi,Xj))Y \sim N\!\left( \sum a_i \mu_i,\; \sum a_i^2 \sigma_i^2 + 2\sum_{i<j} a_i a_j \text{Cov}(X_i, X_j) \right)

This property is unique to the normal family and is crucial for statistical inference (e.g., sums of normal data remain normal).

Exam tip: When a problem states “X1,X2,…X_1, X_2, \dots are independent normal random variables,” any linear combination (like Xˉ\bar{X} or X1−X2X_1 - X_2) is also normal. You only need to compute its mean and variance.


Key Takeaways

  • For any random variables, E[a1X1+… ]E[a_1 X_1 + \dots] is linear; variance adds with squares plus covariance terms if not independent.
  • For normal random variables, the linear combination is also normal — a key result for later use (e.g., Central Limit Theorem).
  • This property does not hold for most other distributions; it is special to the normal.

Linear Combinations of Normal Random Variables

A key property: any linear combination of independent normal random variables is itself normally distributed. This makes it possible to model sums, differences, and averages of normal data with a single normal distribution.

The Basic Result: (Y = aX + b)

If (X \sim N(\mu,\ \sigma^2)) and (Y = aX + b) (with constants (a) and (b)), then
[ Y \sim N\big(a\mu + b,\ a^2\sigma^2\big). ]

Intuition – Shifting and scaling a normal distribution preserves its bell shape; only the mean and variance change linearly.

Standardization as a Special Case

Set (a = \frac{1}{\sigma}) and (b = -\frac{\mu}{\sigma}):

[ Y = \frac{X - \mu}{\sigma} \sim N(0,1) ]

This is the standard normal random variable (Z).

Sum of Two Independent Normals

Let (X_1 \sim N(\mu_1, \sigma_1^2)) and (X_2 \sim N(\mu_2, \sigma_2^2)) be independent. Then
[ Y = X_1 + X_2 \sim N(\mu_1 + \mu_2,\ \sigma_1^2 + \sigma_2^2). ]

If the variables are not independent, the sum is still normal, but the variance includes covariance terms.

General Linear Combination

For (n) independent normal variables (X_i \sim N(\mu_i, \sigma_i^2)) and constants (a_1, \dots, a_n, b):

[ Y = a_1 X_1 + a_2 X_2 + \cdots + a_n X_n + b ]

has distribution

[ Y \sim N!\left( \sum_{i=1}^n a_i \mu_i + b,\ \sum_{i=1}^n a_i^2 \sigma_i^2 \right). ]

Important Special Cases: Sum and Sample Mean

Assume (X_1, X_2, \dots, X_n) are iid (independent and identically distributed) as (N(\mu, \sigma^2)).

QuantityDefinitionDistributionMeanVarianceStandard Deviation
Sum(S_n = \sum_{i=1}^n X_i)(N(n\mu,\ n\sigma^2))(n\mu)(n\sigma^2)(\sqrt{n},\sigma)
Sample mean(\bar{X} = \frac{1}{n}\sum_{i=1}^n X_i)(N(\mu,\ \sigma^2/n))(\mu)(\frac{\sigma^2}{n})(\frac{\sigma}{\sqrt{n}})

Exam tip: Standard deviations do not add up directly. For the sum, (\text{SD}(S_n) = \sqrt{n},\sigma); for the average, (\text{SD}(\bar{X}) = \sigma/\sqrt{n}). The variance reduces by a factor of (n) when averaging – this is the foundation of sampling precision.


Worked Examples

1. Piston–Cylinder Gap

Problem – Piston head radius (X_1 \sim N(30\ \text{mm},\ 0.0025\ \text{mm}^2)), cylinder inside radius (X_2 \sim N(30.25\ \text{mm},\ 0.0036\ \text{mm}^2)). The gap is (Y = X_2 - X_1). Find
(a) (P(Y \le 0)) (piston does not fit)
(b) (P(0.1 \le Y \le 0.35)) (optimal performance)

Solution
(Y) is normal because it is a linear combination of independent normals.

[ \begin{aligned} \mu_Y &= 30.25 - 30 = 0.25 \ \sigma_Y^2 &= 0.0025 + 0.0036 = 0.0061 \ \sigma_Y &= \sqrt{0.0061} \approx 0.078 \end{aligned} ]

(a)
[ P(Y \le 0) = \Phi!\left(\frac{0 - 0.25}{0.078}\right) \approx \Phi(-3.205) \approx 0.0007 ]

(b)
[ P(0.1 \le Y \le 0.35) = \Phi!\left(\frac{0.35-0.25}{0.078}\right) - \Phi!\left(\frac{0.1-0.25}{0.078}\right) \approx \Phi(1.282) - \Phi(-1.923) \approx 0.872 ]

Interpretation – Extremely low chance of misfit; ~87 % chance of optimal gap.

2. Average Height of Cotton Plants

Problem – Individual plant height after two weeks: (X_i \sim N(29.4\ \text{cm},\ 4.41\ \text{cm}^2)). For 20 plants, the sample mean (\bar{X}) is normal. Find the 95th percentile of (\bar{X}).

Solution
[ \bar{X} \sim N!\left(29.4,\ \frac{4.41}{20} = 0.2205\right),\quad \sigma_{\bar{X}} = \sqrt{0.2205} \approx 0.469 ]

The 95th percentile is
[ \mu + z_{0.05},\sigma_{\bar{X}} = 29.4 + 1.645 \times 0.469 \approx 30.17\ \text{cm}. ]

Notice that the standard deviation of the average (0.47 cm) is much smaller than that of an individual plant (2.1 cm) – averaging greatly reduces variability.

3. Stock Returns Comparison

Problem – Stock A: (X_A \sim N(8%,\ 2.25%^2)), Stock B: (X_B \sim N(9.5%,\ 4%^2)), independent.
Define (Y = X_B - X_A). Find:
(a) (P(X_A \text{ moderate})) meaning (5 \le X_A \le 10)
(b) (P(X_B \text{ excellent})) meaning (X_B \ge 10)
(c) (P(X_B > X_A)) i.e., (P(Y > 0))
(d) (P(Y \ge 2))

Solution

(a) (X_A \sim N(8, 2.25)).
[ P(5 \le X_A \le 10) = \Phi!\left(\frac{10-8}{1.5}\right) - \Phi!\left(\frac{5-8}{1.5}\right) \approx \Phi(1.333) - \Phi(-2) \approx 0.886 ]

(b) (X_B \sim N(9.5, 4)).
[ P(X_B \ge 10) = 1 - \Phi!\left(\frac{10-9.5}{2}\right) \approx 1 - \Phi(0.25) \approx 0.4013 ]

(c) (Y = X_B - X_A \sim N(1.5,\ 2.25+4=6.25)), so (\sigma_Y = 2.5).
[ P(Y > 0) = 1 - \Phi!\left(\frac{0-1.5}{2.5}\right) = 1 - \Phi(-0.6) \approx 0.726 ]

(d)
[ P(Y \ge 2) = 1 - \Phi!\left(\frac{2-1.5}{2.5}\right) = 1 - \Phi(0.2) \approx 0.421 ]


Key Takeaways

  • Any linear combination of independent normal random variables is normally distributed.
  • For (Y = \sum a_i X_i + b): mean = (\sum a_i \mu_i + b), variance = (\sum a_i^2 \sigma_i^2) (independence assumed).
  • Sum of iid normals: mean (n\mu), variance (n\sigma^2).
  • Sample mean of iid normals: mean (\mu), variance (\sigma^2/n) – precision improves with (n).
  • Difference of two independent normals: variance adds; subtract the means.
  • Use standardization to compute probabilities for any linear combination.

Chi-Square Distribution

The chi-square distribution arises naturally from squared standard normal variables. If Z∼N(0,1)Z \sim N(0,1), then X=Z2X = Z^2 follows a chi-square distribution with 1 degree of freedom. More generally, if Z1,Z2,…,ZnZ_1, Z_2, \dots, Z_n are independent standard normal random variables, then

X=Z12+Z22+⋯+Zn2X = Z_1^2 + Z_2^2 + \cdots + Z_n^2

follows a chi-square distribution with nn degrees of freedom (denoted χn2\chi^2_n).

Intuitively, it measures the sum of squared deviations; because squares are always non‑negative, the distribution is defined only for positive values.

Properties

  • Single parameter: the degrees of freedom (kk).
  • Mean: μ=k\mu = k.
  • Variance: σ2=2k\sigma^2 = 2k.
  • Shape: right‑skewed for small kk; as kk increases, the skewness decreases and the density approaches a normal distribution (by the Central Limit Theorem).
  • Support: x>0x > 0.

Applications

  • Non‑parametric hypothesis tests (e.g., goodness‑of‑fit).
  • Feature selection and classification.
  • Testing independence in contingency tables (logistic regression).

Critical Values

The chi‑square critical value χα,k2\chi^2_{\alpha, k} is the value such that the right‑tail probability equals α\alpha:

P(X≥χα,k2)=αP(X \geq \chi^2_{\alpha, k}) = \alpha

Unlike the standard normal, critical values depend on the degrees of freedom. The table below gives critical values for several α\alpha and kk (computed using Microsoft Excel’s CHISQ.DIST function).

kkα=0.20\alpha = 0.20α=0.10\alpha = 0.10α=0.05\alpha = 0.05α=0.01\alpha = 0.01
57.39.211.115.1
1013.416.018.323.2
1519.322.325.037.6
2025.028.431.437.6

Exam tip: As kk increases, critical values increase because the distribution spreads out (variance =2k=2k). The values in the last column are the largest because the tail probability α\alpha is smallest.

Computing in Excel

Use CHISQ.DIST(x, k, cumulative):

  • x: value of the chi‑square random variable.
  • k: degrees of freedom.
  • cumulative: TRUE returns the cumulative distribution function; FALSE returns the probability density function.

A template can automate mean, variance, probability calculations, and inverse (critical value) look‑up.

Key takeaways

  • Chi‑square is the sum of kk independent squared standard normals.
  • Mean = kk, variance = 2k2k.
  • Always positive, right‑skewed, converges to normal as k→∞k \to \infty.
  • Critical values increase with kk; used in goodness‑of‑fit and independence tests.

Student’s t‑Distribution

The Student’s t‑distribution (or simply t‑distribution) is defined when a standard normal variable is divided by the square root of an independent chi‑square variable scaled by its degrees of freedom:

T=ZX/kT = \frac{Z}{\sqrt{X / k}}

where Z∼N(0,1)Z \sim N(0,1), X∼χk2X \sim \chi^2_k, and ZZ and XX are independent. TT follows a t‑distribution with kk degrees of freedom.

Intuitively, the t‑distribution is used in place of the standard normal when the population variance is unknown and the sample size is small. It has fatter tails (more probability in the extremes) than the normal.

History

Discovered by William Sealy Gosset while working at the Guinness Brewery. He published under the pseudonym “Student” because his employer prohibited employees from publishing research. Hence the name Student’s t‑distribution.

Properties

  • Symmetric and bell‑shaped, centered at 0.
  • Mean: μ=0\mu = 0 (for k>1k > 1).
  • Variance: σ2=kk−2\sigma^2 = \frac{k}{k-2} (for k>2k > 2; undefined for k≤2k \leq 2).
  • Shape: flatter (more spread) than the standard normal; as k→∞k \to \infty, the t‑distribution converges to N(0,1)N(0,1).
  • Support: −∞<t<∞-\infty < t < \infty.

Applications

  • Hypothesis tests about population means (especially small samples).
  • Comparing means of two populations.
  • Diagnostics in linear regression (e.g., t‑tests for coefficients).

Critical Values

Because the t‑distribution is symmetric, critical values are often defined for two‑tailed regions. The value tα/2,kt_{\alpha/2, k} is the positive number such that:

P(T≥tα/2,k)=α/2P(T \geq t_{\alpha/2, k}) = \alpha/2

By symmetry, P(T≤−tα/2,k)=α/2P(T \leq -t_{\alpha/2, k}) = \alpha/2. Therefore P(−tα/2,k≤T≤tα/2,k)=1−αP(-t_{\alpha/2, k} \leq T \leq t_{\alpha/2, k}) = 1 - \alpha.

The table below gives tα/2,kt_{\alpha/2, k} for common α/2\alpha/2 and degrees of freedom (computed using Excel’s T.DIST function).

kkα/2=0.10\alpha/2 = 0.10α/2=0.05\alpha/2 = 0.05α/2=0.025\alpha/2 = 0.025α/2=0.005\alpha/2 = 0.005
51.482.022.574.03
101.371.812.233.17
151.341.752.132.95
201.331.732.092.85

Notice that as kk increases, the critical values decrease (the distribution becomes less flat) and approach the corresponding standard normal critical values (e.g., for α/2=0.025\alpha/2 = 0.025, z0.025≈1.96z_{0.025} \approx 1.96).

Computing in Excel

Use T.DIST(x, k, cumulative):

  • x: value of the t random variable.
  • k: degrees of freedom.
  • cumulative: TRUE for cumulative probability, FALSE for PDF.

For critical values, the inverse function T.INV or a template can be used.

Exam tip: The t‑distribution is flatter than the normal, so its critical values are larger than the corresponding normal critical values for small kk. As kk grows, they converge.

Key takeaways

  • t‑distribution: T=Z/χk2/kT = Z / \sqrt{\chi^2_k / k}.
  • Symmetric, zero‑mean, fatter tails than normal.
  • Variance =k/(k−2)= k/(k-2); converges to standard normal as k→∞k \to \infty.
  • Critical values decrease with increasing kk; used in mean hypothesis tests and regression diagnostics.

The F Distribution

The F distribution (Fisher–Snedecor distribution, after Ronald Fisher) models the ratio of two independent chi-square random variables, each divided by its own degrees of freedom.

F(k1,k2)=χk12/k1χk22/k2F(k_1, k_2) = \frac{\chi^2_{k_1} / k_1}{\chi^2_{k_2} / k_2}

Here k1k_1 = numerator degrees of freedom and k2k_2 = denominator degrees of freedom. The order matters: F(5,10)≠F(10,5)F(5,10) \neq F(10,5).

Properties

  • Defined only for positive values of the random variable.
  • Unimodal and right-skewed (long right tail).
  • Mean: k2k2−2\displaystyle \frac{k_2}{k_2 - 2} for k2≥2k_2 \geq 2 (when k2≤2k_2 \leq 2, the mean is undefined).
  • Standard deviation decreases as k2k_2 increases.
  • As both k1k_1 and k2k_2 become large, the PDF becomes sharply spiked around 11.

Where it is used

  • Analysis of variance (ANOVA)
  • Feature selection in machine learning
  • Testing overall fit of multiple linear regression models

Excel Function: F.DIST

In Excel, the function F.DIST(x, k1, k2, cumulative) computes probabilities.

ParameterMeaning
xValue of the F random variable
k1Numerator degrees of freedom
k2Denominator degrees of freedom
cumulativeTRUE → cumulative probability; FALSE → probability density

Inverse (critical values) can be obtained via F.INV(alpha, k1, k2) for right‑tail probability α\alpha.

Worked Example: Critical Values (Right‑Tail)

For an F‑distribution with given k1,k2k_1, k_2, the critical value Fα,k1,k2F_{\alpha, k_1, k_2} is the value such that the area to its right equals α\alpha.

Using an F‑distribution template (or Excel), with α=0.2,0.1,0.05,0.01\alpha = 0.2, 0.1, 0.05, 0.01:

(k1,k2)(k_1, k_2)F0.20F_{0.20}F0.10F_{0.10}F0.05F_{0.05}F0.01F_{0.01}
(5, 5)2.233.455.0510.97
(5, 10)1.802.523.335.64
(10, 5)2.193.304.7410.05
(20, 20)1.471.792.212.94

Exam tip: Swapping k1k_1 and k2k_2 changes critical values — always double‑check which is numerator and which is denominator. When k2k_2 is large, critical values approach 1 because the distribution becomes concentrated near 1.

Key takeaways

  • F(k1,k2)=(χk12/k1)/(χk22/k2)F(k_1, k_2) = (\chi^2_{k_1}/k_1) / (\chi^2_{k_2}/k_2) — ratio of independent scaled chi‑squares.
  • Positive, unimodal, right‑skewed; mean =k2/(k2−2)= k_2/(k_2-2) for k2≥2k_2 \geq 2.
  • Order of degrees of freedom matters: F(5,10)≠F(10,5)F(5,10) \neq F(10,5).
  • As k2k_2 increases, variance decreases and distribution peaks nearer 1.
  • Widely used in ANOVA, regression, and machine learning feature selection.
  • Excel: F.DIST(x, k1, k2, TRUE/FALSE) for cumulative/density; F.INV for critical values.