Term 1 · Module 2 of 8

Random Variables and Discrete Probability Distributions

Business Statistics for Entrepreneurs

Introduction to Random Variables

A random variable (RV) assigns a numerical value to each outcome of a random experiment. Think of it as a rule that converts the abstract outcome (e.g., "heads") into a number you can work with (e.g., 00).

Discrete vs. Continuous Random Variables

TypeValuesExamplePractical business example
DiscreteCountable (often integers); can be finite or countably infiniteNumber of heads in 10 coin tosses: {0,1,…,10}\{0,1,\dots,10\}Number of orders from 5 sales calls, number of defective radios in a shipment of 50, number of customers at a restaurant in a day
ContinuousAny value in an interval (including fractions)Time between customer arrivals at a bank: (0,∞)(0,\infty)Actual fluid ounces in a “12 oz” soft‑drink can (between 11.5 and 12.5), temperature at which a chemical reaction occurs (between 100°C and 150°C)

Note: In the same setting (e.g., a bank), the number of customers arriving is a discrete RV, while the time between arrivals is a continuous RV.

Probability Mass Function (PMF)

For a discrete RV XX, the probability mass function (PMF), denoted f(x)f(x), gives the probability that XX equals exactly a particular value xx:

f(x)=P(X=x)f(x) = P(X = x)

  • f(x)≥0f(x) \ge 0 for all xx
  • ∑all xf(x)=1\sum_{\text{all }x} f(x) = 1 (total probability mass = 1)

Why “mass”? Think of 1 unit of probability being distributed across the possible values — the PMF decides how much mass lands on each value.

Example: Fair die

f(x)=16,x=1,2,3,4,5,6f(x) = \frac{1}{6}, \quad x = 1,2,3,4,5,6

Plot: six bars of equal height (1/61/6).

Example: Biased die

xxf(x)f(x)
11/31/3
21/61/6
31/61/6
41/61/6
51/121/12
61/121/12
Sum1

Plot: bars of different heights (tallest at x=1x=1).

Cumulative Distribution Function (CDF) and Tail Probability

The cumulative distribution function (CDF), F(ϕ)F(\phi), gives the probability that XX is less than or equal to ϕ\phi:

F(ϕ)=P(X≤ϕ)=∑x≤ϕf(x)F(\phi) = P(X \le \phi) = \sum_{x \le \phi} f(x)

The tail probability (complement) is:

Fˉ(x)=P(X>x)=1−F(x)\bar{F}(x) = P(X > x) = 1 - F(x)

Worked example: Fair die

  • F(5)=P(X≤5)=16+16+16+16+16=56F(5) = P(X \le 5) = \frac{1}{6} + \frac{1}{6} + \frac{1}{6} + \frac{1}{6} + \frac{1}{6} = \frac{5}{6}
  • P(X>3)=1−F(3)=1−36=12P(X > 3) = 1 - F(3) = 1 - \frac{3}{6} = \frac{1}{2}
  • P(even)=P(X=2,4,6)=36=12P(\text{even}) = P(X=2,4,6) = \frac{3}{6} = \frac{1}{2}
  • P(2≤X≤5)=P(X=2,3,4,5)=46=23P(2 \le X \le 5) = P(X=2,3,4,5) = \frac{4}{6} = \frac{2}{3}

Worked example: Biased die (probabilities as above)

  • P(X=2)=f(2)=16P(X=2) = f(2) = \frac{1}{6}
  • P(X≤5)=f(1)+f(2)+f(3)+f(4)+f(5)=13+16+16+16+112=1112P(X \le 5) = f(1)+f(2)+f(3)+f(4)+f(5) = \frac{1}{3}+\frac{1}{6}+\frac{1}{6}+\frac{1}{6}+\frac{1}{12} = \frac{11}{12} (Equivalently, 1−f(6)=1−112=11121 - f(6) = 1 - \frac{1}{12} = \frac{11}{12})

Expected Value, Variance, and Standard Deviation

Expected value E[X]E[X] is the probability‑weighted average of all possible values:

E[X]=∑all xx⋅f(x)E[X] = \sum_{\text{all }x} x \cdot f(x)

  • Provides the center of the distribution.
  • It need not be a value the RV can actually take (e.g., fair die: E[X]=3.5E[X] = 3.5, yet you never roll a 3.5).

Variance Var(X)\text{Var}(X) measures the spread — the probability‑weighted average of squared deviations from the mean:

Var(X)=∑all x(x−E[X])2⋅f(x)\text{Var}(X) = \sum_{\text{all }x} \big(x - E[X]\big)^2 \cdot f(x)

Standard deviation σX\sigma_X is the square root of the variance (same units as X):

σX=Var(X)\sigma_X = \sqrt{\text{Var}(X)}

Worked example: Fair die

E[X]=1⋅16+2⋅16+⋯+6⋅16=216=3.5E[X] = 1\cdot\frac{1}{6} + 2\cdot\frac{1}{6} + \cdots + 6\cdot\frac{1}{6} = \frac{21}{6} = 3.5

Var(X)=(1−3.5)2⋅16+(2−3.5)2⋅16+⋯+(6−3.5)2⋅16=3512≈2.92\begin{aligned} \text{Var}(X) &= (1-3.5)^2\cdot\frac{1}{6} + (2-3.5)^2\cdot\frac{1}{6} + \cdots + (6-3.5)^2\cdot\frac{1}{6} \\[2pt] &= \frac{35}{12} \approx 2.92 \end{aligned}

σX=3512≈1.71\sigma_X = \sqrt{\frac{35}{12}} \approx 1.71

Worked example: Biased die

From the PMF table above:

E[X]=1⋅13+2⋅16+3⋅16+4⋅16+5⋅112+6⋅112=3312=2.75\begin{aligned} E[X] &= 1\cdot\frac{1}{3} + 2\cdot\frac{1}{6} + 3\cdot\frac{1}{6} + 4\cdot\frac{1}{6} + 5\cdot\frac{1}{12} + 6\cdot\frac{1}{12} \\ &= \frac{33}{12} = 2.75 \end{aligned} Var(X)=(1−2.75)2⋅13+(2−2.75)2⋅16+(3−2.75)2⋅16+(4−2.75)2⋅16+(5−2.75)2⋅112+(6−2.75)2⋅112=2.6875≈2.69\begin{aligned} \text{Var}(X) &= (1-2.75)^2\cdot\frac{1}{3} + (2-2.75)^2\cdot\frac{1}{6} + (3-2.75)^2\cdot\frac{1}{6} \\ &\quad + (4-2.75)^2\cdot\frac{1}{6} + (5-2.75)^2\cdot\frac{1}{12} + (6-2.75)^2\cdot\frac{1}{12} \\ &= 2.6875 \approx 2.69 \end{aligned}

Exam tip: For any discrete RV, always check that the probabilities sum to 1 before computing expected value or variance. Also remember: E[X]E[X] may be a number the RV never actually takes — it describes the long‑run average, not a possible outcome.

Key takeaways

  • A random variable maps random outcomes to numbers; discrete RVs take countable values, continuous RVs take any value in an interval.
  • The PMF f(x)f(x) gives P(X=x)P(X=x) and sums to 1.
  • The CDF F(x)F(x) gives P(X≤x)P(X\le x); tail probability P(X>x)=1−F(x)P(X>x) = 1-F(x).
  • Expected value E[X]=∑x⋅f(x)E[X] = \sum x \cdot f(x) is the probability‑weighted mean.
  • Variance Var(X)=∑(x−E[X])2f(x)\text{Var}(X) = \sum (x-E[X])^2 f(x); standard deviation = Var(X)\sqrt{\text{Var}(X)}.
  • Both fair and biased dice illustrate that probabilities change the center and spread, even though the possible values stay the same.

Key Concepts: PMF, CDF, Tail Probability, Expectation, Variance

For a discrete random variable XX:

  • Probability Mass Function (PMF): f(x)=P(X=x)f(x) = P(X = x). Non‑negative and sums to 1 over all xx.
  • Cumulative Distribution Function (CDF): F(x)=P(X≤x)=∑t≤xf(t)F(x) = P(X \leq x) = \sum_{t \leq x} f(t).
  • Tail probability: Fˉ(x)=P(X>x)=1−F(x)\bar{F}(x) = P(X > x) = 1 - F(x).
  • Expected value: μ=E[X]=∑x⋅f(x)\mu = E[X] = \sum x \cdot f(x).
  • Variance: σ2=Var(X)=∑(x−μ)2f(x)=E[X2]−μ2\sigma^2 = \text{Var}(X) = \sum (x - \mu)^2 f(x) = E[X^2] - \mu^2; standard deviation σ=σ2\sigma = \sqrt{\sigma^2}.

From Data to PMF

When the true PMF is unknown, estimate it using relative frequencies from observed data:

f(x)≈number of occurrences of xtotal observations.f(x) \approx \frac{\text{number of occurrences of } x}{\text{total observations}}.

This gives a valid PMF because frequencies are non‑negative and sum to 1.


Example 1: Biased Die (Reviewing PMF, CDF, Tail)

A biased die has PMF:

f(1)=f(2)=f(3)=f(4)=16,f(5)=f(6)=112.f(1)=f(2)=f(3)=f(4)=\frac{1}{6},\quad f(5)=f(6)=\frac{1}{12}.
  • Probability of even value: P(even)=f(2)+f(4)+f(6)=16+16+112=512P(\text{even}) = f(2)+f(4)+f(6) = \frac{1}{6}+\frac{1}{6}+\frac{1}{12} = \frac{5}{12}.
  • Probability of values between 2 and 5 inclusive: P(2≤X≤5)=f(2)+f(3)+f(4)+f(5)=16+16+16+112=712P(2 \leq X \leq 5) = f(2)+f(3)+f(4)+f(5) = \frac{1}{6}+\frac{1}{6}+\frac{1}{6}+\frac{1}{12} = \frac{7}{12}.
  • Tail probability P(X>3)P(X > 3): P(X>3)=f(4)+f(5)+f(6)=16+112+112=13P(X > 3) = f(4)+f(5)+f(6) = \frac{1}{6}+\frac{1}{12}+\frac{1}{12} = \frac{1}{3}, or 1−F(3)=1−(f(1)+f(2)+f(3))=1−12=131 - F(3) = 1 - (f(1)+f(2)+f(3)) = 1 - \frac{1}{2} = \frac{1}{3}.

Approach to any question:

  1. Translate the question into an event.
  2. Identify the range of xx values that satisfy the event.
  3. Sum the corresponding f(x)f(x).

Example 2: Defective Pumps at Kirloskar Factory

Setup: XX = number of defective pumps per day. Data from 200 days:

xx (defects)DaysEstimated f(x)f(x)
0800.40
1500.25
2400.20
3100.05
4200.10
≥5\geq 500.00

Questions answered using the PMF:

  • P(exactly 2)=f(2)=0.20P(\text{exactly 2}) = f(2) = 0.20.
  • P(≤3)=f(0)+f(1)+f(2)+f(3)=0.40+0.25+0.20+0.05=0.90P(\leq 3) = f(0)+f(1)+f(2)+f(3) = 0.40+0.25+0.20+0.05 = 0.90.
  • Expected value: E[X]=∑xf(x)=0(0.40)+1(0.25)+2(0.20)+3(0.05)+4(0.10)=1.2 defective pumps/day.E[X] = \sum x f(x) = 0(0.40) + 1(0.25) + 2(0.20) + 3(0.05) + 4(0.10) = 1.2 \text{ defective pumps/day}.
  • Variance and standard deviation: Var(X)=(0−1.2)2(0.40)+(1−1.2)2(0.25)+(2−1.2)2(0.20)+(3−1.2)2(0.05)+(4−1.2)2(0.10)=1.66,σ=1.66≈1.29.\begin{aligned} \text{Var}(X) &= (0-1.2)^2(0.40) + (1-1.2)^2(0.25) + (2-1.2)^2(0.20) + (3-1.2)^2(0.05) + (4-1.2)^2(0.10) \\ &= 1.66, \\ \sigma &= \sqrt{1.66} \approx 1.29. \end{aligned}

Exam tip: The steps are identical for any discrete distribution: obtain f(x)f(x) (here from data), then compute probabilities via summation, expectation as a weighted sum, and variance as the weighted squared deviation.


Example 3: Late Take-offs at Nagpur Airport

Setup: XX = number of planes taking off late between 9 pm and 11 pm. Data from 30 days:

xx (late planes)DaysEstimated f(x)f(x)
030.10
160.20
290.30
360.20
460.20
≥5\geq 500.00

The PMF is derived from the given statements: f(0)=0.10f(0)=0.10, f(1)=0.20f(1)=0.20, P(≥2)=0.70P(\geq 2)=0.70, P(≤3)=0.80P(\leq 3)=0.80, E[X]=2.20E[X]=2.20, Var(X)=1.56\text{Var}(X)=1.56.

Questions answered:

  • P(2 or more)=1−f(0)−f(1)=1−0.10−0.20=0.70P(\text{2 or more}) = 1 - f(0) - f(1) = 1 - 0.10 - 0.20 = 0.70.
  • P(4 or less)=F(3)=f(0)+f(1)+f(2)+f(3)=0.10+0.20+0.30+0.20=0.80P(\text{4 or less}) = F(3) = f(0)+f(1)+f(2)+f(3) = 0.10+0.20+0.30+0.20 = 0.80.
  • Expected value: E[X]=0(0.10)+1(0.20)+2(0.30)+3(0.20)+4(0.20)=2.20 late planes.E[X] = 0(0.10) + 1(0.20) + 2(0.30) + 3(0.20) + 4(0.20) = 2.20 \text{ late planes}.
  • Variance and standard deviation: E[X2]=02(0.10)+12(0.20)+22(0.30)+32(0.20)+42(0.20)=6.40,Var(X)=6.40−(2.20)2=1.56,σ=1.56≈1.25.\begin{aligned} E[X^2] &= 0^2(0.10) + 1^2(0.20) + 2^2(0.30) + 3^2(0.20) + 4^2(0.20) = 6.40, \\ \text{Var}(X) &= 6.40 - (2.20)^2 = 1.56, \\ \sigma &= \sqrt{1.56} \approx 1.25. \end{aligned}

Key Takeaways

  • PMF from data: relative frequencies estimate probabilities; sum to 1.
  • Calculating probabilities: add f(x)f(x) for the event’s xx range; use 1−F(x)1-F(x) for tail probabilities.
  • Expectation: weighted average ∑xf(x)\sum x f(x).
  • Variance & std dev: weighted average of squared deviations; σ\sigma gives spread in original units.
  • Procedure is universal – same steps for any discrete random variable once f(x)f(x) is known.

Linear Combinations of Random Variables

A linear combination of random variables is a new random variable formed by multiplying each original variable by a constant and summing them: Y=a1X1+a2X2+⋯+anXnY = a_1 X_1 + a_2 X_2 + \dots + a_n X_n. Why does this matter? Because quantities like total cost, total profit, portfolio returns, and sample averages are all linear combinations. Understanding their expectation and variance lets us predict average outcomes and measure risk.

Linear Transformation of a Single Variable

If Y=aX+bY = aX + b (a linear function of a single random variable XX), then

  • Expectation: E(Y)=a E(X)+bE(Y) = a\,E(X) + b Intuition: scaling and shifting the variable scales and shifts its mean.
  • Variance: Var⁡(Y)=a2Var⁡(X)\operatorname{Var}(Y) = a^2 \operatorname{Var}(X) Intuition: shifting (+b+b) does not affect spread; scaling (aa) multiplies spread by a2a^2.

Example 1: Temperature conversion Ideal frying temperature in Celsius: XX with E(X)=160E(X)=160, sd⁡(X)=5\operatorname{sd}(X)=5. Convert to Fahrenheit: Y=95X+32Y = \frac{9}{5}X + 32. E(Y)=95⋅160+32=320,Var⁡(Y)=(95)2⋅25=81,sd⁡(Y)=9.E(Y)=\frac{9}{5}\cdot160+32=320,\quad \operatorname{Var}(Y)=\left(\frac{9}{5}\right)^2\cdot25=81,\quad \operatorname{sd}(Y)=9.

Example 2: Real estate price (rupees to dollars) Price in rupees: XX with E(X)=36, ⁣000E(X)=36,\!000, sd⁡(X)=9, ⁣600\operatorname{sd}(X)=9,\!600. Exchange rate 1 USD=80 INR1\,\text{USD}=80\,\text{INR} → Y=180X+0Y = \frac{1}{80}X + 0. E(Y)=36, ⁣00080=450,Var⁡(Y)=1802⋅9, ⁣6002=14, ⁣400,sd⁡(Y)=120.E(Y)=\frac{36,\!000}{80}=450,\quad \operatorname{Var}(Y)=\frac{1}{80^2}\cdot9,\!600^2=14,\!400,\quad \operatorname{sd}(Y)=120.

Sum of Two Random Variables

For Y=X1+X2Y = X_1 + X_2:

  • Expectation: E(Y)=E(X1)+E(X2)E(Y) = E(X_1) + E(X_2) (always additive).
  • Variance: Var⁡(Y)=Var⁡(X1)+Var⁡(X2)+2 Cov⁡(X1,X2)\operatorname{Var}(Y) = \operatorname{Var}(X_1) + \operatorname{Var}(X_2) + 2\,\operatorname{Cov}(X_1, X_2) where Cov⁡(X1,X2)=ρ12 σ1σ2\operatorname{Cov}(X_1, X_2) = \rho_{12}\,\sigma_1\sigma_2.

If X1X_1 and X2X_2 are independent, then ρ12=0\rho_{12}=0 and: Var⁡(Y)=Var⁡(X1)+Var⁡(X2).\operatorname{Var}(Y) = \operatorname{Var}(X_1) + \operatorname{Var}(X_2).

Example: Store sales Store A: E(X)=7, ⁣000E(X)=7,\!000, sd⁡(X)=1, ⁣500\operatorname{sd}(X)=1,\!500 Store B: E(Y)=8, ⁣000E(Y)=8,\!000, sd⁡(Y)=2, ⁣000\operatorname{sd}(Y)=2,\!000 Independent.

  • Total sales S=X+YS = X+Y: E(S)=15, ⁣000E(S)=15,\!000, Var⁡(S)=1, ⁣5002+2, ⁣0002=6, ⁣250, ⁣000\operatorname{Var}(S)=1,\!500^2+2,\!000^2=6,\!250,\!000, sd⁡(S)=2, ⁣500\operatorname{sd}(S)=2,\!500.
  • Difference D=X−YD = X-Y: E(D)=−1, ⁣000E(D)=-1,\!000, Var⁡(D)=1, ⁣5002+2, ⁣0002=6, ⁣250, ⁣000\operatorname{Var}(D)=1,\!500^2+2,\!000^2=6,\!250,\!000, sd⁡(D)=2, ⁣500\operatorname{sd}(D)=2,\!500 (same variance because (−1)2=1(-1)^2=1).

Key insight: For independent variables, the variance of a difference is the sum of variances, not the difference.

General Linear Combination

Let Y=a1X1+a2X2+⋯+anXnY = a_1 X_1 + a_2 X_2 + \dots + a_n X_n. Then:

  • Expectation (always): E(Y)=a1E(X1)+a2E(X2)+⋯+anE(Xn).E(Y) = a_1 E(X_1) + a_2 E(X_2) + \dots + a_n E(X_n).

  • Variance (if independent): Var⁡(Y)=a12Var⁡(X1)+a22Var⁡(X2)+⋯+an2Var⁡(Xn).\operatorname{Var}(Y) = a_1^2 \operatorname{Var}(X_1) + a_2^2 \operatorname{Var}(X_2) + \dots + a_n^2 \operatorname{Var}(X_n).

  • Variance (if not independent): add pairwise covariance terms: Var⁡(Y)=∑i=1nai2Var⁡(Xi)+2∑i<jaiajCov⁡(Xi,Xj).\operatorname{Var}(Y) = \sum_{i=1}^n a_i^2 \operatorname{Var}(X_i) + 2\sum_{i<j} a_i a_j \operatorname{Cov}(X_i, X_j).

Sum and Difference of Two Independent Identical Variables

If X1X_1 and X2X_2 are independent, E(X1)=E(X2)=μE(X_1)=E(X_2)=\mu, Var⁡(X1)=Var⁡(X2)=σ2\operatorname{Var}(X_1)=\operatorname{Var}(X_2)=\sigma^2, then:

VariableExpectationVariance
Y1=X1+X2Y_1 = X_1 + X_22μ2\mu2σ22\sigma^2
Y2=X1−X2Y_2 = X_1 - X_2002σ22\sigma^2

Both have the same variance — the sign of the constant does not affect the squared scaling.

Portfolio of Two Stocks (Correlated)

Shweta Kulkarni invests in Siemens (XX) and Trent (YY). E(X)=3%E(X)=3\%, sd⁡(X)=1.5%\operatorname{sd}(X)=1.5\%; E(Y)=5%E(Y)=5\%, sd⁡(Y)=3%\operatorname{sd}(Y)=3\%; correlation ρ=−0.6\rho=-0.6.

Portfolio return: P=wX+(1−w)YP = wX + (1-w)Y, where ww = fraction in Siemens.

Formulas: E(P)=w⋅3+(1−w)⋅5E(P) = w\cdot3 + (1-w)\cdot5 Var⁡(P)=w2(1.52)+(1−w)2(32)+2w(1−w)(−0.6)(1.5)(3)\operatorname{Var}(P) = w^2(1.5^2) + (1-w)^2(3^2) + 2w(1-w)(-0.6)(1.5)(3)

For w=0.5w=0.5 (equal split): E(P)=4%,Var⁡(P)=1.4625,sd⁡(P)≈1.21%.E(P)=4\%,\quad \operatorname{Var}(P)=1.4625,\quad \operatorname{sd}(P)\approx1.21\%.

Exploring other splits:

wwE(P)E(P)sd⁡(P)\operatorname{sd}(P)
0.54.0%1.21%
0.93.2%1.19%

The 90-10 portfolio is dominated by the 50-50 portfolio: same risk but lower return. After removing dominated choices, the remaining set of portfolios are Pareto optimal — each offers a different risk-return trade-off. Shweta’s optimal split depends on her risk appetite.

Exam tip: Negative correlation reduces portfolio variance. Even without correlation, diversification often lowers risk relative to a single asset.

Sum of nn Independent Identical Variables

Let X1,X2,…,XnX_1, X_2, \dots, X_n be i.i.d. with mean μ\mu and variance σ2\sigma^2. Define S=∑i=1nXiS = \sum_{i=1}^n X_i. Then:

E(S)=nμ,Var⁡(S)=nσ2.E(S) = n\mu,\qquad \operatorname{Var}(S) = n\sigma^2.

Example: Luggage weight Per passenger: μ=30\mu=30 kg, σ=3\sigma=3 kg, n=81n=81 independent. E(total)=81⋅30=2, ⁣430E(\text{total})=81\cdot30=2,\!430 kg, sd⁡(total)=81⋅9=27\operatorname{sd}(\text{total})=\sqrt{81\cdot9}=27 kg.

Average of nn Independent Identical Variables

Define Xˉ=1n∑i=1nXi\bar{X} = \frac{1}{n}\sum_{i=1}^n X_i (the sample mean). Then:

E(Xˉ)=μ,Var⁡(Xˉ)=σ2n,sd⁡(Xˉ)=σn.E(\bar{X}) = \mu,\qquad \operatorname{Var}(\bar{X}) = \frac{\sigma^2}{n}, \qquad \operatorname{sd}(\bar{X}) = \frac{\sigma}{\sqrt{n}}.

Example: Exam scores Per student: μ=70\mu=70, σ=20\sigma=20, n=25n=25 independent. E(Xˉ)=70E(\bar{X})=70, Var⁡(Xˉ)=40025=16\operatorname{Var}(\bar{X})=\frac{400}{25}=16, sd⁡(Xˉ)=4\operatorname{sd}(\bar{X})=4.

Key insight: Averaging reduces variance — the more data points, the more precise the estimate of the mean.

Coefficient of Variation

The coefficient of variation (CV) is a standardized measure of dispersion: CV=standard deviationmean.\text{CV} = \frac{\text{standard deviation}}{\text{mean}}.

It expresses risk relative to the average — useful for comparing variability across different scales.

  • Luggage total: CV=272430≈0.011=1%\text{CV} = \frac{27}{2430} \approx 0.011 = 1\%.
  • Exam average: CV=470≈0.057=5.7%\text{CV} = \frac{4}{70} \approx 0.057 = 5.7\%.

Key takeaways

  • E(aX+b)=aE(X)+bE(aX+b) = aE(X)+b; Var⁡(aX+b)=a2Var⁡(X)\operatorname{Var}(aX+b) = a^2\operatorname{Var}(X).
  • For sums: expectations add always; variances add only if independent; with dependence, include covariance.
  • Portfolio return variance is reduced by negative correlation between assets.
  • The sample mean Xˉ\bar{X} has variance σ2/n\sigma^2/n — a foundation for statistical inference.
  • Coefficient of variation standardizes dispersion for comparing risk across different means.

Recap: Discrete Random Variables

A discrete random variable XX takes only integer values (e.g., 0, 1, 2, …). Its behaviour is fully described by the probability mass function (PMF) f(x)=P(X=x)f(x) = P(X = x). From the PMF we derive:

  • Cumulative distribution function (CDF): F(x)=P(X≤x)=∑t≤xf(t)F(x) = P(X \leq x) = \sum_{t \leq x} f(t)
  • Expectation (mean): E[X]=∑xx f(x)E[X] = \sum_{x} x \, f(x)
  • Variance: Var(X)=∑x(x−E[X])2 f(x)\text{Var}(X) = \sum_{x} (x - E[X])^2 \, f(x)
  • Standard deviation: SD(X)=Var(X)\text{SD}(X) = \sqrt{\text{Var}(X)}

Linear Combinations of Random Variables

For constants a1,a2,…,ana_1, a_2, \ldots, a_n and random variables X1,X2,…,XnX_1, X_2, \ldots, X_n, define Y=a1X1+a2X2+⋯+anXnY = a_1 X_1 + a_2 X_2 + \cdots + a_n X_n. Then:

  • Expectation: E[Y]=a1E[X1]+a2E[X2]+⋯+anE[Xn]E[Y] = a_1 E[X_1] + a_2 E[X_2] + \cdots + a_n E[X_n]
  • Variance: Var(Y)=a12Var(X1)+a22Var(X2)+⋯+an2Var(Xn)+extra terms\text{Var}(Y) = a_1^2 \text{Var}(X_1) + a_2^2 \text{Var}(X_2) + \cdots + a_n^2 \text{Var}(X_n) + \text{extra terms}

The “extra terms” are all pairwise covariances: 2aiajCov(Xi,Xj)2 a_i a_j \text{Cov}(X_i, X_j). If the XiX_i are independent, all covariances are zero and the variance simplifies to the sum of the squared-coefficient times variances.

Key insight: Linearity of expectation always holds; variance only simplifies under independence.

Key takeaways

  • PMF gives probabilities for each outcome; CDF accumulates them.
  • E[X]E[X] is the probability‑weighted average; Var(X)\text{Var}(X) measures spread around that average.
  • For linear combinations: expectation is linear, variance adds covariances (which vanish under independence).

Bernoulli Random Variable

A Bernoulli random variable YY models an experiment with exactly two outcomes: success (Y=1Y=1) or failure (Y=0Y=0). Let the success probability be pp, and q=1−pq = 1-p the failure probability. The PMF is:

f(y)={p,y=11−p=q,y=0f(y) = \begin{cases} p, & y = 1\\ 1-p = q, & y = 0 \end{cases}

Why it matters

Before studying the Binomial distribution (which counts successes in multiple trials), we must master the Bernoulli – it is the building block. Many real‑world events can be classified as success/failure:

ContextSuccess (Y=1)Failure (Y=0)pp
Coin toss (fair)HeadsTails0.5
Sales callCustomer purchasesNo purchasegiven
Customer surveySatisfiedNot satisfiedgiven
Daily demand > 10Demand > 10Demand ≤ 10given

Expectation and Variance

Using the definitions for a discrete random variable:

E[Y]=(1×p)+(0×q)=pE[Y] = (1 \times p) + (0 \times q) = p Var(Y)=(1−p)2⋅p+(0−p)2⋅q=pq(p+q)=pq\text{Var}(Y) = (1-p)^2 \cdot p + (0-p)^2 \cdot q = p q (p + q) = p q SD(Y)=pq\text{SD}(Y) = \sqrt{p q}

Exam tip: The expectation of a Bernoulli is simply pp – the proportion of successes. Its variance is p(1−p)p(1-p); memorise these two formulas.

Worked Example: Biased Die

For this biased die, the event “outcome is 5 or 6” is given probability p=16p=\frac{1}{6}. The individual face probabilities are not needed for the Bernoulli calculation.

  • Success: die shows 5 or 6 → p=16p = \frac{1}{6}
  • Failure: die shows 1-4 → q=56q = \frac{5}{6}
E[Y]=16,Var(Y)=16⋅56=536,SD(Y)=56E[Y] = \frac{1}{6}, \quad \text{Var}(Y) = \frac{1}{6}\cdot\frac{5}{6} = \frac{5}{36}, \quad \text{SD}(Y) = \frac{\sqrt{5}}{6}

Examples from Earlier Contexts

In each case the Bernoulli is defined by partitioning outcomes of a known discrete distribution:

ContextSuccess conditionppE[Y]=pE[Y]=pVar(Y)=p(1−p)\text{Var}(Y)=p(1-p)
Unbiased die≤ 21/31/31/31/32/9≈0.2222/9 \approx 0.222
Biased die≥ 51/61/61/61/65/36≈0.1395/36 \approx 0.139
Defective pumps (Sangli)≤ 1 defective0.650.650.650.650.65×0.35=0.22750.65 \times 0.35 = 0.2275
Late planes (Sonegaon)≤ 2 late0.60.60.60.60.6×0.4=0.240.6 \times 0.4 = 0.24
Insurance call (Priya)Purchase0.20.20.20.20.2×0.8=0.160.2 \times 0.8 = 0.16
Star day (Baburao)Sales > ₹10,0000.30.30.30.30.3×0.7=0.210.3 \times 0.7 = 0.21

How to Create a Bernoulli from Any Discrete Random Variable

Any discrete random variable can be reduced to a Bernoulli by grouping outcomes into two categories. The examples above show exactly this: for the number of defective pumps, “≤1” becomes success and “≥2” failure, preserving the original probabilities.

Key takeaways

  • Bernoulli has two outcomes: Y=1Y=1 (success) with pp, Y=0Y=0 (failure) with q=1−pq=1-p.
  • E[Y]=pE[Y] = p, Var(Y)=pq\text{Var}(Y) = p q – derived directly from the definitions.
  • It is the foundation for the Binomial distribution (multiple independent Bernoulli trials).

Binomial Random Variables and its Distribution

The binomial random variable counts the number of successes in a fixed number of independent Bernoulli trials. Intuitively: if you flip a coin 10 times, how many heads do you get? That count is a binomial random variable.

Binomial Experiment

A binomial experiment has four properties:

  1. A sequence of nn identical and independent trials.
  2. Each trial has exactly two outcomes: success (probability pp) and failure (probability q=1−pq = 1-p).
  3. The probability of success pp is constant across all trials.
  4. The random variable YY = number of successes in nn trials.
ParameterMeaningExample (coin)
nnFixed number of trials10 tosses
ppSuccess probability per trialpp = P(heads)
YYObserved count of successes# heads in 10 tosses

Exam tip: The binomial random variable only counts successes – it ignores the order in which they occur. All arrangements of yy successes and n−yn-y failures are grouped into the single outcome Y=yY=y.

Probability Mass Function (PMF)

The PMF of YY gives the probability of exactly yy successes:

P(Y=y)=(ny) p y (1−p) n−y,y=0,1,…,nP(Y = y) = \binom{n}{y} \, p^{\,y} \, (1-p)^{\,n-y}, \quad y = 0,1,\dots,n

where (ny)=n!y! (n−y)!\displaystyle \binom{n}{y} = \frac{n!}{y!\,(n-y)!} counts the number of ways to arrange yy successes among nn trials.

Why this formula works

  • p yp^{\,y} : probability of yy successes.
  • (1−p) n−y(1-p)^{\,n-y} : probability of n−yn-y failures.
  • (ny)\binom{n}{y} : number of distinct sequences containing exactly yy successes.

Examples

ExperimentDefinition of successppnnYY range
10 coin tossesHeadspp100–10
15 tosses of a fair die{2 or less}\{2\text{ or less}\}1/31/3150–15
10 tosses of a biased die{5 or more}\{5\text{ or more}\}1/61/6100–10

Expectation and Variance

  • Expectation: E[Y]=n pE[Y] = n\,p
  • Variance: Var(Y)=n p q=n p (1−p)\text{Var}(Y) = n\,p\,q = n\,p\,(1-p)
  • Standard deviation: n p (1−p)\sqrt{n\,p\,(1-p)}

Intuition: The binomial random variable is the sum of nn independent Bernoulli(pp) random variables (each with mean pp and variance p(1−p)p(1-p)). Hence the mean and variance simply add up.

Worked example (fair die)

For n=15n=15, p=1/3p=1/3:

E[Y]=15×13=5E[Y] = 15 \times \frac{1}{3} = 5 Var(Y)=15×13×23=103≈3.33\text{Var}(Y) = 15 \times \frac{1}{3} \times \frac{2}{3} = \frac{10}{3} \approx 3.33 SD(Y)=103≈1.83\text{SD}(Y) = \sqrt{\frac{10}{3}} \approx 1.83

This means over many sets of 15 rolls, you would expect about 5 successes (rolling ≤2), give or take about 1.8.

Key takeaways

  • A binomial random variable requires: fixed nn, independent trials, two outcomes, constant pp.
  • PMF: P(Y=y)=(ny)py(1−p)n−y\displaystyle P(Y=y) = \binom{n}{y} p^y (1-p)^{n-y}.
  • Mean =np= np; variance =np(1−p)= np(1-p).
  • The binomial counts successes – it does not track which trials succeeded.
  • Common pitfalls: using binomial when trials are not independent (e.g., sampling without replacement) – use hypergeometric instead.

Exam tip: The binomial PMF sums to 1: ∑y=0n(ny)py(1−p)n−y=1\sum_{y=0}^n \binom{n}{y} p^y (1-p)^{n-y} = 1. This is the binomial theorem in action. Memorise the mean and variance formulas – they appear in nearly every exam problem.

Examples of Binomial Distributions

The binomial distribution models the number of successes in a fixed number of independent trials, each with the same success probability pp. The examples that follow show how to compute probabilities, expectations, and perform inverse calculations using a spreadsheet.

Key formulas (for binomial random variable Y∼Bin(n,p)Y \sim \text{Bin}(n,p))

  • PMF: P(Y=y)=(ny)py(1−p)n−yP(Y=y) = \binom{n}{y} p^y (1-p)^{n-y}
  • Expected value: μ=E[Y]=np\mu = E[Y] = np
  • Variance: σ2=np(1−p)\sigma^2 = np(1-p)
  • Standard deviation: σ=np(1−p)\sigma = \sqrt{np(1-p)}

Excel: BINOM.DIST function

ArgumentMeaningExample
ynumber of successescell reference
nnumber of trials10
pprobability of success per trial1/6
cumulativeFALSE → PMF; TRUE → CDFFALSE gives P(Y=y)P(Y=y)

A template table lists yy from 00 to nn, then computes P(Y=y)P(Y=y) (PMF), P(Y≤y)P(Y \le y) (CDF), and P(Y≥y)P(Y \ge y) (complement) for each yy. The spreadsheet also automatically updates the mean, variance, and standard deviation when nn or pp changes.


Example 1: Biased die (10 tosses, success = 5 or more)

  • Trials: n=10n = 10
  • Success: outcome ≥5\ge 5 on a die ⇒p=1/6\Rightarrow p = 1/6
  • Failure: outcome ≤4\le 4 ⇒q=5/6\Rightarrow q = 5/6
  • Random variable: YY = number of successes in 10 tosses

Computed parameters

  • Mean: E[Y]=np=10×16≈1.67E[Y] = np = 10 \times \frac{1}{6} \approx 1.67
  • Variance: npq=10×16×56≈1.39npq = 10 \times \frac{1}{6} \times \frac{5}{6} \approx 1.39
  • Standard deviation: 1.39≈1.18\sqrt{1.39} \approx 1.18

Probability of at least 3 successes

P(Y≥3)=∑y=310P(Y=y)P(Y \ge 3) = \sum_{y=3}^{10} P(Y=y)

From the spreadsheet, summing the PMF values for y=3,4,…,10y=3,4,\dots,10 gives:

P(Y≥3)≈0.225P(Y \ge 3) \approx 0.225

Inverse probability: find pp such that P(Y≥3)≥0.5P(Y \ge 3) \ge 0.5

Starting from p=1/6p=1/6, the probability P(Y≥3)P(Y \ge 3) is 0.2250.225. By incrementally increasing pp in the spreadsheet cell, the desired probability reaches 0.50.5 when p≈0.259p \approx 0.259.

  • p=0.18→P(Y≥3)≈0.263p = 0.18 \rightarrow P(Y \ge 3) \approx 0.263
  • p=0.20→0.322p = 0.20 \rightarrow 0.322
  • p=0.24→0.44p = 0.24 \rightarrow 0.44
  • p=0.25→0.474p = 0.25 \rightarrow 0.474
  • p=0.259→0.501p = 0.259 \rightarrow 0.501

The goal seek feature in Excel automates this.

Exam tip: An inverse probability problem gives you a target probability and asks for the parameter pp (or sometimes nn) that produces it. The spreadsheet approach (trial and error or goal seek) is efficient but on an exam you may need to reason qualitatively or use a calculator.


Example 2: Kedar Apte’s pump defects (14 days, success = day with ≤1 defect)

  • Trials: n=14n = 14 independent days
  • Success: day with one or fewer defective pumps ⇒p=0.65\Rightarrow p = 0.65
  • Failure: day with two or more defects ⇒q=0.35\Rightarrow q = 0.35
  • Random variable: YY = number of success days in 14 days

Computed parameters

  • Mean: E[Y]=14×0.65=9.1E[Y] = 14 \times 0.65 = 9.1
  • Standard deviation: 14×0.65×0.35≈1.78\sqrt{14 \times 0.65 \times 0.35} \approx 1.78

Probability of at least 10 success days

P(Y≥10)=∑y=1014P(Y=y)P(Y \ge 10) = \sum_{y=10}^{14} P(Y=y)

From the spreadsheet:

yyP(Y=y)P(Y=y)
100.202
110.137
120.069
130.023
140.002

P(Y≥10)≈0.423P(Y \ge 10) \approx 0.423

Inverse probability: find pp such that P(Y≥10)≥0.5P(Y \ge 10) \ge 0.5

Current p=0.65p=0.65 gives P(Y≥10)=0.423P(Y \ge 10)=0.423. Increase pp:

  • p=0.66→0.454p = 0.66 \rightarrow 0.454
  • p=0.67→0.486p = 0.67 \rightarrow 0.486
  • p=0.68→0.519p = 0.68 \rightarrow 0.519

Thus a success probability of approximately 0.6750.675 (between 0.67 and 0.68) yields P(Y≥10)P(Y \ge 10) just over 0.50.5.


How the spreadsheet approach works (diagram)


Key takeaways

  • The binomial distribution is fully determined by nn and pp; all probabilities, mean, and variance follow.
  • Spreadsheet templates (with BINOM.DIST) simplify repeated calculations and allow quick parameter sensitivity analysis.
  • An inverse probability problem fixes a tail probability and solves for pp (or nn) – often by trial and error or goal seek.
  • Both examples illustrate the same workflow: define success → set nn and pp → compute probabilities → answer questions about expectation, standard deviation, and tail events.

Binomial Distribution – Examples

A binomial random variable YY counts the number of successes in nn independent identical Bernoulli trials, each with success probability pp (fail probability q=1−pq=1-p).

  • PMF: P(Y=y)=(ny)pyqn−y,y=0,1,…,n\displaystyle P(Y=y) = \binom{n}{y} p^y q^{n-y}, \quad y=0,1,\dots,n
  • Mean: μ=np\mu = np
  • Variance: σ2=npq\sigma^2 = npq

In practice, cumulative probabilities are obtained from a spreadsheet or binomial tables.


Example 1 – Mangesh Nadkarni (Plane Delays)

  • Success = “day with ≤2 delays” → p=0.6p=0.6, q=0.4q=0.4.
  • Trials: n=30n=30 independent days.
  • YY = # of success days.
EventProbabilityCalculation
At least 18 successesP(Y≥18)=0.578P(Y \ge 18)=0.578Sum P(Y=18)P(Y=18) through P(Y=30)P(Y=30)
At most 10 failures   ⟺  \iff at least 20 successesP(Y≥20)=0.291P(Y \ge 20)=0.291Sum P(Y=20)P(Y=20) through P(Y=30)P(Y=30)
Between 15 and 25 successes (inclusive)P(15≤Y≤25)=0.901P(15 \le Y \le 25)=0.901Sum P(Y=15)P(Y=15) through P(Y=25)P(Y=25)

Example 2 – Priya (Insurance Calls)

  • Success = “call leads to purchase” → p=0.2p=0.2, q=0.8q=0.8.
  • Trials: n=12n=12 calls per day.
  • YY = # of successes.

Expected number of successes: μ=np=12×0.2=2.4\mu = np = 12 \times 0.2 = 2.4. Standard deviation: σ=npq=12×0.2×0.8≈1.39\sigma = \sqrt{npq} = \sqrt{12 \times 0.2 \times 0.8} \approx 1.39.

EventProbabilityInterpretation
At least 6 successes (≥50% conversion)P(Y≥6)=0.019P(Y \ge 6)=0.019Very low chance of a “good” day
At most 3 successesP(Y≤3)=0.795P(Y \le 3)=0.795~80% chance of “bad” day (≤25% conversion)

Example 3 – Baburao (Star Days)

  • Success = “sales >10 000” → p=0.3p=0.3, q=0.7q=0.7.
  • Trials: n=10n=10 days.
  • YY = # of star days.

Expected star days: μ=np=10×0.3=3\mu = np = 10 \times 0.3 = 3.

EventProbability
At most 3 star daysP(Y≤3)=0.65P(Y \le 3)=0.65
Between 3 and 6 star days (inclusive)P(3≤Y≤6)=0.607P(3 \le Y \le 6)=0.607 (or 0.989−0.233=0.6060.989 - 0.233 = 0.606 by subtraction)

Redefining Success – The Critical Pitfall

The same problem context can require different definitions of “success”. Always re‑identify pp and nn for each question.

Example – Kedar Apte (Defective Pumps)

Original success: “≤1 defect per day” → p=0.65p=0.65. But new question: “at least 7 days with no defects” → success = “0 defects” → p=0.4p=0.4 (from the distribution).

  • n=14n=14, P(Y≥7)=0.308P(Y \ge 7)=0.308.

Another question: “at most 3 days with ≥3 defects” → success = “≥3 defects” → p=0.15p=0.15.

  • n=7n=7, P(Y≤3)=0.988P(Y \le 3)=0.988.

Yet another: “at most one defect on all 14 days” → success = “≤1 defect” (p=0.65p=0.65), but event is Y=14Y=14.

  • P(Y=14)=0.002P(Y=14)=0.002.

Exam tip: Never carry forward a previous pp without rechecking what “success” means in the current question. The binomial calculation is mechanical; the hard part is mapping the business question to the correct nn and pp.


More Complex Examples – Mangesh Nadkarni (Plane Delays Revisited)

Based on historical data for late planes, with probabilities treated as known:

QuestionSuccess definitionppnnEventProbability
At least 15 days with ≤1 plane late≤1 plane late0.2+0.1=0.30.2+0.1=0.330Y≥15Y\ge 150.0170.017
At most 5 days with exactly 4 planes lateexactly 4 planes late0.20.210Y≤5Y\le 50.9440.944
Exactly 15 days with ≤2 planes late≤2 planes late0.60.630Y=15Y=150.0780.078

The probability of ≥15 successes with p=0.3p=0.3 is low; the manager is optimistic.


Divide-and-Conquer Procedure

Key takeaways

  • Binomial distribution models counts of successes in a fixed number of independent Bernoulli trials with constant pp.
  • The formulas P(Y=y)=(ny)pyqn−yP(Y=y)=\binom{n}{y}p^y q^{n-y}, μ=np\mu=np, σ2=npq\sigma^2=npq are the core tools.
  • The most common errors are misidentifying pp (success definition) and nn (number of trials). Always start by clarifying these.
  • Use complement or subtraction tricks (e.g., P(Y≤a)=1−P(Y≥a+1)P(Y\le a) = 1-P(Y\ge a+1)) when helpful, but spreadsheet sums are straightforward.
  • Cumulative probabilities (at most, at least, between) are the most frequent exam questions – practice translating business language into probability events.

Recap of Binomial Distribution

A binomial experiment consists of a sequence of identical and independent trials. Each trial has exactly two outcomes — success (probability pp) and failure (probability q=1−pq = 1-p). The probability pp remains constant across trials. The binomial random variable YY counts the number of successes in nn such trials. Equivalently, it is the sum of nn identical and independently distributed (IID) Bernoulli random variables.

If any assumption fails — e.g., success probabilities change or trials are dependent — the experiment is no longer binomial. In practice, other discrete distributions are then used to model the situation.

Key takeaways

  • Binomial: fixed nn, constant pp, independence, two outcomes per trial.
  • Y∼Binomial(n,p)Y \sim \text{Binomial}(n,p) with PMF P(Y=y)=(ny)pyqn−yP(Y=y) = \binom{n}{y} p^y q^{n-y}.
  • E[Y]=npE[Y] = np, Var[Y]=npqVar[Y] = npq.
  • When assumptions break, alternative distributions are needed.

Poisson Distribution

The Poisson random variable counts the number of occurrences of an event over a specified interval of time or space. It is a discrete distribution (values 0,1,2,…0,1,2,\dots) and is often used to model random arrivals, defects, or counts that occur at a constant average rate.

Intuition and Examples

  • Number of customers arriving at a store in 30 minutes.
  • Number of phone calls at a call center in 15 minutes.
  • Number of defects in one kilometre of highway.
  • Number of keyword occurrences on a page of text.

Probability Mass Function (PMF)

Let YY be Poisson with parameter Λ>0\Lambda > 0 (the mean number of events in the interval). The probability of exactly yy events is:

P(Y=y)=e−Λ Λyy!,y=0,1,2,…P(Y=y) = \frac{e^{-\Lambda}\,\Lambda^y}{y!}, \quad y = 0,1,2,\dots

e≈2.71828e \approx 2.71828 is the base of the natural logarithm.

Exam tip: Poisson has only one parameter Λ\Lambda, unlike binomial’s two (n,pn,p).

Expectation and Variance

For a Poisson random variable YY:

E[Y]=Λ,Var[Y]=ΛE[Y] = \Lambda, \qquad Var[Y] = \Lambda

Thus the mean equals the variance — a unique property of the Poisson distribution. The standard deviation is Λ\sqrt{\Lambda}.

Properties of a Poisson Process

A Poisson process is a stochastic process where events occur randomly over time, and the number of events in any interval of length tt follows a Poisson distribution. The process satisfies:

  1. Stationarity – The probability of a given number of events in an interval depends only on the interval’s length, not its start time.
  2. Independence – The occurrence or non‑occurrence of events in disjoint intervals are independent.
  3. Mean equals variance – The number of events in any interval has identical mean and variance.

The time between successive events in a Poisson process follows an exponential distribution (a continuous distribution covered later).

Parameter as a Rate

When interest lies in the number of events over an interval of length tt, the Poisson parameter becomes Λ=λt\Lambda = \lambda t, where λ\lambda is the rate (events per unit time). Then YY – the number of events in time tt – is Poisson with mean λt\lambda t:

P(Y=y)=e−λt (λt)yy!P(Y=y) = \frac{e^{-\lambda t}\,(\lambda t)^y}{y!}

Worked Example: Call Centre Arrivals

Suppose customers call a help desk at a rate of λ=10\lambda = 10 per hour.

  • In a t=2t = 2‑hour period, the number of calls YY is Poisson with Λ=10×2=20\Lambda = 10 \times 2 = 20. E[Y]=20E[Y] = 20, Var[Y]=20Var[Y] = 20.
  • In a t=30t = 30‑minute (0.50.5 hour) period, the number of calls ZZ is Poisson with Λ=10×0.5=5\Lambda = 10 \times 0.5 = 5. E[Z]=5E[Z] = 5, Var[Z]=5Var[Z] = 5.

The PMF in each case uses the corresponding Λ\Lambda.

Connection to the Mathematician

The distribution is named after Baron Siméon Denis Poisson (1781–1840), a French mathematician and physicist.

Key takeaways

  • Poisson models counts of rare events over time or space.
  • PMF: P(Y=y)=e−ΛΛy/y!P(Y=y) = e^{-\Lambda}\Lambda^y / y!, with single parameter Λ\Lambda.
  • E[Y]=Var[Y]=ΛE[Y] = Var[Y] = \Lambda.
  • For a process with rate λ\lambda over interval tt, parameter Λ=λt\Lambda = \lambda t.
  • Key properties: stationarity, independence, mean=variance.

Poisson Distribution: Examples and Excel Implementation

The Poisson distribution models the number of events occurring in a fixed interval of time (or space) when events happen at a constant average rate and independently. The distribution has a single parameter λ\lambda (also denoted as the mean and variance). The probability mass function (PMF) is:

P(Y=y)=e−λλyy!,y=0,1,2,…P(Y = y) = \frac{e^{-\lambda} \lambda^y}{y!}, \quad y = 0, 1, 2, \dots

Setting Up Poisson in Excel

Excel does not have a dedicated Poisson function, but the PMF can be built using standard functions:

CellContentPurpose
D3λ\lambda (mean)Input parameter
E3=D3 (variance)Since Var(Y)=λ\text{Var}(Y)=\lambda
F3=SQRT(E3)Standard deviation
B6:B106y=0,1,2,…,100y = 0, 1, 2, \dots, 100Values of the random variable
C6=EXP(-D3) * D3^B6 / FACT(B6)PMF: P(Y=y)P(Y=y)
D6=C6CDF for y=0y=0 (since P(Y≤0)=P(Y=0)P(Y \le 0)=P(Y=0))
D7=C7 + D6 (drag down)CDF: P(Y≤y)=P(Y≤y−1)+P(Y=y)P(Y \le y) = P(Y \le y-1) + P(Y=y)
E6=1Tail probability for y=0y=0: P(Y≥0)=1P(Y \ge 0)=1
E7=E6 - C6Tail: P(Y≥1)=P(Y≥0)−P(Y=0)P(Y \ge 1) = P(Y \ge 0) - P(Y=0) (drag down for y≥2y\ge 2)

This table allows quick computation and plotting of the PMF for any λ\lambda.


Example 1: Customer Arrivals at an Eatery

At Sheetal Dadava, the number of customers arriving in one hour follows a Poisson distribution with λ=10\lambda = 10. Questions and solutions:

QuestionProbability Statementλ\lambdaExcel LookupResult
Exactly 8 customers in 1 hourP(Y=8)P(Y=8)10Column C, y=8y=80.113
More than 10 customers in 1 hourP(Y>10)=P(Y≥11)P(Y > 10) = P(Y \ge 11)10Column E, y=11y=110.417
Exactly 4 customers in 30 minutesP(Y=4)P(Y=4)5*Column C, y=4y=40.175
At most 10 customers in 30 minutesP(Y≤10)P(Y \le 10)5*Column D, y=10y=100.986

*λ\lambda scales with time: for a 30‑minute interval, λ=10×3060=5\lambda = 10 \times \frac{30}{60} = 5 (the process is Poisson; the rate is proportional to interval length).

Exam tip: When the time interval changes, always rescale λ\lambda. “More than 10” means Y>10Y > 10, which equals Y≥11Y \ge 11. Reading the wrong row is a common trap.


Example 2: Social Media Likes

Sriram’s weekly likes follow a Poisson distribution with λ=14\lambda = 14. Similarly, daily likes use λ=14/7=2\lambda = 14 / 7 = 2.

QuestionProbability Statementλ\lambdaExcel LookupResult
At least 10 likes in a weekP(Y≥10)P(Y \ge 10)14Column E, y=10y=100.891
Between 10 and 20 likes in a week (inclusive)P(10≤Y≤20)=∑y=1020P(Y=y)P(10 \le Y \le 20) = \sum_{y=10}^{20} P(Y=y)14Sum of column C for y=10y=10 to 20200.843
At least 5 likes on a given dayP(Y≥5)P(Y \ge 5)2Column E, y=5y=50.053
No likes on a given dayP(Y=0)P(Y=0)2Column C, y=0y=00.135

Exam tip: For “between … and …” always check inclusiveness. If both endpoints are included, sum the PMF values directly. The tail probability column gives P(Y≥y)P(Y \ge y), not P(Y>y)P(Y > y).


Key Takeaways – Poisson Examples

  • The Poisson PMF is P(Y=y)=e−λλy/y!P(Y=y) = e^{-\lambda}\lambda^y / y!; λ\lambda = mean = variance.
  • Excel setup (no built‑in function) uses EXP, ^, and FACT to compute PMF, then cumulative and tail probabilities by iteration.
  • When the time interval changes, λ\lambda scales proportionally (e.g., half the time → half the λ\lambda).
  • Translating a business question into a probability statement is critical: “more than 10” = Y>10=Y≥11Y > 10 = Y \ge 11; “at most 10” = Y≤10Y \le 10; “at least 10” = Y≥10Y \ge 10.
  • For range probabilities (between a and b inclusive), sum the PMF for each yy in that range.

Examples of Poisson Distribution (Call Center)

The Poisson distribution models the number of events (e.g., phone calls) occurring in a fixed interval of time when events happen independently at a constant average rate. Its only parameter is the mean λ\lambda (the average number of events per interval). The probability mass function is

P(Y=y)=e−λλyy!,y=0,1,2,…P(Y = y) = \frac{e^{-\lambda} \lambda^y}{y!}, \quad y = 0,1,2,\dots

A classic application is call‑center staffing. A call center receives calls according to a Poisson process: the number of calls in any interval depends only on the interval length, not on the time of day, and successive intervals are independent. (More advanced models use non‑stationary Poisson processes with time‑varying rates.)

Scaling the Poisson Parameter λ\lambda

The average rate is given for one interval but may be needed for another. If the rate is constant, λ\lambda scales linearly with the interval length.

Example: The call center u.net averages 6 calls per 15 minutes. Rate per minute = 6/15=0.46 / 15 = 0.4 calls/min. For any tt‑minute interval, λ=0.4⋅t\lambda = 0.4 \cdot t.

Intervalλ\lambda calculationλ\lambdaQueryResult (from PMF)
5 minutes6/15×56/15 \times 52P(Y=3)P(Y = 3)0.180.18
15 minutes(given)6P(Y=10)P(Y = 10)0.0410.041
3 minutes6/15×36/15 \times 31.2P(Y=0)P(Y = 0)0.3010.301

Worked Calculations

1. Probability of 3 calls in 5 minutes

λ=2\lambda = 2, y=3y=3:

P(Y=3)=e−2233!=0.1353×86≈0.1804P(Y=3) = \frac{e^{-2} 2^3}{3!} = \frac{0.1353 \times 8}{6} \approx 0.1804

There is roughly an 18% chance of exactly 3 calls in a 5‑minute window.

2. Probability of 10 calls in 15 minutes

λ=6\lambda = 6, y=10y=10:

P(Y=10)=e−661010!≈0.0413P(Y=10) = \frac{e^{-6} 6^{10}}{10!} \approx 0.0413

A 4.1% chance – a relatively low probability, suggesting such high demand may strain staffing.

3. Probability of no calls during a 3‑minute break

λ=1.2\lambda = 1.2, y=0y=0:

P(Y=0)=e−1.2≈0.3012P(Y=0) = e^{-1.2} \approx 0.3012

Only a 30% chance that a 3‑minute break will be uninterrupted. The agent is likely to be called back before the break ends.

Key Takeaways

  • The Poisson distribution depends entirely on the mean λ\lambda for the given interval.
  • Scaling: λ\lambda for a new interval = (original rate per unit time) × (new interval length).
  • Probabilities are computed directly from the PMF (or a spreadsheet).
  • Call‑center applications illustrate how randomness affects service quality and staffing decisions.

Exam tip: Always check that the λ\lambda you use matches the interval of the question. A common mistake is to use the original λ\lambda without rescaling.

Poisson Distribution – Worked Examples

The Poisson distribution models the count of events occurring in a fixed interval of time or space when events happen independently at a constant average rate. The single parameter λ\lambda (often called the rate parameter) equals both the mean and variance:

P(Y=y)=e−λλyy!,y=0,1,2,…E[Y]=λ,Var⁡(Y)=λP(Y = y) = \frac{e^{-\lambda} \lambda^y}{y!}, \quad y = 0,1,2,\dots \qquad \mathbb{E}[Y] = \lambda,\quad \operatorname{Var}(Y) = \lambda

When the interval changes, λ\lambda scales proportionally.

Example 1: Mall Visitor Arrivals

Context: Prozone Mall, Aurangabad. Average visits during busy hour = 180 per hour. Manager Mansoor wants the probability of 10–20 people accumulating in a 10‑minute window.

Adjusting λ\lambda for the new interval

  • 180 per hour →\rightarrow 10 minutes: λ=180×1060=30\lambda = 180 \times \frac{10}{60} = 30 visitors per 10 min
  • Similarly, if average is 150/hr → λ=25\lambda = 25; if 210/hr → λ=35\lambda = 35

Probabilities (from Poisson PMF)

Average per hourλ\lambda (10 min)P(10≤Y≤20)P(10 \le Y \le 20)
150250.73
180300.526
210350.225

Why does the probability decrease as λ\lambda increases? The Poisson distribution becomes roughly symmetric around its mean. As λ\lambda shifts right (25 → 30 → 35), the fixed window [10,20] captures less probability mass because the distribution moves to higher counts.

Exam tip: Always re‑scale λ\lambda to match the interval of interest. The probability over the same numeric range can rise or fall as λ\lambda changes – your intuition about “more arrivals = more in the window” fails when the window is far from the mean.


Example 2: Highway Defects

Context: Maharashtra State Highway No. 3. Average defects = 3 per km. Count of defects in a stretch follows Poisson.

Probabilities for 1 km stretch (λ=3\lambda=3)

  • P(Y≥3)=0.577P(Y \ge 3) = 0.577 (tail probability)
  • P(Y≤5)=0.916P(Y \le 5) = 0.916 (cumulative)

Probabilities for an 8 km stretch (λ=3×8=24\lambda = 3 \times 8 = 24)

  • P(Y≤40)=0.999P(Y \le 40) = 0.999
  • P(Y≥20)=0.82P(Y \ge 20) = 0.82

Common mistake: Trying to compute P(Y≤40 in 8 km)P(Y \le 40 \text{ in 8 km}) as [P(Y≤5 in 1 km)]8[P(Y \le 5 \text{ in 1 km})]^8 (or 0.91650.916^5 as initially attempted). This is wrong because it assumes every 1‑km segment must have ≤5 defects – the total can be ≤40 even if some segments exceed 5, as long as others compensate.

Exam tip: For a Poisson process, the total count over combined independent intervals is Poisson with λtotal=λunit×number of units\lambda_{\text{total}} = \lambda_{\text{unit}} \times \text{number of units}. NEVER multiply probabilities of sub‑intervals.


Example 3: Social Media Flags

Context: NIA analyst Nitin monitors flags per conversation. Average = 5 flags per conversation. Flags occur independently and at constant rate.

Single conversation (λ=5\lambda=5)

  • P(2≤Y≤8)=0.891P(2 \le Y \le 8) = 0.891
  • P(Y≥10)=0.032P(Y \ge 10) = 0.032 (less than 3% – low alert)

Two consecutive conversations (λ=5+5=10\lambda = 5+5 = 10)

  • P(4≤Y≤16)=0.963P(4 \le Y \le 16) = 0.963
  • P(Y<4)=P(Y≤3)=0.01P(Y < 4) = P(Y \le 3) = 0.01 (very unlikely)

Note that “less than 4 flags in two conversations” means Y≤3Y \le 3, not Y≤4Y \le 4.

Result summary table

Scenarioλ\lambdaQueryProbability
1 conversation52≤Y≤82 \le Y \le 80.891
1 conversation5Y≥10Y \ge 100.032
2 conversations104≤Y≤164 \le Y \le 160.963
2 conversations10Y≤3Y \le 30.01

Key Properties of the Poisson Distribution

  • Constant rate – probability of an event is the same in any two intervals of equal length.
  • Independence – occurrence in one interval is independent of occurrence in any other non‑overlapping interval.
  • Mean = Variance = λ\lambda.

These properties make Poisson suitable for counts over time (arrivals, calls, flags) and over space (defects per km, potholes, errors in wafers or text).


Connection Between Binomial and Poisson

When pp (success probability) is small and nn is large, the binomial distribution Bin⁡(n,p)\operatorname{Bin}(n,p) approximates a Poisson with λ=np\lambda = np.

  • Binomial: E[Y]=np\mathbb{E}[Y]=np, Var⁡(Y)=np(1−p)\operatorname{Var}(Y)=np(1-p)
  • Poisson: E[Y]=Var⁡(Y)=λ\mathbb{E}[Y]=\operatorname{Var}(Y)=\lambda

If pp is small, 1−p≈11-p \approx 1, so binomial variance ≈ np=λnp = \lambda, matching Poisson. Many Poisson distributions also appear symmetric and bell‑shaped for moderate λ\lambda – a preview of the Central Limit Theorem.

Exam tip: Use the Poisson approximation to the binomial when nn is large, pp is small, and npnp is moderate (rule of thumb: n≥20n \ge 20, p≤0.05p \le 0.05).


Key Takeaways

  • Poisson models counts of rare, independent events over time/space; one parameter λ\lambda = mean = variance.
  • When interval changes, λ\lambda scales proportionally; always recompute for the exact interval.
  • Probabilities computed via PMF, cumulative (CDF), or tail – spreadsheet tools are helpful.
  • Common pitfalls: multiplying probabilities across intervals instead of summing means; confusing P(Y<k)P(Y < k) with P(Y≤k)P(Y \le k).
  • The binomial approximates Poisson when pp is small (λ=np\lambda = np).