Probability Distributions (Normal and Binomial)

Probability Distributions (Normal and Binomial)

Definition: A probability distribution is a function describing how likely each possible outcome of a random process is; the normal distribution models continuous, symmetric, bell-shaped data, while the binomial distribution models the count of successes across a fixed number of independent yes/no trials.

How It Works

  • The normal (Gaussian) distribution is the familiar bell curve, fully described by just two numbers: its mean μ (where the peak sits) and its standard deviation σ (how spread out it is).
  • The normal distribution is symmetric around μ, and its shape flattens and widens as σ increases, or narrows and heightens as σ decreases, while always enclosing the same total area of 1.
  • The 68-95-99.7 rule (empirical rule) states that for any normal distribution, about 68% of values fall within 1σ of the mean, about 95% within 2σ, and about 99.7% within 3σ.
  • The binomial distribution models the number of successes in n independent trials, each with the same success probability p (a “success” here just means whichever of two outcomes you’re counting, like heads on a coin flip).
  • Each binomial trial must be independent (one outcome doesn’t affect another) and have a constant success probability p across every trial for the distribution to apply correctly.
  • The binomial distribution’s mean is simply n × p, the expected number of successes across all the trials.
  • Unlike the normal distribution’s smooth curve, the binomial distribution is discrete: only whole-number outcomes (0 successes, 1 success, 2 successes, …) are possible, so it’s naturally drawn as a bar chart rather than a continuous curve.
  • As n grows large, a binomial distribution’s shape increasingly resembles a normal distribution, an informal preview of the Central Limit Theorem, one of the most important results in all of statistics.
  • A probability density (like the normal curve’s height) is not itself a probability and can exceed 1; a probability mass (like a single binomial bar) is a true probability and can never exceed 1.
  • Both distributions are used constantly to answer the same basic question in different settings: “how likely is this particular outcome, or range of outcomes, to occur by chance?”

Illustration

Normal Distribution

μ = 0.0, σ = 1.0 68% within ±1.0, 95% within ±2.0, 99.7% within ±3.0

Toggle between the normal curve (adjust μ and σ) and the binomial bar chart (adjust n and p). Both redraw live from the actual formulas, not a lookup table.

var muIn = document.getElementById(‘dist-mu’), sigmaIn = document.getElementById(‘dist-sigma’); var nIn = document.getElementById(‘dist-n’), pIn = document.getElementById(‘dist-p’); var muOut = document.getElementById(‘dist-mu-out’), sigmaOut = document.getElementById(‘dist-sigma-out’); var nOut = document.getElementById(‘dist-n-out’), pOut = document.getElementById(‘dist-p-out’); var muLabel = document.getElementById(‘dist-mu-label’), sigmaLabel = document.getElementById(‘dist-sigma-label’); var nLabel = document.getElementById(‘dist-n-label’), pLabel = document.getElementById(‘dist-p-label’); var toggleBtn = document.getElementById(‘dist-toggle’); var resetBtn = document.getElementById(‘dist-reset’);

function fmt(n) { return (Math.round(n * 100) / 100).toFixed(2).replace(/.00$/, ‘.0’); }

var scaleX = 18, ox = 210, oy = 380; var densityScale = 190;

function normalPdf(x, mu, sigma) { var coeff = 1 / (sigma * Math.sqrt(2 * Math.PI)); var exponent = -((x - mu) * (x - mu)) / (2 * sigma * sigma); return coeff * Math.exp(exponent); }

// Iterative binomial coefficient C(n,k) — avoids computing raw factorials // of n up to 30, which stays accurate but grows unnecessarily large. // Verified: choose(5,2) below traces to 5 -> 10, the correct value. function choose(n, k) { if (k < 0 || k > n) return 0; k = Math.min(k, n - k); var result = 1; for (var i = 0; i < k; i++) { result = result * (n - i) / (i + 1); } return result; }

function binomialPmf(n, p, k) { return choose(n, k) * Math.pow(p, k) * Math.pow(1 - p, n - k); }

function drawNormal() { var mu = parseFloat(muIn.value), sigma = parseFloat(sigmaIn.value); muOut.textContent = fmt(mu); sigmaOut.textContent = fmt(sigma);

var pts = [];
for (var x = -11; x <= 11.001; x += 0.2) {
  var y = normalPdf(x, mu, sigma);
  var px = ox + x * scaleX;
  var py = oy - y * densityScale;
  pts.push((pts.length === 0 ? 'M' : 'L') + px.toFixed(1) + ',' + py.toFixed(1));
}
curve.setAttribute('d', pts.join(' '));
var meanPx = ox + mu * scaleX;
meanLine.setAttribute('x1', meanPx); meanLine.setAttribute('x2', meanPx);
meanLine.setAttribute('y1', 40); meanLine.setAttribute('y2', 380);

line1.textContent = 'μ = ' + fmt(mu) + ', σ = ' + fmt(sigma);
line2.textContent = '68% within ±' + fmt(sigma) + ', 95% within ±' + fmt(2 * sigma) + ', 99.7% within ±' + fmt(3 * sigma);

}

function drawBinomial() { var n = parseInt(nIn.value, 10), p = parseFloat(pIn.value); nOut.textContent = n; pOut.textContent = fmt(p);

while (binomialGroup.firstChild) binomialGroup.removeChild(binomialGroup.firstChild);

var probs = [];
for (var k = 0; k <= n; k++) probs.push(binomialPmf(n, p, k));
var maxProb = Math.max.apply(null, probs);
var heightFactor = maxProb > 0 ? 330 / maxProb : 0;

var left = 30, right = 400, totalWidth = right - left;
var slot = totalWidth / (n + 1);
var barWidth = Math.max(1, slot * 0.75);

var mostLikelyK = 0;
for (k = 0; k <= n; k++) if (probs[k] > probs[mostLikelyK]) mostLikelyK = k;

for (k = 0; k <= n; k++) {
  var barX = left + k * slot + (slot - barWidth) / 2;
  var barHeight = probs[k] * heightFactor;
  var barY = oy - barHeight;
  var rect = document.createElementNS(svgNS, 'rect');
  rect.setAttribute('x', barX.toFixed(1));
  rect.setAttribute('y', barY.toFixed(1));
  rect.setAttribute('width', barWidth.toFixed(1));
  rect.setAttribute('height', Math.max(0, barHeight).toFixed(1));
  rect.setAttribute('fill', k === mostLikelyK ? 'var(--brass-bright)' : 'var(--accent)');
  binomialGroup.appendChild(rect);

  // Label every bar when there is room; thin them out for large n so
  // labels don't collide into an unreadable smear.
  if (n <= 15 || k % Math.ceil(n / 15) === 0) {
    var label = document.createElementNS(svgNS, 'text');
    label.setAttribute('x', (barX + barWidth / 2).toFixed(1));
    label.setAttribute('y', (oy + 12).toFixed(1));
    label.setAttribute('text-anchor', 'middle');
    label.setAttribute('class', 'math-label-sm');
    label.textContent = String(k);
    binomialGroup.appendChild(label);
  }
}

line1.textContent = 'n = ' + n + ', p = ' + fmt(p) + ', mean (np) = ' + fmt(n * p);
line2.textContent = 'Most likely outcome: k = ' + mostLikelyK + ' (P = ' + fmt(probs[mostLikelyK]) + ')';

}

function update() { if (mode === ‘normal’) { normalGroup.style.display = ”; binomialGroup.style.display = ‘none’; titleTxt.textContent = ‘Normal Distribution’; drawNormal(); } else { normalGroup.style.display = ‘none’; binomialGroup.style.display = ”; titleTxt.textContent = ‘Binomial Distribution’; drawBinomial(); } }

toggleBtn.addEventListener(‘click’, function () { mode = mode === ‘normal’ ? ‘binomial’ : ‘normal’; toggleBtn.textContent = mode === ‘normal’ ? ‘Switch to Binomial’ : ‘Switch to Normal’; muLabel.style.display = mode === ‘normal’ ? ” : ‘none’; sigmaLabel.style.display = mode === ‘normal’ ? ” : ‘none’; nLabel.style.display = mode === ‘binomial’ ? ” : ‘none’; pLabel.style.display = mode === ‘binomial’ ? ” : ‘none’; update(); });

[muIn, sigmaIn, nIn, pIn].forEach(function (el) { el.addEventListener(‘input’, update); }); resetBtn.addEventListener(‘click’, function () { muIn.value = 0; sigmaIn.value = 1; nIn.value = 10; pIn.value = 0.5; update(); });

update(); })();

Under the Hood

The two governing formulas, stated precisely:

Normal PDF:     f(x) = ( 1 / (sigma * sqrt(2*pi)) ) * e^( -(x-mu)^2 / (2*sigma^2) )
Binomial PMF:   P(X = k) = C(n, k) * p^k * (1-p)^(n-k)

where C(n, k) = n! / (k! (n-k)!) is the number of ways to choose k successes among n trials.

Worked Example 1: Binomial probability by hand. Given: 5 coin flips, probability of exactly 3 heads (n=5, p=0.5, k=3). Step 1: C(5,3) = 5!/(3!2!) = 10. Step 2: P(X=3) = 10 × (0.5)³ × (0.5)² = 10 × 0.125 × 0.25. Answer: P(X=3) = 0.3125 (31.25%).

Worked Example 2: Applying the 68-95-99.7 rule. Given: test scores are normally distributed with μ=100, σ=15. Step 1: 68% of scores fall within 1σ, so between 100-15=85 and 100+15=115. Step 2: 95% of scores fall within 2σ, so between 100-30=70 and 100+30=130. Answer: about 95% of test-takers scored between 70 and 130.

History

  • Jacob Bernoulli developed the mathematics of the binomial distribution, published posthumously in 1713 in his work “Ars Conjectandi.”
  • Abraham de Moivre discovered in 1733 that the binomial distribution, for a large number of trials, could be closely approximated by a smooth bell-shaped curve, an early hint of what would become the normal distribution.
  • Carl Friedrich Gauss independently derived and popularized the normal distribution in 1809 while analyzing errors in astronomical measurements, which is why it is often called the “Gaussian” distribution.
  • Adolphe Quetelet applied the normal distribution broadly to human measurements (height, chest size) in the 1830s, popularizing the idea of an “average man” and helping spread the normal distribution’s use far beyond astronomy.
  • The formal Central Limit Theorem, explaining precisely why so many different distributions converge toward the normal shape, was proven in increasingly general forms through the 19th and early 20th centuries.

Why It Matters

  • Standardized testing (SAT, IQ scores) is calibrated and reported against a normal distribution, with percentile scores derived directly from it.
  • Manufacturing quality control uses the 68-95-99.7 rule to set tolerance limits and flag products that fall statistically outside the expected range.
  • Clinical trials and A/B testing frequently model pass/fail or yes/no outcomes with the binomial distribution to judge whether an observed effect is likely real or due to chance.
  • Insurance and actuarial science rely on both distributions to price risk, from modeling claim amounts (often approximately normal) to modeling event counts (often binomial or related discrete distributions).
  • Machine learning models frequently assume normally distributed noise or errors as a simplifying, and often reasonably accurate, assumption.

Common Pitfalls

  • Assuming real-world data is always normally distributed. Many common measurements, like income or city population, are heavily skewed instead.
  • Applying the binomial distribution when trials are not actually independent or when p changes between trials, both of which break its core assumptions.
  • Misremembering the 68-95-99.7 percentages, or applying them to a distribution that isn’t actually normal.
  • Confusing a probability density (the normal curve’s height, which can exceed 1) with an actual probability (which never can).
  • Forgetting that the binomial distribution is discrete. Asking for “the probability of exactly 3.5 successes” is meaningless; only whole-number outcomes exist.

Comparison

FeatureNormalBinomial
TypeContinuousDiscrete
Parametersμ (mean), σ (spread)n (trials), p (success probability)
ShapeAlways a smooth bell curveBar chart, bell-shaped only for large n
Y-axis meaningDensity (can exceed 1)Probability mass (0 to 1)

FAQ

Why can the height of a normal curve exceed 1 if probabilities can’t? The curve’s height is a probability density, not a probability itself. Probability corresponds to the area under a stretch of the curve, not the height at any single point, and total area under the whole curve is always exactly 1 no matter how tall or short the peak is.

When does a binomial distribution start looking normal? Informally, once n is reasonably large and p isn’t too close to 0 or 1 (a common rule of thumb is np and n(1-p) both at least 5 or so), the binomial’s bar-chart shape becomes visually very close to a smooth bell curve.

Example

Flipping a fair coin 20 times and counting heads follows a binomial distribution centered at a mean of 10 (n×p = 20×0.5); with n this large, the bar chart already looks noticeably bell-shaped, a small-scale preview of the Central Limit Theorem at work.

Dig deeper