Probability Theory

Probability Theory

Definition: Probability is a number from 0 to 1 that measures how likely an event is to occur, where 0 means impossible and 1 means certain, assigned to outcomes drawn from a well-defined sample space of everything that could happen.

How It Works

  • The sample space (S) is the set of every possible outcome of a random process — {1, 2, 3, 4, 5, 6} for one die roll — and an event is any subset of that sample space, such as “rolling an even number” = {2, 4, 6}.
  • A probability is always a number from 0 to 1: 0 means an event never happens, 1 means it always happens, and everything else falls somewhere between rare and near-certain.
  • For equally likely outcomes, P(event) = (outcomes in the event) / (total outcomes in the sample space) — e.g. P(even) = 3/6 = 0.5 for a fair die.
  • The addition rule gives the probability that A or B (or both) occurs: P(A or B) = P(A) + P(B) - P(A and B). Subtracting P(A and B) corrects for counting any shared outcome twice.
  • If A and B can never happen together, they are mutually exclusive, P(A and B) = 0, and the addition rule simplifies to P(A or B) = P(A) + P(B).
  • Two events are independent when one occurring has no effect on the probability of the other; for independent events specifically, P(A and B) = P(A) × P(B), the multiplication rule.
  • Conditional probability, written P(A|B) and read “the probability of A given B,” is A’s probability once B is already known to have happened: P(A|B) = P(A and B) / P(B).
  • Independence has an equivalent conditional definition: A and B are independent exactly when P(A|B) = P(A), meaning learning that B happened doesn’t move the needle on A at all.
  • The complement rule states P(not A) = 1 - P(A). An event and its complement together account for the entire sample space and never overlap, so their probabilities must sum to exactly 1.
  • Mutually exclusive and independent describe two different, almost opposite relationships between events; outside one degenerate case, a pair of events cannot be both at once (developed fully below).
  • Every outcome’s probability in a sample space must sum to exactly 1 — a fair six-sided die assigns 1/6 to each of six faces, and six copies of 1/6 add up to 1.

Illustration

Event A Event B A only 0.28 A∩B 0.12 B only 0.18 P(A) = 0.40 P(B) = 0.30 P(A∩B) = 0.12 P(A∪B) = 0.58 P(A|B) = 0.40 Independent
Drag the sliders to change P(A), P(B), and their overlap P(A∩B); the region values and readout recompute instantly. The two circles' size and position stay fixed — making the drawn area exactly proportional to a probability is a hard inverse-geometry problem, so the true numbers are shown as text instead of as area. The overlap slider is a proxy too: it clamps to whatever range is actually valid for the current P(A) and P(B), including a floor, since once P(A) + P(B) exceeds 1 the two events are mathematically forced to overlap by at least a little.

function fmt(n) { return (Math.round(n * 100) / 100).toFixed(2).replace(/.00$/, ‘.0’); }

function update() { var pA = parseFloat(aIn.value); var pB = parseFloat(bIn.value); var rawCap = parseFloat(capIn.value);

// The overlap can never exceed the smaller of P(A) and P(B) -- a ceiling
// -- and can never drop below P(A)+P(B)-1 either -- a floor. Below that
// floor, P(A or B) would have to exceed 1, which no real probability
// can do. Clamp the raw slider value into that valid window instead of
// fighting the <input>'s own min/max attributes.
var maxCap = Math.min(pA, pB);
var minCap = Math.max(0, pA + pB - 1);
var pCap = Math.min(Math.max(rawCap, minCap), maxCap);

aOut.textContent = fmt(pA);
bOut.textContent = fmt(pB);
capOut.textContent = fmt(pCap);

var aOnly = pA - pCap;
var bOnly = pB - pCap;
var union = pA + pB - pCap;
var cond = pCap / pB;

aOnlyVal.textContent = fmt(aOnly);
bOnlyVal.textContent = fmt(bOnly);
capVal.textContent = fmt(pCap);

paTxt.textContent = 'P(A) = ' + fmt(pA);
pbTxt.textContent = 'P(B) = ' + fmt(pB);
pcapTxt.textContent = 'P(A∩B) = ' + fmt(pCap);
unionTxt.textContent = 'P(A∪B) = ' + fmt(union);
condTxt.textContent = 'P(A|B) = ' + fmt(cond);

var independent = Math.abs(pCap - pA * pB) <= 0.02;
var mutex = pCap <= 0.01;
var label = independent ? 'Independent' : 'Not independent';
if (mutex) label += ' · Mutually exclusive';
classifyTxt.textContent = label;

}

[aIn, bIn, capIn].forEach(function (el) { el.addEventListener(‘input’, update); }); resetBtn.addEventListener(‘click’, function () { aIn.value = 0.4; bIn.value = 0.3; capIn.value = 0.12; update(); });

update(); })();

Mutually exclusive and independent sound like they should be close cousins — both seem to describe two events that have nothing to do with each other — but they capture opposite relationships, and mixing them up is one of the most common mistakes in introductory probability.

Two events are mutually exclusive when they share no outcomes: if A happens, B is guaranteed not to, and vice versa. Knowing A occurred tells you everything about B — that it certainly did not happen. That is about as far from “no effect on each other” as two events can get. Independence requires the opposite: knowing A occurred must tell you nothing about B, formally P(A|B) = P(A).

The two conditions actually collide algebraically. Mutually exclusive means P(A and B) = 0. Independent means P(A and B) = P(A) × P(B). If an event pair satisfies both at once, then P(A) × P(B) = 0. A product of two real numbers is zero only when at least one factor is zero, so this forces P(A) = 0 or P(B) = 0 — one of the two events has to be impossible from the start. Outside that degenerate edge case, any two events that both have a real chance of happening can be mutually exclusive, or they can be independent, but never both. Try it on the sliders above: push P(A∩B) down toward zero (mutually exclusive) while P(A) and P(B) both sit at healthy values like 0.4 and 0.3, and the “Independent” label always switches off, because 0.4 × 0.3 = 0.12, not 0.

Under the Hood

Precisely stated, the three central rules of two-event probability:

P(A or B) = P(A) + P(B) - P(A and B)        addition rule (always true)
P(A or B) = P(A) + P(B)                     special case: A, B mutually exclusive, so P(A and B) = 0

P(A|B) = P(A and B) / P(B)                  conditional probability (requires P(B) > 0)
P(B|A) = P(A and B) / P(A)                  the reverse conditional (requires P(A) > 0) -- not generally equal to P(A|B)

P(A and B) = P(A) × P(B)                    multiplication rule -- valid ONLY if A and B are independent
Independence test:  P(A and B) = P(A) × P(B)   <=>   P(A|B) = P(A)   <=>   P(B|A) = P(B)

Worked Example 1: Addition rule with real overlap. Given: in a class of 40 students, 18 play soccer (A), 14 play basketball (B), and 6 play both. Step 1: P(A) = 18/40 = 0.45, P(B) = 14/40 = 0.35, P(A and B) = 6/40 = 0.15. Step 2: P(A or B) = P(A) + P(B) - P(A and B) = 0.45 + 0.35 - 0.15. Answer: P(A or B) = 0.65 — 65% of the class (26 of 40 students) plays at least one of the two sports.

Worked Example 2: Conditional probability from an inspection log. Given: of 200 manufactured parts, 150 pass a first inspection (B); of those 150, 120 also pass a stricter secondary test (A and B together, since the secondary test only runs on parts that already passed). Step 1: P(B) = 150/200 = 0.75. P(A and B) = 120/200 = 0.60. Step 2: P(A|B) = P(A and B) / P(B) = 0.60 / 0.75. Answer: P(A|B) = 0.80 — given that a part passed the first inspection, there is an 80% chance it also clears the stricter second one.

Worked Example 3: Cards, independence, and mutual exclusivity together. Given: one card drawn from a standard 52-card deck. Let A = “King,” B = “Heart,” C = “Queen.” Step 1: P(A) = 4/52, P(B) = 13/52, P(A and B) = P(King of Hearts) = 1/52. Step 2: Test independence of A and B: P(A) × P(B) = (4/52) × (13/52) = 1/52, which equals P(A and B) exactly — A and B are independent (a card’s rank and suit don’t affect each other in a fair deck). Step 3: Now compare A and C: P(A and C) = 0 (a card can’t be a King and a Queen at once, so they’re mutually exclusive), but P(A) × P(C) = (4/52) × (4/52) = 1/169 ≈ 0.0059, which is not 0. Answer: King-and-Heart are independent; King-and-Queen are mutually exclusive but, exactly as the theory above predicts, not independent.

History

  • In 1654, a gambler known as the Chevalier de Méré brought a practical question — how to fairly divide the stakes of an interrupted game of chance, the “problem of points” — to Blaise Pascal, who worked it out through a series of letters with Pierre de Fermat; that correspondence is usually credited as probability theory’s formal starting point.
  • Christiaan Huygens turned the Pascal-Fermat letters into the first published probability treatise, De Ratiociniis in Ludo Aleae (“On Reasoning in Games of Chance”), in 1657.
  • Jacob Bernoulli’s Ars Conjectandi, published posthumously in 1713, introduced the law of large numbers, connecting a theoretical probability to what actually happens on average over many repeated trials.
  • Pierre-Simon Laplace’s Théorie analytique des probabilités (1812) formalized the classical “favorable outcomes over total outcomes” definition and extended probability well beyond gambling, into astronomy, error measurement, and demographics.
  • Andrey Kolmogorov finally gave probability a rigorous axiomatic foundation in 1933, built on measure theory, placing the whole subject on the same solid logical footing as geometry or calculus — over 250 years after Pascal and Fermat’s informal beginnings.
  • That axiomatic version, not the gambling-table version, is what modern statistics, finance, and machine learning are built on today.

Why It Matters

  • Insurance and actuarial science are built entirely on probability: insurers estimate the likelihood and expected cost of claims across large pools of people to price every policy they sell.
  • Interpreting a medical test result correctly requires conditional probability, not just the test’s advertised accuracy — when a condition is rare, even a highly accurate test can flag more false positives than true cases, a direct consequence of how P(disease|positive test) and P(positive test|disease) differ.
  • Weather forecasts (“70% chance of rain”) are direct probability statements, produced by models trained on historical atmospheric patterns.
  • Genetics is a textbook probability application: Mendelian inheritance and Punnett squares predict the probability that offspring inherit a particular combination of alleles from their parents.
  • Modern machine learning and statistics are probability from the ground up — classifiers output probabilities rather than certainties, and most learning algorithms are, under the hood, probability models fit to data.
  • Manufacturing quality control uses probability to set acceptable defect rates and to design sampling inspections that catch problems without testing every single unit off the line.

Common Pitfalls

  • The gambler’s fallacy: believing that past independent outcomes affect future ones, like assuming a coin is “due” for tails after five straight heads. The coin has no memory; P(tails) is still 0.5 no matter what came before.
  • Confusing P(A|B) with P(B|A): these can be very different numbers. If A = “King” and B = “face card” in a standard deck, P(A|B) = 1/3 (only a third of face cards are Kings) but P(B|A) = 1 (every King is automatically a face card) — swapping the two is a common and consequential error.
  • Assuming mutually exclusive means independent, or the reverse: as shown above, the truth runs the other way — two events that both have a real chance of happening can be mutually exclusive or independent, but essentially never both.
  • Forgetting to subtract the overlap in the addition rule: adding P(A) + P(B) when the events can happen together double-counts every outcome in P(A and B), inflating the result — sometimes past 1, which is an immediate sign of a dropped overlap term.
  • Treating a conditional probability as if it applies unconditionally: P(A|B) only describes the world once B is already known to be true; using it in place of the plain P(A) silently smuggles in information you may not actually have.
  • Reading “independent” as “unrelated in real life”: independence is a precise numerical statement, P(A and B) = P(A) × P(B); two obviously connected real-world events can happen to satisfy that equation, and two seemingly unrelated events can fail it.

Comparison

PropertyIndependent EventsMutually Exclusive Events
DefinitionOne event’s occurrence carries no information about the otherThe events cannot occur at the same time
Can they share outcomes?Yes — typically some overlap, unless P(A) or P(B) is 0No — by definition they share zero outcomes
Defining formulaP(A and B) = P(A) × P(B)P(A and B) = 0
Effect of knowing B happenedLeaves P(A) unchanged: P(A|B) = P(A)Rules A out completely: P(A|B) = 0
Can an event pair be both?Only if P(A) = 0 or P(B) = 0Only if P(A) = 0 or P(B) = 0

FAQ

What’s the real difference between P(A|B) and P(B|A)? They condition on different information, and are generally not equal. Let A = “the card is a King” and B = “the card is a face card” (Jack, Queen, or King) in a standard deck. P(A|B) = P(King and face) / P(face) = (4/52) / (12/52) = 1/3, since only a third of face cards are Kings. But P(B|A) = P(face and King) / P(King) = (4/52) / (4/52) = 1, since every King is automatically a face card. Same two events, two very different conditional probabilities.

Can two events with a real chance of happening ever be both mutually exclusive and independent? No — only in the degenerate case where one of them already has probability 0. Mutually exclusive forces P(A and B) = 0; independent forces P(A and B) = P(A) × P(B). Together those force P(A) × P(B) = 0, which means P(A) = 0 or P(B) = 0. See the full argument above the widget.

If P(A and B) = 0, doesn’t that mean A and B don’t affect each other, i.e. they’re independent? No — it usually means the opposite. Two events that can never occur together are about as dependent as two events can be: the instant you learn one happened, you know with certainty the other did not. “P(A and B) = 0” and independence only coincide when one of the events was already impossible.

Example

A weather app quoting a 30% chance of rain (A) and a 20% chance of high wind (B), with the two historically occurring together 8% of the time, lets a user combine both forecasts into one number with a single addition-rule calculation: P(A or B) = 0.30 + 0.20 - 0.08 = 0.42 — a 42% chance of at least one adverse condition today, computed the same way as the class or card examples above.

Dig deeper