Probability Theory
Probability Theory
Definition: Probability is a number from 0 to 1 that measures how likely an event is to occur, where 0 means impossible and 1 means certain, assigned to outcomes drawn from a well-defined sample space of everything that could happen.
How It Works
- The sample space (S) is the set of every possible outcome of a random process — {1, 2, 3, 4, 5, 6} for one die roll — and an event is any subset of that sample space, such as “rolling an even number” = {2, 4, 6}.
- A probability is always a number from 0 to 1: 0 means an event never happens, 1 means it always happens, and everything else falls somewhere between rare and near-certain.
- For equally likely outcomes, P(event) = (outcomes in the event) / (total outcomes in the sample space) — e.g. P(even) = 3/6 = 0.5 for a fair die.
- The addition rule gives the probability that A or B (or both) occurs: P(A or B) = P(A) + P(B) - P(A and B). Subtracting P(A and B) corrects for counting any shared outcome twice.
- If A and B can never happen together, they are mutually exclusive, P(A and B) = 0, and the addition rule simplifies to P(A or B) = P(A) + P(B).
- Two events are independent when one occurring has no effect on the probability of the other; for independent events specifically, P(A and B) = P(A) × P(B), the multiplication rule.
- Conditional probability, written P(A|B) and read “the probability of A given B,” is A’s probability once B is already known to have happened: P(A|B) = P(A and B) / P(B).
- Independence has an equivalent conditional definition: A and B are independent exactly when P(A|B) = P(A), meaning learning that B happened doesn’t move the needle on A at all.
- The complement rule states P(not A) = 1 - P(A). An event and its complement together account for the entire sample space and never overlap, so their probabilities must sum to exactly 1.
- Mutually exclusive and independent describe two different, almost opposite relationships between events; outside one degenerate case, a pair of events cannot be both at once (developed fully below).
- Every outcome’s probability in a sample space must sum to exactly 1 — a fair six-sided die assigns 1/6 to each of six faces, and six copies of 1/6 add up to 1.
Illustration
function fmt(n) { return (Math.round(n * 100) / 100).toFixed(2).replace(/.00$/, ‘.0’); }
function update() { var pA = parseFloat(aIn.value); var pB = parseFloat(bIn.value); var rawCap = parseFloat(capIn.value);
// The overlap can never exceed the smaller of P(A) and P(B) -- a ceiling
// -- and can never drop below P(A)+P(B)-1 either -- a floor. Below that
// floor, P(A or B) would have to exceed 1, which no real probability
// can do. Clamp the raw slider value into that valid window instead of
// fighting the <input>'s own min/max attributes.
var maxCap = Math.min(pA, pB);
var minCap = Math.max(0, pA + pB - 1);
var pCap = Math.min(Math.max(rawCap, minCap), maxCap);
aOut.textContent = fmt(pA);
bOut.textContent = fmt(pB);
capOut.textContent = fmt(pCap);
var aOnly = pA - pCap;
var bOnly = pB - pCap;
var union = pA + pB - pCap;
var cond = pCap / pB;
aOnlyVal.textContent = fmt(aOnly);
bOnlyVal.textContent = fmt(bOnly);
capVal.textContent = fmt(pCap);
paTxt.textContent = 'P(A) = ' + fmt(pA);
pbTxt.textContent = 'P(B) = ' + fmt(pB);
pcapTxt.textContent = 'P(A∩B) = ' + fmt(pCap);
unionTxt.textContent = 'P(A∪B) = ' + fmt(union);
condTxt.textContent = 'P(A|B) = ' + fmt(cond);
var independent = Math.abs(pCap - pA * pB) <= 0.02;
var mutex = pCap <= 0.01;
var label = independent ? 'Independent' : 'Not independent';
if (mutex) label += ' · Mutually exclusive';
classifyTxt.textContent = label;
}
[aIn, bIn, capIn].forEach(function (el) { el.addEventListener(‘input’, update); }); resetBtn.addEventListener(‘click’, function () { aIn.value = 0.4; bIn.value = 0.3; capIn.value = 0.12; update(); });
update(); })();
Mutually exclusive and independent sound like they should be close cousins — both seem to describe two events that have nothing to do with each other — but they capture opposite relationships, and mixing them up is one of the most common mistakes in introductory probability.
Two events are mutually exclusive when they share no outcomes: if A happens, B is guaranteed not to, and vice versa. Knowing A occurred tells you everything about B — that it certainly did not happen. That is about as far from “no effect on each other” as two events can get. Independence requires the opposite: knowing A occurred must tell you nothing about B, formally P(A|B) = P(A).
The two conditions actually collide algebraically. Mutually exclusive means P(A and B) = 0. Independent means P(A and B) = P(A) × P(B). If an event pair satisfies both at once, then P(A) × P(B) = 0. A product of two real numbers is zero only when at least one factor is zero, so this forces P(A) = 0 or P(B) = 0 — one of the two events has to be impossible from the start. Outside that degenerate edge case, any two events that both have a real chance of happening can be mutually exclusive, or they can be independent, but never both. Try it on the sliders above: push P(A∩B) down toward zero (mutually exclusive) while P(A) and P(B) both sit at healthy values like 0.4 and 0.3, and the “Independent” label always switches off, because 0.4 × 0.3 = 0.12, not 0.
Under the Hood
Precisely stated, the three central rules of two-event probability:
P(A or B) = P(A) + P(B) - P(A and B) addition rule (always true)
P(A or B) = P(A) + P(B) special case: A, B mutually exclusive, so P(A and B) = 0
P(A|B) = P(A and B) / P(B) conditional probability (requires P(B) > 0)
P(B|A) = P(A and B) / P(A) the reverse conditional (requires P(A) > 0) -- not generally equal to P(A|B)
P(A and B) = P(A) × P(B) multiplication rule -- valid ONLY if A and B are independent
Independence test: P(A and B) = P(A) × P(B) <=> P(A|B) = P(A) <=> P(B|A) = P(B)
Worked Example 1: Addition rule with real overlap. Given: in a class of 40 students, 18 play soccer (A), 14 play basketball (B), and 6 play both. Step 1: P(A) = 18/40 = 0.45, P(B) = 14/40 = 0.35, P(A and B) = 6/40 = 0.15. Step 2: P(A or B) = P(A) + P(B) - P(A and B) = 0.45 + 0.35 - 0.15. Answer: P(A or B) = 0.65 — 65% of the class (26 of 40 students) plays at least one of the two sports.
Worked Example 2: Conditional probability from an inspection log. Given: of 200 manufactured parts, 150 pass a first inspection (B); of those 150, 120 also pass a stricter secondary test (A and B together, since the secondary test only runs on parts that already passed). Step 1: P(B) = 150/200 = 0.75. P(A and B) = 120/200 = 0.60. Step 2: P(A|B) = P(A and B) / P(B) = 0.60 / 0.75. Answer: P(A|B) = 0.80 — given that a part passed the first inspection, there is an 80% chance it also clears the stricter second one.
Worked Example 3: Cards, independence, and mutual exclusivity together. Given: one card drawn from a standard 52-card deck. Let A = “King,” B = “Heart,” C = “Queen.” Step 1: P(A) = 4/52, P(B) = 13/52, P(A and B) = P(King of Hearts) = 1/52. Step 2: Test independence of A and B: P(A) × P(B) = (4/52) × (13/52) = 1/52, which equals P(A and B) exactly — A and B are independent (a card’s rank and suit don’t affect each other in a fair deck). Step 3: Now compare A and C: P(A and C) = 0 (a card can’t be a King and a Queen at once, so they’re mutually exclusive), but P(A) × P(C) = (4/52) × (4/52) = 1/169 ≈ 0.0059, which is not 0. Answer: King-and-Heart are independent; King-and-Queen are mutually exclusive but, exactly as the theory above predicts, not independent.
History
- In 1654, a gambler known as the Chevalier de Méré brought a practical question — how to fairly divide the stakes of an interrupted game of chance, the “problem of points” — to Blaise Pascal, who worked it out through a series of letters with Pierre de Fermat; that correspondence is usually credited as probability theory’s formal starting point.
- Christiaan Huygens turned the Pascal-Fermat letters into the first published probability treatise, De Ratiociniis in Ludo Aleae (“On Reasoning in Games of Chance”), in 1657.
- Jacob Bernoulli’s Ars Conjectandi, published posthumously in 1713, introduced the law of large numbers, connecting a theoretical probability to what actually happens on average over many repeated trials.
- Pierre-Simon Laplace’s Théorie analytique des probabilités (1812) formalized the classical “favorable outcomes over total outcomes” definition and extended probability well beyond gambling, into astronomy, error measurement, and demographics.
- Andrey Kolmogorov finally gave probability a rigorous axiomatic foundation in 1933, built on measure theory, placing the whole subject on the same solid logical footing as geometry or calculus — over 250 years after Pascal and Fermat’s informal beginnings.
- That axiomatic version, not the gambling-table version, is what modern statistics, finance, and machine learning are built on today.
Why It Matters
- Insurance and actuarial science are built entirely on probability: insurers estimate the likelihood and expected cost of claims across large pools of people to price every policy they sell.
- Interpreting a medical test result correctly requires conditional probability, not just the test’s advertised accuracy — when a condition is rare, even a highly accurate test can flag more false positives than true cases, a direct consequence of how P(disease|positive test) and P(positive test|disease) differ.
- Weather forecasts (“70% chance of rain”) are direct probability statements, produced by models trained on historical atmospheric patterns.
- Genetics is a textbook probability application: Mendelian inheritance and Punnett squares predict the probability that offspring inherit a particular combination of alleles from their parents.
- Modern machine learning and statistics are probability from the ground up — classifiers output probabilities rather than certainties, and most learning algorithms are, under the hood, probability models fit to data.
- Manufacturing quality control uses probability to set acceptable defect rates and to design sampling inspections that catch problems without testing every single unit off the line.
Common Pitfalls
- The gambler’s fallacy: believing that past independent outcomes affect future ones, like assuming a coin is “due” for tails after five straight heads. The coin has no memory; P(tails) is still 0.5 no matter what came before.
- Confusing P(A|B) with P(B|A): these can be very different numbers. If A = “King” and B = “face card” in a standard deck, P(A|B) = 1/3 (only a third of face cards are Kings) but P(B|A) = 1 (every King is automatically a face card) — swapping the two is a common and consequential error.
- Assuming mutually exclusive means independent, or the reverse: as shown above, the truth runs the other way — two events that both have a real chance of happening can be mutually exclusive or independent, but essentially never both.
- Forgetting to subtract the overlap in the addition rule: adding P(A) + P(B) when the events can happen together double-counts every outcome in P(A and B), inflating the result — sometimes past 1, which is an immediate sign of a dropped overlap term.
- Treating a conditional probability as if it applies unconditionally: P(A|B) only describes the world once B is already known to be true; using it in place of the plain P(A) silently smuggles in information you may not actually have.
- Reading “independent” as “unrelated in real life”: independence is a precise numerical statement, P(A and B) = P(A) × P(B); two obviously connected real-world events can happen to satisfy that equation, and two seemingly unrelated events can fail it.
Comparison
| Property | Independent Events | Mutually Exclusive Events |
|---|---|---|
| Definition | One event’s occurrence carries no information about the other | The events cannot occur at the same time |
| Can they share outcomes? | Yes — typically some overlap, unless P(A) or P(B) is 0 | No — by definition they share zero outcomes |
| Defining formula | P(A and B) = P(A) × P(B) | P(A and B) = 0 |
| Effect of knowing B happened | Leaves P(A) unchanged: P(A|B) = P(A) | Rules A out completely: P(A|B) = 0 |
| Can an event pair be both? | Only if P(A) = 0 or P(B) = 0 | Only if P(A) = 0 or P(B) = 0 |
FAQ
What’s the real difference between P(A|B) and P(B|A)? They condition on different information, and are generally not equal. Let A = “the card is a King” and B = “the card is a face card” (Jack, Queen, or King) in a standard deck. P(A|B) = P(King and face) / P(face) = (4/52) / (12/52) = 1/3, since only a third of face cards are Kings. But P(B|A) = P(face and King) / P(King) = (4/52) / (4/52) = 1, since every King is automatically a face card. Same two events, two very different conditional probabilities.
Can two events with a real chance of happening ever be both mutually exclusive and independent? No — only in the degenerate case where one of them already has probability 0. Mutually exclusive forces P(A and B) = 0; independent forces P(A and B) = P(A) × P(B). Together those force P(A) × P(B) = 0, which means P(A) = 0 or P(B) = 0. See the full argument above the widget.
If P(A and B) = 0, doesn’t that mean A and B don’t affect each other, i.e. they’re independent? No — it usually means the opposite. Two events that can never occur together are about as dependent as two events can be: the instant you learn one happened, you know with certainty the other did not. “P(A and B) = 0” and independence only coincide when one of the events was already impossible.
Example
A weather app quoting a 30% chance of rain (A) and a 20% chance of high wind (B), with the two historically occurring together 8% of the time, lets a user combine both forecasts into one number with a single addition-rule calculation: P(A or B) = 0.30 + 0.20 - 0.08 = 0.42 — a 42% chance of at least one adverse condition today, computed the same way as the class or card examples above.