A-B Testing
A/B Testing
Definition: An experiment that randomly shows two (or more) versions of a product to different user groups, then compares a target metric between them to decide which version actually performs better.
How It Works
- Users are randomly split into groups: one sees the current version (control), another sees the change (variant)
- Randomization is the key mechanic. Without it, any difference between groups could come from who they are, not what they saw
- A single primary metric is defined before the test starts (conversion rate, click-through rate, revenue per visitor)
- The test runs until it hits a pre-registered sample size or a pre-set time window, not until the result “looks good”
- Statistical significance testing checks whether the observed gap between control and variant is bigger than random chance alone would produce
- Guardrail metrics (page load time, support tickets, unsubscribe rate) are watched alongside the primary metric so a “win” on one number isn’t a loss somewhere else
- Tests typically run for at least one full week (often two), so day-of-week effects like weekend shopping behavior don’t skew the result
- Assignment is sticky: the same user sees the same variant on every visit, usually via a hashed user ID or cookie, so results aren’t contaminated by people bouncing between versions
- Most tests use a two-tailed significance test, since the variant could turn out worse than control, not just better
- Sequential or “always valid” testing methods exist specifically to let teams check results early without inflating the false-positive rate, unlike a naive daily peek at a fixed-horizon test
Under the Hood
Each box above maps to a real decision point: skipping the sample-size step is how teams end up staring at inconclusive dashboards for months, and skipping the pre-registration of what counts as “significant” is how peeking creeps in.
Worked example 1: how big a sample do you need?
Given:
- Baseline conversion rate p = 4% (0.04)
- Minimum effect worth detecting: 0.6 percentage points (0.046 vs 0.04)
- Target: 80% statistical power, 95% confidence (alpha = 0.05)
Step, using the common rule-of-thumb approximation n = 16 x p(1-p) / d^2 per variant:
n = 16 x 0.04 x 0.96 / 0.006^2
n = 16 x 0.0384 / 0.000036
n = 0.6144 / 0.000036
n = 17,067
Answer: roughly 17,000 users per variant (34,000 total) before the test even starts, which drives how long the team plans to run it.
Worked example 2: is the result real?
Given, after running the test with 20,000 users per variant:
- Control: 800 conversions out of 20,000 = 4.0%
- Variant: 920 conversions out of 20,000 = 4.6%
Step 1, pooled conversion rate:
p_pool = (800 + 920) / (20,000 + 20,000) = 1,720 / 40,000 = 0.043
Step 2, standard error:
SE = sqrt(p_pool x (1 - p_pool) x (1/n1 + 1/n2))
SE = sqrt(0.043 x 0.957 x 0.0001)
SE = sqrt(0.0000041151) = 0.00203
Step 3, z-score:
z = (0.046 - 0.04) / 0.00203 = 2.96
Answer: |z| = 2.96 exceeds the 1.96 threshold for 95% confidence (p is about 0.003, two-tailed), so the lift is statistically significant. The variant ships.
Reading the threshold:
- 1.96 marks the outer 2.5% of each tail of a normal distribution, together the 5% (alpha = 0.05) considered too unlikely to be chance in a two-tailed test
- A z-score below 1.96 doesn’t prove the variant is worthless, it just means the test hasn’t produced enough evidence yet
- Statistical significance is not the same as practical significance: a 0.1-point lift can be statistically real at huge sample sizes but too small to justify the engineering cost of shipping it
Why It Matters
- Replaces internal opinion and debate with evidence of what real users actually do, and it regularly contradicts what the team expected
- Turns product decisions into a repeatable process instead of a one-off argument won by whoever is most senior in the room
- Limits the blast radius of a bad idea: a losing variant only ran on part of traffic for a few weeks, not on every user forever
- Surfaces effect sizes, not just direction, so a team can weigh a confirmed 0.6-point lift against the engineering cost of shipping it
- Builds an institutional record of what’s already been tried, so the same “let’s try a bigger button” idea doesn’t get re-litigated every year
Common Pitfalls
- Peeking at results daily and stopping the instant it “looks like a winner,” before reaching the pre-registered sample size
- Running too many variants at once, which shrinks the traffic (and statistical power) each one gets and inflates the odds something wins by chance
- Ignoring the novelty effect: a redesign can spike engagement for a week simply because it’s different, then fade back to baseline
- Testing a change too small to matter, then reading a null result as “the idea failed” when the test was underpowered to detect it
- Chasing a secondary metric that happened to move while the primary metric didn’t
- Running the test on too small a slice of the audience (a single country or device type) and then rolling the result out globally as if it generalized
- Sample ratio mismatch: the actual traffic split ends up far from the intended 50/50 (say 48/52) because of a bug in the assignment logic, which quietly invalidates the whole test
- Comparing results across tests run in different seasons or campaigns as if conditions were held constant, when the underlying traffic mix had already shifted
Comparison
| A/B Testing | Multivariate Testing | Bandit Algorithms | |
|---|---|---|---|
| What it compares | Two (or a few) full versions | Multiple elements combined (headline x image x button) | Multiple variants, adjusted live |
| Traffic split | Fixed, usually even | Fixed, split across combinations | Shifts toward the current best performer over time |
| Sample size needed | Moderate | High, scales with combinations | Lower per arm, but needs many rounds |
| Best for | Clear either/or decisions | Isolating which element drives an effect | Ongoing optimization where losing traffic is costly |
| Statistical rigor | High, easy to explain | High, but interaction effects get complex | Weaker for one-time definitive claims |
| Cost of a losing variant | Bounded and known upfront | Higher, spread across many combinations | Self-limiting, the algorithm shifts traffic away automatically |
Pick multivariate testing when several elements might interact (a headline’s wording can change which image performs best). Pick a bandit when the test itself is costly to run on a loser, like pricing or a homepage that gets one-time visitors who won’t come back for a second, cleaner comparison.
Example
The Obama 2008 presidential campaign ran systematic A/B tests on its donation and signup page. The campaign’s analytics team, led by Dan Siroker, found that pairing a family photo with a “Learn More” button outperformed the campaign’s original media and button-text combination, an improvement widely reported to have lifted conversion by roughly 40%, translating into millions of additional email signups and donations over the campaign.
Before that, Siroker worked on Google’s search results page, which became widely known for testing 41 different shades of blue for its links to see which one users clicked more, a story reported by the New York Times in 2009 as an example of how far data-driven testing extended into even small visual decisions at scale.
Related Terms
Referenced by