A-B Testing

A/B Testing

Definition: An experiment that randomly shows two (or more) versions of a product to different user groups, then compares a target metric between them to decide which version actually performs better.

How It Works

  • Users are randomly split into groups: one sees the current version (control), another sees the change (variant)
  • Randomization is the key mechanic. Without it, any difference between groups could come from who they are, not what they saw
  • A single primary metric is defined before the test starts (conversion rate, click-through rate, revenue per visitor)
  • The test runs until it hits a pre-registered sample size or a pre-set time window, not until the result “looks good”
  • Statistical significance testing checks whether the observed gap between control and variant is bigger than random chance alone would produce
  • Guardrail metrics (page load time, support tickets, unsubscribe rate) are watched alongside the primary metric so a “win” on one number isn’t a loss somewhere else
  • Tests typically run for at least one full week (often two), so day-of-week effects like weekend shopping behavior don’t skew the result
  • Assignment is sticky: the same user sees the same variant on every visit, usually via a hashed user ID or cookie, so results aren’t contaminated by people bouncing between versions
  • Most tests use a two-tailed significance test, since the variant could turn out worse than control, not just better
  • Sequential or “always valid” testing methods exist specifically to let teams check results early without inflating the false-positive rate, unlike a naive daily peek at a fixed-horizon test

Under the Hood

Each box above maps to a real decision point: skipping the sample-size step is how teams end up staring at inconclusive dashboards for months, and skipping the pre-registration of what counts as “significant” is how peeking creeps in.

Worked example 1: how big a sample do you need?

Given:

  • Baseline conversion rate p = 4% (0.04)
  • Minimum effect worth detecting: 0.6 percentage points (0.046 vs 0.04)
  • Target: 80% statistical power, 95% confidence (alpha = 0.05)

Step, using the common rule-of-thumb approximation n = 16 x p(1-p) / d^2 per variant:

n = 16 x 0.04 x 0.96 / 0.006^2
n = 16 x 0.0384 / 0.000036
n = 0.6144 / 0.000036
n = 17,067

Answer: roughly 17,000 users per variant (34,000 total) before the test even starts, which drives how long the team plans to run it.

Worked example 2: is the result real?

Given, after running the test with 20,000 users per variant:

  • Control: 800 conversions out of 20,000 = 4.0%
  • Variant: 920 conversions out of 20,000 = 4.6%

Step 1, pooled conversion rate:

p_pool = (800 + 920) / (20,000 + 20,000) = 1,720 / 40,000 = 0.043

Step 2, standard error:

SE = sqrt(p_pool x (1 - p_pool) x (1/n1 + 1/n2))
SE = sqrt(0.043 x 0.957 x 0.0001)
SE = sqrt(0.0000041151) = 0.00203

Step 3, z-score:

z = (0.046 - 0.04) / 0.00203 = 2.96

Answer: |z| = 2.96 exceeds the 1.96 threshold for 95% confidence (p is about 0.003, two-tailed), so the lift is statistically significant. The variant ships.

Reading the threshold:

  • 1.96 marks the outer 2.5% of each tail of a normal distribution, together the 5% (alpha = 0.05) considered too unlikely to be chance in a two-tailed test
  • A z-score below 1.96 doesn’t prove the variant is worthless, it just means the test hasn’t produced enough evidence yet
  • Statistical significance is not the same as practical significance: a 0.1-point lift can be statistically real at huge sample sizes but too small to justify the engineering cost of shipping it

Why It Matters

  • Replaces internal opinion and debate with evidence of what real users actually do, and it regularly contradicts what the team expected
  • Turns product decisions into a repeatable process instead of a one-off argument won by whoever is most senior in the room
  • Limits the blast radius of a bad idea: a losing variant only ran on part of traffic for a few weeks, not on every user forever
  • Surfaces effect sizes, not just direction, so a team can weigh a confirmed 0.6-point lift against the engineering cost of shipping it
  • Builds an institutional record of what’s already been tried, so the same “let’s try a bigger button” idea doesn’t get re-litigated every year

Common Pitfalls

  • Peeking at results daily and stopping the instant it “looks like a winner,” before reaching the pre-registered sample size
  • Running too many variants at once, which shrinks the traffic (and statistical power) each one gets and inflates the odds something wins by chance
  • Ignoring the novelty effect: a redesign can spike engagement for a week simply because it’s different, then fade back to baseline
  • Testing a change too small to matter, then reading a null result as “the idea failed” when the test was underpowered to detect it
  • Chasing a secondary metric that happened to move while the primary metric didn’t
  • Running the test on too small a slice of the audience (a single country or device type) and then rolling the result out globally as if it generalized
  • Sample ratio mismatch: the actual traffic split ends up far from the intended 50/50 (say 48/52) because of a bug in the assignment logic, which quietly invalidates the whole test
  • Comparing results across tests run in different seasons or campaigns as if conditions were held constant, when the underlying traffic mix had already shifted

Comparison

A/B TestingMultivariate TestingBandit Algorithms
What it comparesTwo (or a few) full versionsMultiple elements combined (headline x image x button)Multiple variants, adjusted live
Traffic splitFixed, usually evenFixed, split across combinationsShifts toward the current best performer over time
Sample size neededModerateHigh, scales with combinationsLower per arm, but needs many rounds
Best forClear either/or decisionsIsolating which element drives an effectOngoing optimization where losing traffic is costly
Statistical rigorHigh, easy to explainHigh, but interaction effects get complexWeaker for one-time definitive claims
Cost of a losing variantBounded and known upfrontHigher, spread across many combinationsSelf-limiting, the algorithm shifts traffic away automatically

Pick multivariate testing when several elements might interact (a headline’s wording can change which image performs best). Pick a bandit when the test itself is costly to run on a loser, like pricing or a homepage that gets one-time visitors who won’t come back for a second, cleaner comparison.

Example

The Obama 2008 presidential campaign ran systematic A/B tests on its donation and signup page. The campaign’s analytics team, led by Dan Siroker, found that pairing a family photo with a “Learn More” button outperformed the campaign’s original media and button-text combination, an improvement widely reported to have lifted conversion by roughly 40%, translating into millions of additional email signups and donations over the campaign.

Before that, Siroker worked on Google’s search results page, which became widely known for testing 41 different shades of blue for its links to see which one users clicked more, a story reported by the New York Times in 2009 as an example of how far data-driven testing extended into even small visual decisions at scale.

Dig deeper