Descriptive Statistics

Descriptive Statistics

Definition: Descriptive statistics summarize a dataset with a small number of representative values: measures of central tendency (mean, median, mode) that describe a “typical” value, and measures of spread (range, variance, standard deviation) that describe how much the data varies around that typical value.

How It Works

  • The mean (arithmetic average) adds up every value in a dataset and divides by how many values there are. It uses the exact magnitude of every single data point.
  • The median is the middle value once the data is sorted: for an odd count of values it’s the single middle value, for an even count it’s the average of the two middle values.
  • The mode is the most frequently occurring value. A dataset can have one mode, several tied modes (multimodal), or no mode at all if every value appears exactly once.
  • The mean is sensitive to every value, including extreme ones, so a single very large or very small outlier can pull it substantially away from where most of the data actually sits.
  • The median only cares about rank, not magnitude, so it barely moves when an extreme value is added or changed, which makes it the more robust choice for skewed data.
  • The range is the simplest measure of spread: the maximum value minus the minimum. It’s easy to compute but depends entirely on the two most extreme values.
  • The variance measures spread by averaging the squared distance of every value from the mean. Squaring makes every deviation positive, so distances above and below the mean don’t cancel out.
  • The standard deviation is the square root of the variance, which brings the units back to the original scale (dollars instead of dollars-squared, for example), making it far more interpretable than variance alone.
  • A population statistic describes an entire group of interest; a sample statistic estimates that same quantity from only part of the group. Sample variance divides by n−1 instead of n (Bessel’s correction), because a sample’s own mean is always at least as close to that sample’s data as the true population mean would be, which makes the plain n-divisor an underestimate of the true spread.
  • Quartiles split sorted data into four equal parts: Q1 (25th percentile), the median (50th percentile), and Q3 (75th percentile).
  • The five-number summary (minimum, Q1, median, Q3, maximum) condenses a whole dataset into five values, and is exactly what a box plot draws: a box from Q1 to Q3 with a line at the median, and whiskers reaching out to the min and max.

Illustration

Sorted: 4, 7, 9, 12, 15, 18 Mean: 10.83 Median: 10.50 Mode: No mode (all distinct) Range: 14.0 Variance (population): 22.47 Std Dev (population): 4.74 mean median 0 5 10 15 20
Drag the six sliders to move each data point. The dot plot, the mean (red dashed line), and the median (gold dashed line) redraw live, along with the full summary statistics above the axis. Identical values stack vertically so repeats are still visible.

var meanLine = document.getElementById(‘stats-mean-line’); var medianLine = document.getElementById(‘stats-median-line’); var meanLabel = document.getElementById(‘stats-mean-label’); var medianLabel = document.getElementById(‘stats-median-label’);

var resetBtn = document.getElementById(‘stats-reset’);

var ox = 40, scale = 17.5, dotBase = 196; // pixels per data unit; axis runs 0-20

function toX(v) { return ox + v * scale; } function fmt(n) { return (Math.round(n * 100) / 100).toFixed(2).replace(/.00$/, ‘.0’); }

function modeText(sorted) { var freq = {}; sorted.forEach(function (v) { freq[v] = (freq[v] || 0) + 1; }); var maxFreq = 0; Object.keys(freq).forEach(function (k) { if (freq[k] > maxFreq) maxFreq = freq[k]; }); if (maxFreq === 1) return ‘No mode (all distinct)’; var modes = Object.keys(freq).filter(function (k) { return freq[k] === maxFreq; }).map(Number).sort(function (a, b) { return a - b; }); if (modes.length === 1) return modes[0] + ’ (appears ’ + maxFreq + ‘x)’; return modes.join(’, ’) + ’ (each appears ’ + maxFreq + ‘x)’; }

function update() { var vals = sliders.map(function (el) { return parseInt(el.value, 10); }); vals.forEach(function (v, i) { outs[i].textContent = v; });

var sorted = vals.slice().sort(function (a, b) { return a - b; });
var n = vals.length;
var sum = vals.reduce(function (a, b) { return a + b; }, 0);
var mean = sum / n;
var median = (sorted[n / 2 - 1] + sorted[n / 2]) / 2;
var range = sorted[n - 1] - sorted[0];
var sumSqDev = vals.reduce(function (acc, v) { return acc + (v - mean) * (v - mean); }, 0);
var variance = sumSqDev / n;
var stdDev = Math.sqrt(variance);

sortedTxt.textContent = 'Sorted: ' + sorted.join(', ');
meanTxt.textContent = 'Mean: ' + fmt(mean);
medianTxt.textContent = 'Median: ' + fmt(median);
modeTxt.textContent = 'Mode: ' + modeText(sorted);
rangeTxt.textContent = 'Range: ' + fmt(range);
varTxt.textContent = 'Variance (population): ' + fmt(variance);
stdTxt.textContent = 'Std Dev (population): ' + fmt(stdDev);

var stackCount = {};
vals.forEach(function (v, i) {
  var k = stackCount[v] || 0;
  stackCount[v] = k + 1;
  dots[i].setAttribute('cx', toX(v).toFixed(1));
  dots[i].setAttribute('cy', (dotBase - k * 14).toFixed(1));
});

var meanX = toX(mean).toFixed(1);
var medianX = toX(median).toFixed(1);
meanLine.setAttribute('x1', meanX); meanLine.setAttribute('x2', meanX);
medianLine.setAttribute('x1', medianX); medianLine.setAttribute('x2', medianX);
meanLabel.setAttribute('x', meanX);
medianLabel.setAttribute('x', medianX);

}

sliders.forEach(function (el) { el.addEventListener(‘input’, update); }); resetBtn.addEventListener(‘click’, function () { sliders.forEach(function (el, i) { el.value = defaults[i]; }); update(); });

update(); })();

Under the Hood

Population (the entire group you care about), N values:
  μ  = (Σ x_i) / N                    population mean
  σ² = (Σ (x_i − μ)²) / N             population variance
  σ  = √σ²                            population standard deviation

Sample (a subset used to estimate the population), n values:
  mean = (Σ x_i) / n                  identical formula to the population mean
  s²   = (Σ (x_i − mean)²) / (n − 1)  sample variance — divide by n − 1, not n (Bessel's correction)
  s    = √s²                          sample standard deviation
  • The only formula that actually changes between population and sample is the divisor on variance: N for a population, n−1 for a sample. That single change is Bessel’s correction, and it exists because a sample’s own mean is pulled from, and therefore sits unusually close to, that sample’s own data — which makes dividing by n alone an underestimate of the true population variance.

Worked Example 1: Mean, median, and mode by hand. Given: the dataset 2, 4, 4, 7, 9 (n = 5). Step 1: mean = (2 + 4 + 4 + 7 + 9) / 5 = 26 / 5. Step 2: the data is already sorted; with an odd count of 5, the median is simply the single middle value, the 3rd one. Step 3: 4 is the only value that repeats (twice), so it is the mode. Answer: mean = 5.2, median = 4, mode = 4.

Worked Example 2: Variance and standard deviation step by step. Given: the dataset 2, 4, 6, 8 (n = 4), treated as a population. Step 1: mean = (2 + 4 + 6 + 8) / 4 = 20 / 4 = 5. Step 2: deviations from the mean: −3, −1, 1, 3. Squared: 9, 1, 1, 9. Step 3: sum of squared deviations = 9 + 1 + 1 + 9 = 20; population variance = 20 / 4 = 5. Answer: variance = 5, standard deviation = √5 ≈ 2.24. (If this same data were a sample instead of the full population, sample variance would divide by n−1 = 3 instead: 20 / 3 ≈ 6.67, giving a sample standard deviation of about 2.58 — larger, because dividing by a smaller number always increases the result.)

Worked Example 3: One outlier, mean vs. median. Given: the dataset 4, 5, 6, 7, 8 (n = 5), then the highest value is replaced with an outlier: 4, 5, 6, 7, 50. Step 1: original mean = (4 + 5 + 6 + 7 + 8) / 5 = 30 / 5 = 6; original median (sorted, middle of 5) = 6. Step 2: new mean = (4 + 5 + 6 + 7 + 50) / 5 = 72 / 5 = 14.4. Step 3: new median: sorted is 4, 5, 6, 7, 50; the middle value is still 6. Answer: swapping one value for an extreme outlier dragged the mean from 6 all the way to 14.4, while the median did not move at all — it is still 6, because the outlier’s huge size is irrelevant to the median. Only its rank (still the largest) matters.

History

  • Averaging repeated measurements to cancel out random error is an old practical trick from trade and astronomy; astronomers such as Tycho Brahe were systematically averaging repeated observations by the late 16th century, long before anyone had a formal theory for why it worked.
  • Carl Friedrich Gauss’s early-19th-century work on the normal distribution and the method of least squares (around 1809) showed that the arithmetic mean is the single value that minimizes total squared deviation from a dataset, linking the mean and variance mathematically.
  • Adolphe Quetelet extended averaging from astronomy into social science in the 1830s, applying it to human measurements and popularizing the idea of a statistically “average” person.
  • Francis Galton’s studies of inherited height in the 1880s, running alongside this tradition, introduced regression toward the mean and correlation, closely related tools for describing how data relates and spreads.
  • Karl Pearson coined the term “standard deviation” in 1893, giving the measure of spread the name still used today.
  • Ronald Fisher formalized much of modern statistical methodology in the 1910s-20s, coining the term “variance” itself (1918) and putting sample-based estimation, including the n−1 correction, on rigorous theoretical footing.

Why It Matters

  • Standardized test reporting (SAT, IQ tests) expresses individual scores as percentiles or as a number of standard deviations from the mean, letting one score be compared against an entire population instantly.
  • Manufacturing quality control tracks the mean and standard deviation of a production process continuously; a shift in either signals the process is drifting out of spec before defective parts pile up.
  • A/B testing and experiment analysis in tech and business compare the means of a metric between two groups, using each group’s spread to judge whether an observed difference is real or just noise.
  • Financial risk is routinely measured as the standard deviation of returns (“volatility”): a higher standard deviation means a wider range of plausible outcomes, and therefore a riskier asset, for the same average return.
  • Scientific measurements are reported as a mean plus-or-minus a standard deviation (or standard error) specifically to communicate how much uncertainty surrounds the reported value, not just the value itself.
  • Climate scientists report temperature “anomalies” as deviations from a long-term mean baseline, because the deviation is more informative and comparable across regions than the raw temperature itself.

Common Pitfalls

  • Reporting the mean on heavily skewed data without checking the distribution’s shape first. Income data is the classic case: a handful of extremely high earners pull the mean well above what a “typical” person earns, so the median is usually the more honest summary.
  • Confusing variance with standard deviation. Variance is in squared units (dollars², meters²), which usually has no intuitive real-world meaning; standard deviation is back in the original units and is what should be quoted or plotted.
  • Assuming every dataset has exactly one mode. A dataset can be multimodal (two or more values tied for most frequent), or have no mode at all if every value is unique.
  • Using the population formula (dividing by n) on what is actually a sample meant to estimate a larger population’s variance. That systematically underestimates the true spread, which is exactly why the n−1 sample correction exists.
  • Judging spread from the range alone. The range uses only the two most extreme values and ignores everything in between, so it is far more sensitive to a single outlier than variance or standard deviation.
  • Treating “average” as a synonym for the mean by default. Colloquial “average” often really means “typical,” and depending on the data’s shape, the median or mode can be the more appropriate measure of that.

Comparison

MeasureWhat It MeasuresSensitivity to OutliersWhen It’s Preferred
MeanThe arithmetic center of gravity of every valueHigh — every value pulls it, extremes most of allSymmetric data with no major outliers
MedianThe middle value once the data is sortedLow — only rank matters, not magnitudeSkewed data or data with outliers (income, home prices)
ModeThe single most frequent valueUnaffected by magnitude, but touchy with small samplesCategorical or discrete data; finding the most common outcome

FAQ

Why divide by n−1 for a sample instead of n? A sample’s own mean is calculated from, and therefore sits closer to, that sample’s data points than the true (unknown) population mean would. Dividing by n alone systematically underestimates the population’s true variance. Dividing by n−1 (Bessel’s correction) corrects for that bias, which is most noticeable in small samples.

Can a dataset have more than one mode, or none at all? Yes to both. A dataset is “bimodal” or “multimodal” when two or more values are tied for the highest frequency, and has no mode at all when every value appears exactly once, as in the widget’s default six sliders before you move any of them to match.

Which measure should I trust if the mean and median disagree by a lot? A large gap between mean and median is itself useful information: it signals the data is skewed or contains outliers. In that situation, the median usually better describes a “typical” value, while the size of the gap hints at how extreme the skew is.

Example

A teacher records six quiz scores: 62, 74, 78, 81, 85, 91. These average to a mean of 78.5 and a median of 79.5, close together since the scores are fairly symmetric. If the lowest score had instead been a single badly misbubbled exam scored as a 12 instead of a 62, the mean would fall all the way to about 70.17, while the median would not move at all, still 79.5, because the two middle-ranked scores (78 and 81) never changed. That gap between a shaken mean and an unmoved median is exactly the signature a single outlier leaves behind.

Dig deeper