Overfitting vs Underfitting
Overfitting vs Underfitting
Definition: Overfitting is when a model memorizes training data noise instead of general patterns, performing well on training data but poorly on new data. Underfitting is when a model is too simple to capture the underlying pattern at all.
How It Works
- Overfitting: training accuracy high, validation/test accuracy low — the model learned the training set’s quirks, noise, and outliers rather than the signal that generalizes
- Underfitting: both training and validation accuracy are low — the model lacks capacity, features, or training time to represent the underlying relationship
- A well-fit model sits between the two: training and validation performance are both reasonably high and reasonably close to each other
- The gap between training and validation error is the diagnostic signal — a large gap points to overfitting, a small gap with poor absolute performance points to underfitting
- As model complexity increases (more parameters, more depth, more training epochs), training error monotonically decreases, but validation error follows a U-shape — dropping, then rising again once the model starts fitting noise
- Overfitting can happen gradually within a single training run — early epochs generally improve both train and validation performance together, then at some point validation performance plateaus or degrades while training performance keeps climbing
- Underfitting is sometimes called “high bias” in older statistics literature, and overfitting “high variance” — the two terms are used interchangeably across ML and classical statistics, so expect both vocabularies in papers and interviews
Under the Hood
- This is a direct manifestation of the Bias-Variance Tradeoff: underfitting = high bias (systematic error from an overly simple model), overfitting = high variance (predictions swing wildly depending on which training set you happened to sample)
- Expected test error decomposes as
Error = Bias^2 + Variance + Irreducible Noise— you cannot drive both bias and variance to zero simultaneously with a fixed amount of data; reducing one usually raises the other - Learning curves make this diagnosable: plot training and validation loss against training set size or epoch count — overfitting curves diverge (train keeps dropping, val flattens or rises), underfitting curves converge to a shared high error
- Effective model capacity isn’t just “parameter count” — a network can have millions of parameters yet still underfit if the learning rate is too low or training stops too early, and a tiny model can overfit a tiny dataset
- Double descent is a more recent, counterintuitive finding: for some very high-capacity models (especially deep networks), test error can rise then fall again as capacity keeps increasing past the point of perfectly fitting the training data — this complicates the classic single-U-shaped-curve intuition but doesn’t invalidate the underlying bias-variance logic for typical model sizes
Worked Example: Decomposing Error into Bias and Variance
Abstract formulas are easier to trust once you’ve seen them work on made-up numbers. Suppose the true value we’re trying to predict at some fixed test point is y = 10, and we retrain two candidate models on three different resampled versions of the same training set, then check what each predicts at that one test point every time.
Model A (too simple — underfit): predictions across the three retrains are 7.1, 6.8, 7.3.
- Mean prediction:
(7.1 + 6.8 + 7.3) / 3 = 7.07 - Bias = mean prediction minus true value =
7.07 - 10 = -2.93, soBias^2 ≈ 8.60 - Variance = average squared deviation from the mean ≈
0.04 - Error is almost entirely bias — the model is consistently wrong in the same direction regardless of which training set it saw, the signature of underfitting
Model B (too complex — overfit): predictions across the same three retrains are 12.4, 6.1, 11.2.
- Mean prediction:
(12.4 + 6.1 + 11.2) / 3 = 9.9 - Bias =
9.9 - 10 = -0.1, soBias^2 ≈ 0.01 - Variance = average squared deviation from the mean ≈
7.46 - Error is almost entirely variance — the model is right on average but wildly inconsistent from one training set to the next, the signature of overfitting
Neither model is obviously “worse” from bias or variance alone — a low-bias, high-variance model and a high-bias, low-variance model can land at a similar total error, which is exactly why the diagnostic table below looks at the train/validation gap, not a single error number, to tell them apart.
Why It Matters
- The central tension in ML model design — balancing complexity against generalization is basically the whole job
- Directly informs decisions about regularization, model size, and how much training data you need
- A model that overfits looks impressive on the metric you’re watching (training accuracy) while being silently useless in production — this is one of the most common ways ML projects fail after deployment
- Understanding which failure mode you’re in changes your next move entirely: overfitting calls for more data or Regularization (L1, L2, Dropout); underfitting calls for more capacity, better features, or more training
- Every model selection decision — how many trees in a random forest, how many layers in a network, how long to train — is implicitly a decision about where to sit on this spectrum
What Underfitting and Overfitting Look Like Structurally
Independent of any specific algorithm, the two failure modes have a recognizable shape when you compare what the model is against what the data actually needs:
The two branches are drawn as near mirror images on purpose — underfitting and overfitting aren’t two unrelated bugs, they’re the same underlying tension (how much flexibility to give the model relative to how much signal is actually in the data) resolved in opposite directions.
Diagnosing Overfitting vs Underfitting
| Symptom | Training performance | Validation performance | Diagnosis |
|---|---|---|---|
| Both low | Low | Low | Underfitting |
| Gap is large | High | Low | Overfitting |
| Both high and close | High | High | Good fit |
| Both improving together | Rising | Rising | Still training — not yet converged |
| Validation worse than a naive baseline | Any | Below baseline | Likely a bug (data leakage, label error), not just under/overfitting |
| High variance across CV folds | Consistently high | Swings widely fold to fold | Overfitting to fold-specific noise, even if the averaged score looks fine |
Fixes By Failure Mode
- If overfitting:
- Get more training data, or augment existing data (crops/flips/noise for images, synonym swaps for text)
- Add Regularization (L1, L2, Dropout) — L1/L2 weight penalties, dropout, or early stopping
- Reduce model capacity — fewer layers, fewer parameters, simpler features
- Use Cross-Validation to get a more reliable estimate of generalization performance
- Simplify the feature set — remove noisy or redundant features that give the model more surface area to memorize on
- If underfitting:
- Increase model capacity — more layers, more parameters, a more expressive architecture
- Train longer or raise the learning rate if training hasn’t converged
- Add better or more informative features (Feature Engineering)
- Reduce regularization strength if it’s currently too aggressive
- Check for implementation bugs — a genuinely powerful model that underfits badly sometimes indicates a broken loss function, frozen weights, or a data pipeline bug rather than a true capacity problem
- Remove or reduce aggressive regularization that may have been added preemptively — a common mistake is copying a heavily-regularized config from a much larger model or dataset onto a smaller problem where it’s simply too strong
Neither failure mode is a life sentence — moving between them is exactly what the fixes above are for:
Code Example
from sklearn.model_selection import learning_curve
import numpy as np
train_sizes, train_scores, val_scores = learning_curve(
estimator=model, X=X, y=y, cv=5,
train_sizes=np.linspace(0.1, 1.0, 10), scoring="accuracy"
)
train_mean = train_scores.mean(axis=1)
val_mean = val_scores.mean(axis=1)
# Large, persistent gap between train_mean and val_mean -> overfitting
# Both curves converge at a low score -> underfitting
# In deep learning, the equivalent check is tracking loss per epoch:
# for epoch in range(num_epochs):
# train_loss = run_epoch(model, train_loader, train=True)
# val_loss = run_epoch(model, val_loader, train=False)
# history.append((epoch, train_loss, val_loss))
# then plot history and watch for the divergence point
Interactive Example — Polynomial Degree vs Generalization Error
The U-shaped generalization curve, made concrete: fit polynomials of increasing degree to the same small noisy dataset (true relationship is quadratic), and compare training error against held-out error at each degree. No libraries — just a hand-rolled least-squares solver.
The true relationship is a degree-2 polynomial, so watch for training error dropping smoothly through every degree while held-out error bottoms out around degree 2-3 and then climbs — often sharply — as higher-degree terms start fitting the noise instead of the signal.
Real-World Example
- A résumé-screening model trained on 500 examples from one company can overfit to idiosyncratic details (a specific university’s name, a particular phrasing) rather than learning genuinely predictive signal, then perform poorly when applied to résumés from a different hiring pool
- A linear model predicting housing prices from only square footage will underfit in a market where location, school district, and age of the property all matter — it captures a coarse trend but misses most of the real variance
- Early stopping in production ML pipelines (e.g., an ad click-through-rate model) is often the single most impactful regularizer, since these models are typically trained on truly enormous datasets where overfitting still creeps in after enough passes over the data
- Google Flu Trends (2008-2013) is one of the most widely cited real cautionary tales: a model that nowcast flu prevalence from aggregated search-query patterns performed well for several years, then, in February 2013, was reported to be predicting more than double the flu-like-illness rate the CDC actually observed. A 2014 Science paper by Lazer, Kennedy, King, and Vespignani (“The Parable of Google Flu: Traps in Big Data Analysis”) traced much of the failure to the model latching onto search terms that were correlated with flu during the training window but not causally related to it — a textbook case of a model fitting transient noise in its training data rather than the true underlying signal
- Kaggle leaderboard “shake-ups” are a live, recurring demonstration of overfitting to a validation signal: competitors iteratively tune models against public-leaderboard feedback computed on a held-out slice of data, and a model that climbs the public leaderboard by exploiting quirks of that specific slice often falls sharply once rankings are recomputed on the separate private leaderboard after the competition closes — the public leaderboard score itself becomes something you can overfit to, even without ever training on it directly
- Backtest overfitting in quantitative finance is a well-documented professional hazard: a trading strategy tuned against historical price data through many rounds of trial and error can look consistently profitable in backtests while having simply memorized noise in that specific historical window. Marcos López de Prado’s Advances in Financial Machine Learning (2018) discusses this at length and proposes statistical corrections (like the deflated Sharpe ratio) specifically to account for how many variations were tried before landing on the “winning” one
Common Pitfalls
- Judging a model only by training accuracy, missing an overfitting problem entirely
- Adding regularization or simplifying a model that’s actually underfitting, making it worse
- Tuning hyperparameters against the test set instead of a separate validation set — this leaks test-set information into your decisions and produces an overly optimistic final estimate (a subtle form of overfitting to the test set itself)
- Assuming more training epochs always help — past a certain point, extra epochs just memorize the training set harder
- Comparing models trained on different data splits or preprocessing pipelines and attributing the difference to architecture instead of the actual confound
- Mistaking data leakage (e.g., a feature that indirectly encodes the label) for a great fit — suspiciously perfect validation performance is often a leakage bug, not a well-fit model
- Using a random train/validation split on data with temporal structure — for time-series problems, a random shuffle can leak future information into training, producing validation scores that look like a well-fit model but collapse once deployed against genuinely future data
- Judging fit quality from accuracy alone on an imbalanced dataset — a badly underfit model that always predicts the majority class can still post deceptively high accuracy; see Precision, Recall, and F1 Score and Confusion Matrix for metrics that don’t hide this
- Treating the bias-variance decomposition as if bias and variance must always trade off one-for-one — double descent shows overparameterized models can sometimes reduce both simultaneously past the interpolation threshold, so the classical tradeoff is a strong default intuition, not an ironclad law
Best Practices
- Always hold out a validation set (or use Cross-Validation) distinct from both training and final test data
- Use early stopping: track validation loss during training and stop once it stops improving, even if training loss keeps falling
- Plot learning curves as a routine diagnostic, not just a final-report artifact
- Start simple, then add capacity only as needed — it’s easier to detect and fix underfitting than to unwind overfitting after the fact
- Re-evaluate the fit/overfit balance whenever you change the data pipeline, not just when you change the model — new features or preprocessing can shift a model from underfit to overfit or vice versa
- Vary training-set size, not just epoch count, when plotting learning curves — a gap that closes as you add more data points to a fixed-capacity model tells you something a per-epoch loss curve can’t: whether you’re data-starved rather than capacity-starved
- Fix the train/validation split (or CV folds) when comparing candidate models head-to-head, so that differences in score reflect the model, not random split variance
- Track a task-specific metric alongside loss, not loss alone — a model can show a healthy, smoothly-converging loss curve while a metric that actually matters for the task (F1, calibration, ranking quality) quietly degrades in the background
History
- The bias-variance decomposition was formalized in a widely-cited 1992 paper by Geman, Bienenstock, and Doursat, giving statistical language to a tension practitioners had long observed informally
- Vladimir Vapnik and Alexey Chervonenkis developed VC (Vapnik-Chervonenkis) dimension theory in the 1970s, providing a formal measure of a model class’s capacity and a mathematical basis for why unconstrained capacity leads to poor generalization — this theory underlies why techniques like regularization and cross-validation work, not just that they work
- The double descent phenomenon, where test error can fall, rise, then fall again as model capacity keeps increasing past the interpolation threshold, was popularized in deep learning contexts around 2018-2019 and forced a partial rethink of the classical single-U-shaped bias-variance story for very overparameterized models
- Occam’s razor — “prefer the simplest explanation that fits the facts” — is the philosophical ancestor of regularization, dating back centuries before it was formalized mathematically as a penalty term in a loss function
- The mathematical seed of L2 regularization traces to Andrey Tikhonov’s 1943 work on stabilizing ill-posed inverse problems; Arthur Hoerl and Robert Kennard applied the same idea to statistical regression in 1970, naming it ridge regression, decades before “L2 regularization” became standard deep learning vocabulary
- Cross-validation’s formal statistical grounding dates to Mervyn Stone’s 1974 paper on cross-validatory assessment of predictions and Seymour Geisser’s 1975 paper on predictive sample reuse — both predate its everyday use as a machine learning practice by decades
- Leo Breiman’s work on bagging (1996) and random forests (2001) gave practitioners a direct, algorithmic way to trade variance for stability by averaging many independently overfit models, rather than constraining any single model’s capacity
FAQ
Q: Is a small train/validation gap always fine? Only if absolute performance is also acceptable. A small gap with poor performance on both sets is underfitting, not success.
Q: Can a model overfit and underfit at the same time? Yes, in different regions of the input space — e.g., it can memorize dense regions of the training distribution while failing to capture the pattern in sparse regions.
Q: Does more data always fix overfitting? Usually helps, but not always — if the new data comes from a different distribution, or the model architecture is fundamentally too flexible for the problem, more data alone won’t close the gap.
Q: Does overfitting only happen with too little data? No — a sufficiently flexible model (like a very deep, wide neural network) can overfit even large datasets if left training long enough without regularization or early stopping; “too little data” is one common cause, not the only one.
Q: Is overfitting always a bad thing? In almost every applied setting, yes — the entire point of a model is to generalize to data it hasn’t seen. The one common exception is one-off curve-fitting or interpolation over data you already have in full and never intend to generalize beyond, where “overfitting” in the statistical sense is simply not the objective being optimized for.
Common Interview Questions
Q: How would you tell, from a single number, whether a model is overfitting? You can’t from one number — you need at least a training-set score and a validation-set score to compute the gap. A single accuracy figure alone is uninformative about generalization.
Q: Why does cross-validation help detect overfitting better than a single train/validation split? A single split gives one noisy estimate of generalization error, which can look fine or bad by chance. Cross-validation averages over multiple splits, giving a more stable estimate and revealing high variance across folds if the model is overly sensitive to which data it sees.
Q: What’s the relationship between overfitting and the number of training epochs? Training loss almost always keeps dropping with more epochs. Validation loss drops too, up to a point, then typically rises again as the model starts fitting training-set-specific noise — this rising point is exactly what early stopping is designed to catch.
Q: If you had to pick just one technique to reduce overfitting, what would you pick and why? Getting more (or better) training data is usually the highest-leverage fix, since it directly attacks the root cause — insufficient signal relative to model flexibility — rather than compensating for it after the fact the way regularization does. When more data isn’t available, early stopping is the cheapest and least invasive second choice.
Q: How does model ensembling relate to the overfitting/underfitting tradeoff? Averaging predictions from multiple models trained with different random seeds, data subsets, or architectures reduces variance (since individual models’ errors partially cancel out) without necessarily increasing bias — see Ensemble Methods. This is why ensembles often outperform any single constituent model on held-out data even when no individual member is regularized particularly hard.
Q: A model has 99% training accuracy and 97% validation accuracy — overfitting or not? Judge the gap in context, not in isolation: a 2-point gap is unremarkable for most problems and likely just the sweet spot, but if 97% is far below a domain-appropriate ceiling (or below a simpler baseline), the small gap can coexist with a real, separate problem — the numbers alone don’t tell the whole story without knowing what “good” looks like for that task.
Q: Why might a model with excellent cross-validation scores still fail in production? Cross-validation only measures generalization to data drawn from the same distribution as the training set. If production data drifts from that distribution over time, or the CV folds themselves shared some leakage (e.g., duplicate or near-duplicate rows split across folds), strong CV scores can coexist with a model that was effectively overfit to assumptions that no longer hold once deployed.
Overfitting and Underfitting Beyond Classic Supervised Learning
- Reinforcement learning: an agent can overfit to the specific quirks of its training environment — memorizing a fixed set of game levels or simulator physics rather than learning a strategy that generalizes — which is exactly why procedurally-generated benchmarks like Procgen were built: they measure whether an agent trained on some levels can perform on entirely new ones, the RL analogue of a held-out validation set. See Reinforcement Learning
- LLM benchmark contamination: when a benchmark’s test questions leak into a model’s pretraining data (directly or via near-duplicate text scraped from the web), the model can score well not because it generalizes but because it has effectively seen the answers — a dataset-level form of overfitting that’s become a major methodological concern for evaluating neural network-based language models honestly
- Underfitting at scale — the Chinchilla finding: Hoffmann et al.’s 2022 DeepMind paper found that many large language models of that era were undertrained relative to their parameter count — that is, underfit relative to what their capacity could achieve given enough data. Their 70-billion-parameter Chinchilla model, trained on 1.4 trillion tokens, outperformed the 280-billion-parameter Gopher model trained on roughly 300 billion tokens, using comparable total compute — reallocated toward more training data rather than more parameters. This reframed “underfitting” as a compute-allocation problem, not just a model-size problem
Related Terms
- Bias-Variance Tradeoff
- Regularization (L1, L2, Dropout)
- Cross-Validation
- Hyperparameter Tuning
- Precision, Recall, and F1 Score
- Ensemble Methods
- Feature Engineering
- Reinforcement Learning
- Neural Network
Deeper Dive: Reading a Loss Curve
- A textbook overfitting curve: training loss decreases smoothly and continuously, validation loss decreases alongside it initially, then bottoms out and starts climbing while training loss keeps falling — the widening gap after that bottom point is the overfitting region
- A textbook underfitting curve: both training and validation loss decrease together but plateau early at a high value, with little to no gap between them — more training time alone won’t close that gap because the model has run out of capacity to represent the pattern, not out of time to learn it
- A healthy curve: both losses decrease together and plateau close to each other at a low value — this is the target state, and it’s what early stopping tries to freeze the model at before the validation curve turns upward
The same three curves, viewed as a single training run and the stopping point you choose along it, rather than three separate curves:
Example
A model that gets 99% accuracy on training data but only 60% on a test set is overfitting — it memorized rather than learned to generalize. Conversely, a linear regression trying to fit a clearly curved (quadratic) relationship in the data will underfit: no matter how much data you give it, a straight line simply cannot capture the curve, so both its training and test error stay stubbornly high. The fix in the first case is more data, regularization, or a smaller model; the fix in the second case is a model with more expressive power, such as adding polynomial features or switching to a nonlinear model.
A practical middle-ground case: a decision tree with no depth limit will keep splitting until every training leaf is pure, achieving near-perfect training accuracy by essentially memorizing individual examples — capping max_depth or setting a minimum samples-per-leaf constraint trades away some training accuracy for a tree that captures the general shape of the data instead of its exact contents, which is why tree-based models expose these knobs directly as regularization hyperparameters.
Referenced by
- AI Bias and Fairness
- Bias-Variance Tradeoff
- Computer Vision
- Confusion Matrix
- Cross-Validation
- Ensemble Methods
- Epoch, Batch, and Iteration
- Explainable AI (XAI)
- Feature Engineering
- Fine-Tuning
- Hallucination
- Hyperparameter Tuning
- Learning Rate
- Loss Function
- Machine Learning and Deep Learning Terms MOC
- Neural Network
- Object Detection
- Precision, Recall, and F1 Score
- Prompt Engineering
- Regularization (L1, L2, Dropout)
- Supervised Learning