Precision, Recall, and F1 Score
Precision, Recall, and F1 Score
Definition: Classification metrics derived from the confusion matrix. Precision measures how many predicted positives were actually correct; recall measures how many actual positives were correctly found; F1 is their harmonic mean, balancing both.
How It Works
- Precision = True Positives / (True Positives + False Positives) — “when the model says yes, how often is it right?”
- Recall = True Positives / (True Positives + False Negatives) — “of all actual positives, how many did the model catch?”
- F1 = 2 × (Precision × Recall) / (Precision + Recall) — the harmonic mean, useful when you need a single balanced number
- All three metrics come directly from the four cells of a Confusion Matrix: True Positive (TP), False Positive (FP), True Negative (TN), False Negative (FN)
- Precision and recall are in tension — pushing a classifier’s decision threshold to catch more positives (raising recall) almost always lets in more false positives (lowering precision), and vice versa
- Both metrics ignore true negatives entirely, which is exactly why they’re preferred over accuracy on imbalanced problems where true negatives can be overwhelmingly numerous and uninformative
- Both are computed at a specific decision threshold — a probabilistic classifier (logistic regression, a neural net with a sigmoid/softmax output) produces a continuum of scores, and precision/recall change as that threshold slides, which is why a single reported number is only ever a snapshot at one operating point
Two Questions, One Confusion Matrix
Precision and recall are computed from the exact same four counts — they just ask different questions of them, which is precisely why a model can score well on one while doing poorly on the other:
Precision never looks at FN, and recall never looks at FP — each metric is blind to exactly the error type the other one measures. F1 exists because neither question alone answers “is this model actually good,” and averaging the two answers back together into one number needs to be done carefully, which is what the harmonic mean is for (see Under the Hood below).
Under the Hood — Worked Example
Say a spam classifier is evaluated on 1,000 emails, 50 of which are actually spam. The model predicts:
| Predicted Spam | Predicted Not Spam | |
|---|---|---|
| Actually Spam | TP = 40 | FN = 10 |
| Actually Not Spam | FP = 30 | TN = 920 |
- Precision = 40 / (40 + 30) = 40/70 ≈
0.571— of the emails flagged as spam, 57.1% actually were - Recall = 40 / (40 + 10) = 40/50 =
0.80— the model caught 80% of all real spam - F1 = 2 × (0.571 × 0.80) / (0.571 + 0.80) ≈
0.667 - Accuracy, for comparison, is (TP+TN)/(TP+TN+FP+FN) = 960/1000 =
0.96— deceptively high because the negative class (920 non-spam) dominates the denominator; accuracy barely reflects how well spam itself is detected - The harmonic mean (F1) punishes imbalance between precision and recall more harshly than a plain average would — if precision = 1.0 and recall = 0.01, the arithmetic mean is ~0.5 but F1 is ~0.02, correctly signaling the model is nearly useless despite one perfect-looking number
- Specificity, a related but distinct metric, uses the two cells precision/recall ignore:
Specificity = TN / (TN + FP)= 920/950 ≈0.968— “of all actual non-spam, how many were correctly left alone” - Negative predictive value (NPV), the mirror image of precision, is
TN / (TN + FN)= 920/930 ≈0.989— “when the model says not-spam, how often is that correct” — rarely reported but occasionally important when a false negative is the costly error, such as in medical rule-out testing - G-Mean for this same model is
sqrt(recall × specificity)=sqrt(0.80 × 0.968)≈0.880— notice it sits between recall and specificity rather than being dragged down toward whichever is smaller as harshly as F1 would be, because it’s balancing performance on the positive and negative classes rather than precision and recall
Sliding the Threshold: Precision vs. Recall
A single confusion matrix is a snapshot at one threshold. To see the actual tradeoff, sweep the threshold across the same set of scores and watch precision and recall move in opposite directions. Take 8 emails a model has scored with a probability of being spam:
| True label | Score | |
|---|---|---|
| 1 | Spam | 0.90 |
| 2 | Spam | 0.75 |
| 3 | Spam | 0.55 |
| 4 | Spam | 0.35 |
| 5 | Not spam | 0.65 |
| 6 | Not spam | 0.45 |
| 7 | Not spam | 0.25 |
| 8 | Not spam | 0.10 |
Classify “spam” whenever the score meets or exceeds the threshold, and recompute the confusion matrix and every metric at each threshold:
| Threshold | TP | FP | FN | TN | Precision | Recall | F1 |
|---|---|---|---|---|---|---|---|
| 0.30 (permissive) | 4 | 2 | 0 | 2 | 0.667 | 1.000 | 0.800 |
| 0.50 (default) | 3 | 1 | 1 | 3 | 0.750 | 0.750 | 0.750 |
| 0.70 (conservative) | 2 | 0 | 2 | 4 | 1.000 | 0.500 | 0.667 |
At threshold 0.30, emails 4, 5, and 6 all cross into “predicted spam” — every real spam email is now caught (recall = 1.0), but 2 of the 6 flagged emails (5 and 6) aren’t spam at all, so precision drops to 0.667. At threshold 0.70, only the two highest-confidence predictions count as spam — both really are spam, so precision is a perfect 1.0, but two real spam emails (3 and 4) score below the bar and are missed, so recall falls to 0.5. F1 happens to peak at the middle threshold here, but that’s an artifact of this particular data, not a general rule — the right threshold is whichever one matches the real cost of a false positive against the cost of a false negative, not whichever one maximizes F1. This table is a hand-computed slice of the full precision-recall curve mentioned below: plot precision and recall at every achievable threshold instead of just three, and the resulting curve is the standard tool for picking an operating point instead of guessing.
The two extremes are worth knowing by heart because they bound every curve. At threshold 0 (or below every score), the model predicts “positive” for everything: recall is always exactly 1.0 (nothing is missed), and precision falls to the dataset’s raw positive rate (4/8 = 0.5 here) since every example, relevant or not, gets flagged. At threshold 1 (or above every score), the model predicts “negative” for everything: recall collapses to 0.0 (nothing is caught), and precision becomes undefined — there are no predicted positives to be right or wrong about, which is exactly the discontinuity that makes the left edge of a plotted precision-recall curve look different from every other point on it.
Variants
- Fβ score — a generalized F1 that weights recall β times as important as precision:
Fβ = (1+β²) × (precision × recall) / (β²×precision + recall). F2 favors recall (e.g., disease screening), F0.5 favors precision (e.g., spam filtering where false positives are costly) - Macro-averaging — compute precision/recall/F1 per class independently, then take the unweighted mean across classes; treats every class as equally important regardless of size
- Micro-averaging — pool all TP/FP/FN counts across classes first, then compute one global precision/recall/F1; dominated by the performance on large classes
- Weighted averaging — like macro, but each class’s score is weighted by its support (number of true instances), balancing between macro and micro
- Precision-Recall curve — plots precision vs recall across every possible decision threshold; the area under this curve (PR-AUC) summarizes performance across all thresholds at once, and is more informative than ROC-AUC on heavily imbalanced data
- Matthews Correlation Coefficient (MCC) — a single-number metric that uses all four confusion matrix cells symmetrically, ranging from -1 (total disagreement) to +1 (perfect prediction); often considered more robust than F1 on imbalanced binary classification since it accounts for true negatives too
- G-Mean — the geometric mean of recall (sensitivity) and specificity,
sqrt(recall × specificity); common in the imbalanced-learning literature because, like MCC, it collapses toward zero if the model does badly on either class instead of letting good performance on one class mask poor performance on the other
Why It Matters
- Plain accuracy is misleading on imbalanced datasets (e.g., 99% accuracy on 1%-fraud data by always predicting “not fraud”)
- The right metric to optimize for depends entirely on the cost of false positives vs false negatives in your specific problem
- These metrics let you explicitly choose a tradeoff via the decision threshold, rather than being stuck with whatever a single fixed cutoff (typically 0.5) happens to produce
- Multi-class problems (10 classes, 1000 classes) still reduce to these same building blocks per-class, then get aggregated — understanding the aggregation method matters as much as the base metric
- These are the metrics stakeholders outside ML actually understand and can weigh in on — “we’d rather miss a few more fraud cases than annoy this many legitimate customers” is a business conversation that precision/recall makes possible, unlike a single opaque accuracy number
- Comparing two models fairly requires agreeing on how each was tuned — a model with better raw ranking ability (higher ROC-AUC or PR-AUC) can still lose on F1 at the specific threshold each was independently calibrated to, so “which model is better” depends partly on the tuning procedure, not purely on what each model is capable of
Code Example
from sklearn.metrics import precision_score, recall_score, f1_score, classification_report
y_true = [1, 0, 1, 1, 0, 1, 0, 0, 1, 0]
y_pred = [1, 0, 0, 1, 0, 1, 1, 0, 1, 0]
print(precision_score(y_true, y_pred)) # 0.8
print(recall_score(y_true, y_pred)) # 0.8
print(f1_score(y_true, y_pred)) # 0.8
# Multi-class: choose an averaging strategy explicitly
print(f1_score(y_true_multiclass, y_pred_multiclass, average="macro"))
print(f1_score(y_true_multiclass, y_pred_multiclass, average="weighted"))
print(classification_report(y_true, y_pred)) # per-class precision/recall/F1 + support
# sweeping thresholds to find the best operating point for your use case
from sklearn.metrics import precision_recall_curve
precisions, recalls, thresholds = precision_recall_curve(y_true, model_probabilities)
Interactive: Precision, Recall, and F1 From a Confusion Matrix
Same four counts as the worked example above (1,000 emails, 50 actually spam) — computed directly instead of by hand:
Comparison: Precision vs Recall by Use Case
| Use case | Prioritize | Why |
|---|---|---|
| Cancer screening | Recall | Missing a real case (FN) is far worse than a false alarm that gets ruled out later |
| Spam filtering | Precision | A false positive (real email in spam folder) is worse than letting one spam email through |
| Fraud detection | Recall (often, with human review) | Missing fraud is costly, but flagged cases get manually reviewed to catch FPs |
| Search/recommendation ranking | Precision (top-k) | Users only see the top few results; low-ranked false positives barely matter |
| Legal e-discovery | Recall | Missing a relevant document can be legally consequential |
| Manufacturing defect detection | Recall | A shipped defective part is typically costlier than a false alarm that triggers extra inspection |
| Content moderation | Depends on category | Legal-risk categories favor recall even with more false takedowns; borderline creative content favors precision to avoid over-removal |
| Loan default prediction | Depends on cost | A missed default (FN) costs the lender the loan amount; a false positive (FP) costs a creditworthy customer — the ratio of those two costs sets the threshold |
| Customer churn prediction | Recall (often) | Missing a customer about to churn forecloses any chance to intervene; a false positive just means an unnecessary retention offer |
Real-World Example
- A hospital deploying a sepsis-prediction model deliberately tunes the threshold toward higher recall, accepting more false alarms, because a missed sepsis case can be fatal while a false alarm just triggers an extra check by a nurse
- An email provider tunes spam filtering toward higher precision — Google has reported Gmail blocking upward of 99.9% of spam while explicitly holding the line on false positives, because users are far more upset by a legitimate email lost in spam than by an occasional spam email reaching the inbox
- A/B testing a new fraud model often reports precision and recall separately to different stakeholders — the fraud operations team cares about recall (catch more fraud), while customer support cares about precision (fewer angry calls about blocked legitimate transactions)
- Content moderation systems on social platforms typically report precision and recall by policy category separately (hate speech, spam, violence), since the acceptable tradeoff differs by category — a platform may tolerate lower precision (more false takedowns) for content with legal risk while requiring high precision for borderline creative content
- Algorithmic risk scoring and fairness (COMPAS). ProPublica’s 2016 analysis of the COMPAS criminal recidivism-risk tool found its errors weren’t evenly distributed by race: among defendants who did not reoffend, Black defendants were nearly twice as likely to have been wrongly flagged high-risk (a false positive) as white defendants, while white defendants who did reoffend were more often wrongly scored low-risk (a false negative) than Black defendants were. Northpointe, the tool’s vendor, countered that its scores were equally well-calibrated across race. Both claims held simultaneously — Kleinberg, Mullainathan, and Raghavan’s 2016 paper “Inherent Trade-Offs in the Fair Determination of Risk Scores” proved that matching false-positive rates and matching calibration across groups are mathematically incompatible whenever the underlying base rates differ, meaning which confusion-matrix-derived metric you choose to equalize is a value judgment baked into the model, not a detail that can be optimized away
Common Pitfalls
- Optimizing accuracy alone on an imbalanced dataset, producing a model that’s technically accurate but practically useless
- Chasing high recall without considering precision cost (e.g., a cancer screening tool that flags everyone as positive has perfect recall but is useless)
- Reporting a single F1 score on a multi-class problem without specifying macro/micro/weighted — these can differ substantially and aren’t interchangeable
- Comparing F1 scores across datasets with different class balances as if they were on the same scale
- Picking the default 0.5 classification threshold without checking whether it actually matches the precision/recall tradeoff the business needs
- Evaluating on the same data used for threshold tuning, inflating the reported metric
- Reporting precision/recall computed on a validation set that was also used for feature selection or model/hyperparameter selection, which inflates the numbers beyond what a genuinely held-out test set would show
- Treating “equalize precision across groups” and “equalize recall across groups” as simultaneously achievable goals — as the COMPAS case above shows, they generally aren’t once groups have different base rates
- Forgetting that precision is undefined (division by zero) when a model predicts zero positives — a naive implementation that doesn’t special-case this can crash or silently return a misleading default instead of flagging that the model made no positive predictions at all
Best Practices
- Always report precision, recall, and a confusion matrix together — F1 alone hides which type of error dominates
- Choose the averaging strategy (macro/micro/weighted) based on whether classes matter equally or in proportion to their frequency, and state which one you used
- Tune the decision threshold explicitly using a validation set and a precision-recall curve, rather than accepting the default 0.5
- For heavily imbalanced problems, prefer PR-AUC over ROC-AUC as the summary metric — ROC-AUC can look good even when precision is poor
- Track these metrics per class in multi-class problems, not just the aggregate — a good weighted average can hide one class performing terribly
- Guard against the divide-by-zero case explicitly (a model predicting no positives at all) rather than relying on a library’s silent default, since that default can mask a completely broken model as merely “undefined”
- Recompute and monitor precision/recall after deployment, not just at training time — real-world class balance and input distributions drift, and a threshold tuned on last quarter’s data may no longer match this quarter’s tradeoff
- When a metric must be equalized across subgroups, decide explicitly which one — false-positive rate, false-negative rate, or calibration — matters most for the application, since the fairness literature shows equalizing all of them at once generally isn’t possible
- Report variance across cross-validation folds alongside point estimates of precision/recall/F1, especially on small test sets where a handful of examples can swing the numbers substantially
- Sanity-check the two extreme thresholds (predict-all-positive and predict-all-negative) before trusting a curve or a tool’s threshold sweep — a bug in label encoding or a flipped positive class often shows up immediately as those extremes not matching the expected base-rate precision and undefined precision described above
FAQ
Q: Why not just always maximize F1? F1 assumes precision and recall are equally important, which is a business decision, not a mathematical fact. Use Fβ or a custom cost function when the tradeoff is asymmetric.
Q: Why “harmonic” mean instead of a regular average? The harmonic mean stays close to the smaller of the two numbers, so F1 only looks good when both precision and recall are reasonably high — a regular average would let one very high value mask one very low value.
Q: Can precision and recall both be 1.0? Yes, if the classifier makes zero false positives and zero false negatives — a perfect classifier on that dataset.
Q: What does it mean if recall is high but F1 is close to zero? Precision must be extremely low. F1’s harmonic mean collapses toward zero whenever either input is near zero, so high recall paired with a near-zero F1 always points to a precision problem — typically a model flagging nearly everything as positive.
Q: Is it possible for accuracy to go up while F1 goes down? Yes, easily, on imbalanced data. A threshold change that trades a few true positives for many more true negatives can raise accuracy (dominated by the large negative class) while lowering recall enough that F1 drops — another reason not to treat accuracy as a stand-in for these metrics.
Q: Does a higher F1 always mean a better model? Not necessarily. F1 assumes equal weight on precision and recall and ignores true negatives entirely, so two models with identical F1 can behave very differently — a model with a slightly lower F1 but a precision/recall balance that better matches the application’s actual costs can be the genuinely better choice.
Common Interview Questions
Q: A model has 95% accuracy on a dataset that’s 95% negative class. Is it good? Not necessarily — a classifier that always predicts “negative” would also score 95% accuracy while being useless. You need to check precision and recall on the positive class specifically to know anything meaningful.
Q: How do you choose a classification threshold in practice? Plot the precision-recall curve across thresholds, then pick the operating point that matches the actual cost ratio of false positives to false negatives for the business problem — not a default value like 0.5.
Q: What’s the difference between macro-F1 and micro-F1 on an imbalanced dataset, concretely? Macro-F1 treats a rare class’s poor performance as equally important as a common class’s performance, so it drops sharply if the model fails on rare classes. Micro-F1 is dominated by the common classes and can stay high even if rare-class performance is poor.
Q: If precision and recall are both 0.6, is F1 higher, lower, or equal to 0.6? Equal — when precision and recall are identical, their harmonic mean equals that same value. The harmonic mean only pulls below the arithmetic mean when the two inputs differ from each other.
Q: How would you explain precision and recall to a non-technical stakeholder? Precision answers “when the model raises a flag, how often is it actually right?” Recall answers “of all the real cases out there, how many did the model actually catch?” A model can score well on one while doing poorly on the other, so ask which mistake is more expensive for the business before picking which one to prioritize.
Q: Two models have the same F1 score — how do you tell which one actually fits your use case better? Look past F1 to each model’s individual precision and recall — one might reach that F1 with a balanced tradeoff while the other reaches it lopsided toward one side; compare each against the real cost of a false positive versus a false negative for your application before choosing.
Q: Why does scikit-learn warn about “ill-defined” precision when a model predicts no positives at all? Because precision’s denominator, TP + FP, is zero when nothing is predicted positive — a division by zero. Libraries typically substitute 0 by convention and emit a warning rather than raising an error, which is why silently ignoring that warning can hide a model that has effectively stopped making positive predictions altogether.
History
- Precision and recall originate in information retrieval research — the Cranfield experiments, run by Cyril Cleverdon starting in 1957 with US National Science Foundation funding, were the first to formalize the two as evaluation metrics for document search systems: “precision” meant relevant-among-retrieved, “recall” meant retrieved-among-relevant, the same definitions used in ML today
- The F-measure was introduced by Cornelis Joost van Rijsbergen in the 1970s as a way to combine precision and recall into a single tunable score via the beta parameter, generalizing what’s now commonly used as the unweighted F1
- These metrics migrated from information retrieval into general statistical classification and later machine learning as the field grew, replacing or supplementing pure accuracy once practitioners repeatedly ran into the class-imbalance problem accuracy can’t handle
- The Cranfield methodology’s evaluation tradition continued through NIST’s Text REtrieval Conference (TREC), launched in 1992, which scaled precision/recall evaluation up to far larger, more realistic document collections and shaped how retrieval and ranking systems are still benchmarked today
Deeper Dive: The ROC Curve and AUC
- The ROC (Receiver Operating Characteristic) curve plots the true positive rate (recall) against the false positive rate (
FP / (FP + TN)) across every threshold, and originated in WWII-era signal detection theory for radar operators distinguishing real targets from noise - ROC-AUC (area under that curve) summarizes ranking quality independent of any single threshold, and is popular for balanced classification problems
- On heavily imbalanced data, ROC-AUC can look deceptively good because the false positive rate denominator (TN-heavy) stays small even when precision (which uses FP directly against a much smaller TP count) is poor — this is exactly why PR-AUC is generally preferred over ROC-AUC once the positive class becomes rare
- A model can have a high ROC-AUC (~0.95) while still having mediocre precision at any threshold a business could realistically operate at — always sanity-check ROC-AUC against the precision-recall curve on imbalanced problems before trusting it as the headline metric
- A useful sanity anchor: a random classifier’s PR curve is a flat horizontal line at precision equal to the positive class’s base rate, while a random classifier’s ROC curve is always the diagonal from (0,0) to (1,1) regardless of base rate — which is exactly why PR-AUC is far more sensitive to class imbalance than ROC-AUC is
- Computing PR-AUC via naive linear interpolation between two achieved points can overstate performance, because unlike ROC space, precision doesn’t interpolate linearly between points — this is why libraries typically compute “average precision” (a step-function summation over the actual achieved points) rather than trapezoidal interpolation, for a more honest PR-AUC estimate
Related Terms
- Confusion Matrix
- Supervised Learning
- Overfitting vs Underfitting
- Cross-Validation
- Hyperparameter Tuning
- Bias-Variance Tradeoff
- Feature Engineering
Example
A fraud detector with high recall but low precision catches almost all fraud but also falsely flags many legitimate transactions — a tradeoff that must match the business’s tolerance for each type of error. If a bank sets the model’s threshold low to maximize recall (catch every possible fraud case), it accepts more false positives and routes more legitimate transactions to manual review; raising the threshold trades some missed fraud for fewer customer complaints about blocked purchases. The right operating point isn’t a modeling question at all — it’s determined by comparing the dollar cost of a missed fraud case against the cost (in support tickets and lost trust) of a wrongly blocked transaction.
Referenced by