Signal archive

Signal

Issue 9 · September 9, 2026 · 6 min read

Precision and Recall: The Metric Costing Banks Billions

Precision and recall explained with the formula, multiclass averaging, PR vs ROC curves, and a graded project that prices a 92% false-positive alert queue.

"To measure is to know." - Lord Kelvin (William Thomson)

A fraud model with 99.8% accuracy can be worthless. A model with 60% accuracy can be worth millions. The difference is which mistakes you count.

This week we're talking about precision and recall, the two numbers that decide whether a classifier ships, and yes, this can get you a very comfortable job.

Here's the idea: precision and recall are the two classification metrics that split a model's mistakes into false alarms and misses. Precision is the share of positive predictions that were right. Recall is the share of real positives the model caught. Accuracy hides the trade-off between them; precision and recall expose it.

Now look at what a precision problem costs. Industry reports on anti-money-laundering transaction monitoring put the false positive rate of legacy alert systems at 85% to 95%. Facctum's 2026 AML report and coverage in Yahoo Finance both use the 95% figure. Each alert costs roughly $30 to $70 for a human analyst to clear. A mid-size bank generating 100,000 alerts a year is spending $3 million to $7 million investigating them, and 95 cents of every dollar goes to false alarms. That is a precision of about 0.05. Nobody at that bank cares about accuracy, because a model that flags nothing scores 99%+ on accuracy and gets the bank fined. They care about recall on real laundering and precision on the alert queue, and they will pay for an engineer who can move both.

The same trade-off runs through every field where positives are rare. Cancer screening wants recall, because a miss is a death and a false alarm is a follow-up scan. Spam filters want precision, because a lost invoice is worse than one spam email. Retrieval-augmented generation has its own version: context precision and context recall measure whether the chunks you retrieved were relevant and whether you retrieved all the relevant ones. If you learned statistics for ML, this is where it turns into money.

How do you calculate precision and recall for multiclass classification?

Build the confusion matrix, then treat each class as its own binary problem. For one class, its true positives are the diagonal entry, its false positives are the rest of that column, and its false negatives are the rest of that row. Macro averaging takes the mean of per-class scores and treats every class equally. Micro averaging pools all the true positives, false positives, and false negatives first and is dominated by big classes. Weighted averaging scales each class by its support. The 248,000-view StackExchange question on exactly this exists because people report "precision" without saying which average they used.

Should you use precision and recall or the ROC curve?

Use the precision-recall curve when positives are rare. ROC-AUC uses the false positive rate, whose denominator is the huge negative class, so a model can add thousands of false alarms and barely move on the ROC plot. Precision feels every one of them. Saito and Rehmsmeier (2015) showed this formally on imbalanced data. Average precision, the area under the PR curve, is the number to report for fraud, defects, and rare-event detection.

PrecisionRecall
Question it answersWhen the model says yes, how often is it right?Of all real positives, how many did it catch?
DenominatorAll predicted positivesAll actual positives
Error it punishesFalse positivesFalse negatives
Raise it byRaising the thresholdLowering the threshold
Prioritize whenEach alert is expensive to handleEach miss is expensive to suffer
Also calledPositive predictive valueSensitivity, true positive rate

Now I'm going to explain what this means for you:

The Skill Employers EXPECT

Learn to compute precision and recall by hand and from a confusion matrix. Interview loops assume you can:

  • Write precision, recall, and F1 from the confusion matrix counts without looking anything up
  • Explain what happens to each metric when you move the decision threshold
  • Handle the zero-denominator case (no predicted positives) and say what scikit-learn does about it
  • Compute macro, micro, and weighted averages for a 3-class confusion matrix on a whiteboard
  • Say why accuracy is the wrong headline metric for a 0.2% positive rate

If your answer to "which one matters here?" is "both," you have not thought about the cost of each error.

The Skill That Separates You

Learn to pick the threshold from the cost of mistakes, not from 0.5.

Strong engineers turn precision and recall into a business decision. You should know how to:

  • Sweep thresholds, plot the precision-recall curve, and report average precision
  • Build a cost matrix (a false negative costs 20x a false positive) and pick the threshold that minimizes expected cost
  • Bootstrap confidence intervals on precision and recall so a 0.02 change is not mistaken for a win
  • Compute precision@k and recall@k for retrieval and RAG pipelines, where "positives" are relevant documents
  • Watch precision drift in production as the base rate changes, even when the model has not

The Project To Learn These Skills This Week

Build every classification metric from scratch, then use them to price a real alert queue. I built this one into GRADuateML as a graded Pro project: AML Alert Triage: Precision vs Recall With Real Costs. You get a 60,000-row synthetic transaction-monitoring queue where about 8% of legacy alerts are confirmed suspicious, a legacy risk score to beat, a hidden test set, and instant grading on every milestone. If you are not on Pro, the free challenge at the bottom of this email covers the core metric by hand.

Requirements (each one is a graded milestone):

  • Size the queue: confirmation rate, the legacy false-positive rate, and which rule is noisiest.
  • Implement precision, recall, F1, and average precision in NumPy only on the legacy score at threshold 70. Match the reference within 1e-4. Average precision must be computed over distinct thresholds, the way scikit-learn defines it.
  • Break it on purpose: auto-close every alert and record the 92% accuracy next to the 0% recall.
  • Train a ranking model and beat the legacy engine: PR-AUC of at least 0.36 against the legacy score's 0.17.
  • Fit the analyst queue: recall of at least 0.62 while alerting on no more than 25% of the queue.
  • Price the threshold: a confirmed case is worth $1,500 in avoided regulatory exposure, a false alarm costs $50 of analyst time. Sweep thresholds on a held-out split and hit at least $1.25M net benefit.
  • Score the full batch in under 8 seconds, then write the memo: the precision and recall at your threshold, the alerts per day it produces, the cases you will miss, and what changes if analyst capacity doubles.

This project teaches a practical lesson for production ML:

A metric is a decision about which mistake you are willing to pay for. Pick it before you train.

If you want to sharpen your machine learning skills even more, I also selected a free challenge problem for you this week. It is the multiclass macro-versus-micro question from above, by hand:

If you learned something from this newsletter, make sure to forward it to a friend.

Precision and Recall from a Confusion Matrix

Easy · ~15 min

Concept: model evaluation

GRADuateML

The full AML alert triage build, with hidden-test grading on every milestone, is part of GRADuateML Pro.

Ready to practice?

Turn weekly insights into hands-on ML skills on GRADuateML.