A hypothetical alert model reports 99% accuracy. That sounds reassuring until the evaluation set turns out to contain 990 ordinary events and just 10 events that require an alert. Predicting “ordinary” for every event already gets 990 of 1,000 labels right. It finds none of the 10 cases the alert exists to catch.
Here is that hypothetical evaluation as a confusion matrix. Rows are actual labels; columns are predictions, following scikit-learn’s convention.
| Actual / predicted | Ordinary | Alert |
|---|---|---|
| Ordinary | 990 | 0 |
| Alert | 10 | 0 |
The 990 ordinary events correctly dismissed are true negatives. The 10 missed alerts are false negatives. There are zero true positives and zero false positives. Accuracy is (990 + 0) / 1,000 = 99%. Positive-class recall is 0 / (0 + 10) = 0%. With no predicted positives, precision has a zero denominator; report it as undefined unless the evaluation software explicitly applies a convention.
Now imagine a candidate model on the same set: 970 true negatives, 20 false positives, 4 false negatives and 6 true positives. Its accuracy falls to (970 + 6) / 1,000 = 97.6%, while recall rises to 6 / 10 = 60%. Of its 26 alerts, 6 are correct, giving precision 6 / 26, about 23%. These are arithmetic examples, not benchmark results. They expose the decision that accuracy alone conceals: the candidate catches six events while creating 20 false alarms.
The class prevalence matters because each correct ordinary prediction contributes to accuracy. If ordinary events dominate the set, the majority class can dominate the score as well. A useful first comparison is an always-negative rule. scikit-learn’s DummyClassifier can implement a most_frequent baseline, provided ordinary is the majority label in the training data. It ignores the input features, so a sophisticated model that barely exceeds its accuracy has not yet shown much value on that measure.
For a rare-positive task, record the four counts, the positive prevalence, recall and precision alongside accuracy. Balanced accuracy can also expose an all-negative baseline: it averages recall for the positive and negative classes, giving this baseline 50% in the binary example. Its limitation is that equal class weighting may differ from the actual cost of missed alerts and false alarms. Set the operating threshold and acceptance criteria from those costs, then evaluate on a held-out set with the prevalence the deployment decision needs to represent.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.