A classifier assigns 0.62 to a support message that needs urgent human review. At a 0.70 cutoff, the message is classified as routine; at 0.50, it is flagged. The model’s score did not move. The action did. That distinction is central when an AI triage feature is judged by F1.
A binary classifier often outputs a probability estimate or decision score. A threshold converts that score to a positive or negative label. As the scikit-learn threshold guide explains, changing the threshold after fitting can change labels while leaving the underlying scores and the ranking of examples intact. Its usual probability cutoff of 0.5 is a default decision rule, not an optimum for every application.
Precision is the fraction of predicted positives that truly are positive. Recall is the fraction of actual positives found. F1 is their harmonic mean: 2 × precision × recall / (precision + recall), when the denominator is nonzero. Lowering the threshold usually flags more cases, tending to raise recall and to admit more false positives; precision may fall. Raising it usually does the reverse. These are tendencies rather than guarantees for every discrete threshold step. F1 changes only when one or more examples cross the threshold and the confusion counts change.
Suppose a validation set has 20 genuinely urgent messages. At one cutoff, the classifier flags 10 messages, eight correctly. Precision is 8/10, recall is 8/20, and F1 is about 0.53. At a lower cutoff, it flags 24 messages, 15 correctly. Precision is 15/24, recall is 15/20, and F1 is about 0.68. The second cutoff improves F1 in this invented example, though it also sends nine routine messages to reviewers instead of two. If reviewer capacity is limited, that operational cost deserves attention alongside F1.
Choose the target metric and positive class before searching thresholds. On a validation set separate from model training, inspect precision, recall, and F1 across candidate cutoffs, then select a cutoff consistent with the cost of missed and extra alerts. Scikit-learn provides TunedThresholdClassifierCV for cross-validated threshold selection and warns against tuning a fitted classifier’s threshold on the same observations used to train it. When using a separately trained estimator with cv="prefit", supply fresh validation data for the cutoff search.
Keep a final test set untouched during both model fitting and threshold selection. Evaluate the selected model and cutoff once on that set. Otherwise repeated test-set adjustments turn the reported F1 into another tuning result. Save the chosen cutoff with the classifier: a deployed decision rule needs both the scoring model and the threshold that converts scores into actions.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.