Imagine a fictional set of 1,000 transactions. Ten are labelled fraudulent and 990 are labelled legitimate. A rule predicts “legitimate” for every transaction. Its accuracy is 99%.
A manager proposes using the rule to decide which transactions to investigate because its accuracy is 99%. Explain why that score is insufficient for this task. Name two pieces of decision evidence you would request before choosing a detector; do not assume that a high accuracy score justifies blocking transactions.
The rule gets all 990 legitimate cases right and all ten fraud cases wrong. It finds no labelled fraud, so its high accuracy does not demonstrate the capability the investigation team needs. Useful decision evidence includes the errors by class, what happens after a flag, the cost of each error and the team's review capacity. The arithmetic is sound; the conclusion that it detects fraud well is not. Accuracy measures the proportion of correct predictions. It does not tell you which kinds of prediction were correct. scikit-learn’s accuracy documentation
By the end of this example, you will be able to turn a headline score into a small error table and explain what you would need to know before choosing a model. You need only division and percentages.
Compare with a simple baseline
The always-legitimate rule is a baseline: an easy result that a more complicated model should be compared with. In this dataset, legitimate transactions are the majority class, so predicting that class produces a high score without learning to identify fraud.
Suppose a trained model marks 28 transactions for review. Eight of those are fraudulent; twenty are legitimate. It misses two of the ten fraud cases. The remaining 970 legitimate cases are correctly left unflagged.
| Actual label | Flagged as fraud | Predicted legitimate | Total |
|---|---|---|---|
| Fraud | 8 | 2 | 10 |
| Legitimate | 20 | 970 | 990 |
| Total | 28 | 972 | 1,000 |
This is a confusion matrix. Read the row labels as what the example says actually happened and the columns as what the model predicted. Here, “fraud” is the positive class. A true positive is correctly flagged fraud; a false positive is a legitimate transaction flagged as fraud. A false negative is fraud the model misses.
The trained model's accuracy is (8 + 970) ÷ 1,000 = 97.8%. That is lower than the baseline's 99%, yet it finds eight fraud cases the baseline misses. We still cannot choose it from that fact alone: its twenty false alarms also matter.
Ask two different questions about the flags
Precision asks: of the cases flagged as fraud, how many are actually fraud? Here that is 8 ÷ 28, or about 28.6%. A review team would inspect twenty legitimate transactions alongside eight fraudulent ones. Precision definition
Recall asks: of all the fraud cases, how many did the model find? Here that is 8 ÷ 10, or 80%. Two fraud cases remain unflagged. Recall definition
These percentages have different denominators. Saying “the model is 80% reliable” loses the distinction. A useful report names the metric, the positive class and the counts that produced it.
Balanced accuracy offers another view by averaging recall across the classes. The baseline has 0% recall for fraud and 100% for legitimate transactions, so its balanced accuracy is 50%. This exposes the class it ignores. It still does not express how costly either type of error is for a particular application. Balanced accuracy definition
Connect the errors to the decision
What happens after a flag? If it only invites a person to inspect a case, a false alarm consumes review capacity. If it blocks a transaction, the same false alarm has a different consequence. Your evaluation needs that context.
For a deliberately simplified exercise, assign 100 cost units to each missed fraud case and one unit to each false alarm. The baseline costs 10 × 100 = 1,000 units. The trained model costs 2 × 100 + 20 × 1 = 220 units. Under those invented assumptions, the trained model has the lower cost despite lower accuracy.
Now change the false-alarm cost to 50 units. The trained model costs 1,200 units, so the comparison reverses. These are teaching assumptions, not estimates of real fraud losses. Their purpose is to show why the preferred model depends on the action and its consequences.
If the model produces a score, the cutoff used to turn it into a flag is part of that decision. Evaluate the chosen cutoff using development data and the error tradeoff you care about; do not select it by repeatedly inspecting the final test set.
Check whether the evaluation resembles the future task
A helpful metric cannot repair a misleading test. Suppose a feature records an investigation result that only becomes available after the transaction. It may predict the label very well while being unavailable when a real prediction is needed.
Also keep evaluation data out of fitting decisions. Fitting preprocessing or selecting features using the test data can make the reported score too optimistic. A pipeline can help keep learned transformations within the training process. scikit-learn’s leakage guidance
Choose the split around the intended use. Repeated records from the same customer may need to stay together; predicting later events calls for an appropriate time-based evaluation. A random split is not automatically suitable for dependent or time-ordered observations. Cross-validation guidance
Ten positive cases also give a limited picture: missing one additional fraud case changes recall by ten percentage points. Do not turn this small fictional example into a claim about future reliability.
Try another error table
Keep the same 1,000 transactions and ten fraud cases. A second model flags nine fraud cases and sixteen legitimate cases. It misses one fraud case.
A team can review at most twenty-six flagged transactions from this batch. Every flag must be reviewed; it cannot choose only some of a model's flags. Compare the always-legitimate baseline, the first model and this second model. Use 100 cost units per missed fraud and one per false alarm, with no other costs in this simplified decision.
Calculate the second model's accuracy, precision, recall and cost. Recommend the feasible option with the lowest cost, showing the flag counts and costs that justify your choice. Then assess the claim, “90% recall means nine out of ten flagged transactions are fraud.” Finally, name two kinds of evidence you would need before claiming this choice will work as well on future transactions.
It correctly leaves 974 legitimate cases unflagged. Its accuracy is (9 + 974) ÷ 1,000 = 98.3%; precision is 9 ÷ 25 = 36%; recall is 9 ÷ 10 = 90%. Its simplified cost is 1 × 100 + 16 × 1 = 116 units.
The baseline is feasible with zero flags but costs 1,000 units. The first model costs 220 units but sends twenty-eight flags, exceeding capacity. The second sends twenty-five flags and costs 116 units, so it is the lowest-cost feasible option under the stated assumptions. Its 90% recall is nine of the ten actual fraud cases; only nine of twenty-five flags are fraud, giving 36% precision.
If you used 1,000 as the precision denominator, revisit what precision asks. If you ignored the review limit, include feasibility before comparing costs. Before extending the choice to future transactions, request a representative evaluation without unavailable future features and check realistic error costs and review capacity. With only ten fraud cases here, one more miss changes recall by ten percentage points. These fictional records and costs do not establish future performance.
Keep the error table with your score. It lets another reader see what the model found, what it missed and what your choice depends on.
Continue with What your data can and cannot tell you to examine whether the records support the conclusion you want to draw.