Data Engineering For Ml

Handling Imbalanced and Messy Data: Why Your 99% Accuracy Is a Lie

0 of 27 complete

0%

Contents

Back|Data Engineering For MlHandling Imbalanced and Messy Data: Why Your 99% Accuracy Is a Lie
1/27
65 min left
  1. Home
  2. AI Engineering: Data, RAG and Agents
  3. Data Engineering for ML
  4. Handling Imbalanced and Messy Data: Why Your 99% Accuracy Is a Lie

Where this shows up in interviews

This idea carries a full system design question on its own. Each walks through the full answer.

  • →Design a Fraud Detection System
Prerequisites
Data Labeling and Annotation: The Expensive Bottleneck of Supervised MLrequired
Related Topics
Train/Serve Skew: One Input Computed Two WaysWhy Production BreaksFailing Silently: Which Checks Catch a Broken Input Before the Answers ArriveWhy Production BreaksFine-Tuning vs RAG vs Prompting: Choosing Your ApproachLLM and GenAI OpsParameter-Efficient Fine-Tuning: LoRA and QLoRALLM and GenAI OpsEvaluating LLMs in Production: Grading Answers That Have No Right AnswerLLM and GenAI Ops
Previous lessonData Labeling and Annotation: The Expensive Bottleneck of Supervised MLNext lessonSynthetic Data Generation: When Fake Data Beats No Data

System Design

  • Foundation
  • Intermediate
  • Advanced
  • Capstone

AI Engineering

  • Foundation
  • Data, RAG and Agents
  • Evaluation, LLM Ops and Security

systemdesign.academy

  • Home
  • Glossary
  • Interview prep
  • Reviews
  • About
  • Privacy
  • Terms
1 of 27

The Forecaster Who Never Says Rain

Picture a weather forecaster in a desert town. Every morning he says the same thing: no rain today. It rains on maybe two days a year. So he is right on about 363 days out of 365, more than 99 percent of the time. He has never once warned anyone to bring an umbrella.

Is he a good forecaster? By the count of right answers, he is excellent. By the one thing people need from him, he is useless. The rare day is the whole reason to have a forecaster.

An illustration of a woman in a knitted sweater, sitting in an armchair with a laptop on her lap, beside text. Headed a score that looks perfect, titled 99.827% right, and no fraud caught. Beside her: the model on her screen scores 99.827% accuracy on 85,443 card payments; a model that answers legit to every payment gets exactly that score on this data. Then: 148 of those payments were fraud, and it caught none of them. Beneath: accuracy counts every right answer the same; on rare events, almost all the right answers are free.

Machine learning has the same trap. When one kind of case is rare, a model can score very well by never naming it. In this lesson I measure that on real card payments. A model that calls every payment legit scores 99.827 percent on the test set and catches none of the 148 frauds.

In real work the rare case is often the one that matters: fraud, a rare disease, a broken part on a factory line. This lesson is about how to tell when a score is lying to you, which number to read instead, and what the common fixes really do. It ends with the other half of the title: messy data, such as copied rows and gaps, which can quietly bend every score you read.

Where This Lesson Starts

This is the eighth lesson of the chapter on data engineering for machine learning. Lesson 7, on data labeling, was about where labels come from. This lesson is about what happens when one label is far rarer than the other.

Two later lessons in this chapter measure parts of this story in depth, and I build on them here rather than repeat them. Lesson 12, rebalance or threshold, compared rebalancing the data with simply moving the model's cut-off. On average, a plain model with a well-chosen cut-off beat rebalanced models left at the default. Lesson 13, leakage before the split, measured how much each preparation step leaks when it runs before the data is split.

What this rebuild corrected. An earlier version of this lesson made claims its sources do not support. Each is fixed where it comes up, and the full record is in results/ibl-factcheck.json.

  • It opened with a fraud team's story. It was made up, and so were the confusion matrices and curves that followed it.
  • It said fraud is "roughly 1 transaction in 1000" in a typical payments stream. I found no source. The real public data used here is 0.173 percent.
  • It said ROC-AUC is "flatteringly high" on rare classes. More exactly, it does not move when only the rarity changes. That is a different problem, measured below.
  • It said fitting an imputer on all rows turned a 0.95 score into 0.62 in production. Those numbers were invented, and lesson 13 measured that rescaling, a similar step, barely leaked.
  • It said to start with class weights or SMOTE and tune the threshold last. Lesson 12 measured the opposite order working better.

A flowchart headed the plan of this lesson, titled five questions, in order. Is accuracy hiding the rare class? leads to which number to read instead: ROC-AUC, PR-AUC, precision and recall. That leads to does the test set hold enough rare rows? Then where does the threshold go? and does rebalancing help, and what does it cost? Then what else in the data bends the score: copies, gaps, leaks. Beneath: every answer is measured on one real dataset, card payments, 0.173% fraud.

The plan follows that chart from top to bottom. First the trap itself, then the two numbers people use in place of accuracy and how they behave as the rare class gets rarer. Then the split, the cut-off and rebalancing, each with a real measurement. Last, the messy data: copies, gaps and leaks.

Words for This Lesson

A hand-drawn list headed ten words for this lesson, titled the words the numbers need. Accuracy: the share of all answers that were right. Rare class: the thing you want to catch, when it is a small share of the data. Precision: of the rows the model flagged, the share that really were rare. Recall: of the rows that really were rare, the share the model flagged. Threshold: the cut-off that turns a probability into a yes or no; the default is 0.5. ROC-AUC: the chance a random rare row is ranked above a random common one. PR-AUC: how precise the model stays as it flags more and more of the rare rows. Stratified split: a split that keeps the same share of rare rows on each side. Class weights: making each rare row count more while the model learns. Calibrated: a probability of 0.2 really comes true about 2 times in 10. Beneath: none of these is new maths; each answers a different question.

Accuracy is the share of all answers that were right. The rare class is the thing you want to catch, when it makes up a small share of the data. Here it is fraud.

Precision is, of the rows the model flagged as fraud, the share that really were fraud. Low precision means many false alarms. Recall is, of the rows that really were fraud, the share the model flagged. Low recall means many misses.

Most models do not answer yes or no. They give a probability, a number from 0 to 1. The threshold is the cut-off: a row above it is flagged. scikit-learn uses strictly above; this lesson's lab counts at or above, which differs only for a row exactly on the line. In scikit-learn, a popular Python library for machine learning, the default threshold is 0.5.

ROC-AUC and PR-AUC grade the ranking itself, across every threshold. I explain both properly on their own slide. A stratified split is a split that keeps the same share of rare rows on each side. Class weights make each rare row count more while the model learns. A probability is calibrated when it means what it says: rows given 0.2 turn out rare about 2 times in 10.

The Data, the Model and the Runs

I wrote the lab's design into the docstring of its script, imbalance_demo.py, on 1 October 2026, before it ran. The script is the one you can copy from this lesson. Before writing the design, I loaded the data once to see its columns and count exact copies. I also fitted two models on one split to check the plan could run. Nothing else ran first.

An editorial page in four labelled zones, headed what the lab ran: imbalance_demo.py, designed before it ran, titled one real dataset, one simple model, thirty seeds. The data: OpenML 1597, 284,807 card payments by European cardholders over two days in September 2013; 492 are fraud, 0.173%; 29 columns, V1 to V28 and Amount. The model: StandardScaler, then logistic regression; both fit on the training rows only. The runs: 30 seeds for the main parts, 100 for the split; every result is a mean, with its lowest and highest seed. The five parts: the trap; the same model at four rarities; plain or stratified split; class weights and the probability; copies across the split. Beneath: no timings are printed; the copies with the Time column and three seed-by-seed counts came after the results.

The data is a public set of card payments, published by a research group at the Université Libre de Bruxelles with the payments company Worldline. In its own words: "transactions made by credit cards in September 2013 by european cardholders", over two days, with "492 frauds out of 284,807 transactions". That is 0.173 percent fraud.

The columns are not readable. To protect the cardholders, the publishers turned the original fields into 28 new columns, V1 to V28. They used a method called PCA, which mixes the original columns into new ones. Only the time and the amount were left as they were. OpenML, the site the script downloads from, marks the time column as a row label, so scikit-learn leaves it out.

The model is logistic regression, which is a simple model that draws one straight boundary between the classes and turns each row's distance from it into a probability. Before it, a step called StandardScaler puts every column on the same scale. Both are fitted on the training rows only.

The runs. Each part runs with 30 different seeds. A seed is a starting number for the random choices, so each seed splits the data differently. One split is one draw, and one draw is not a finding, so every number below is a mean with its lowest and highest seed.

The Model That Does Nothing Scores 99.8 Percent

Here is the trap, measured. With seed 0, the test set held 85,443 payments, and 148 of them were fraud.

Two panels headed seed 0, test set of 85,443 payments, 148 fraud, titled almost the same accuracy, very different models. Always legit: 99.827%, caught 0 of 148. The model at 0.5: 99.915%, caught 89 of 148, with 14 false alarms. Beneath: accuracy moved by 0.088 points; fraud caught moved from 0 to 89.

A "model" that answers legit to every payment scored 99.827 percent accuracy and caught nothing. The trained model, at the default threshold of 0.5, scored 99.915 percent. It caught 89 of the 148 frauds and raised 14 false alarms. Over all 30 seeds, the trained model averaged 99.918 percent and always-legit 99.827 percent.

So accuracy moved by less than a tenth of a point, while the number of frauds caught went from 0 to 89. If you only had the accuracy, you could barely tell a working model from one that does nothing. The dataset's own page on Kaggle, a site for sharing data, warns about exactly this: "Confusion matrix accuracy is not meaningful for unbalanced classification."

A table headed seed 0's test set, counted, titled where the 85,443 answers went. Columns: the model said, really fraud, really legit. Always legit, flagged: 0 and 0. Always legit, passed: 148 and 85,295. The model at 0.5, flagged: 89 and 14. The model at 0.5, passed: 59 and 85,281. Beneath: precision 89 of 103 flagged is 0.864; recall 89 of 148 is 0.601.

The table splits every answer by what the model said and what was true. This kind of table is called a confusion matrix. Accuracy only adds up the two "right" cells: flagged fraud and passed legit. Precision and recall look only at the fraud side of it, which is where all the useful information is.

The Trap, Worked by Hand

Formulas are easier to trust once you fill them in yourself. Every number here comes from the table above.

Always legit. It flags nothing. So it gets every legit payment right, 85,295 of them, and every fraud wrong. Accuracy is 85,295 ÷ 85,443 = 0.99827, or 99.827 percent. Recall is 0 ÷ 148 = 0. Precision has no answer, because it flagged nothing: 0 ÷ 0.

The model at 0.5. It flagged 89 + 14 = 103 payments. Right answers are the 89 frauds it caught plus the legit payments it passed: 85,295 − 14 = 85,281. So accuracy is (89 + 85,281) ÷ 85,443 = 85,370 ÷ 85,443 = 0.99915, or 99.915 percent.

Now compare the two by right answers. The model got 85,370 right, and always-legit got 85,295. The difference is 75 answers, and 75 is exactly 89 − 14: it gained 89 by catching frauds and lost 14 by raising false alarms. Out of 85,443, 75 answers is less than a tenth of a percent.

Precision is 89 ÷ 103 = 0.864: about 86 of every 100 flags were real fraud. Recall is 89 ÷ 148 = 0.601: it caught about 60 of every 100 frauds. Those two numbers say what the model does. The accuracy mostly says how rare fraud is.

Two Ways to Grade a Ranking

Precision and recall depend on the threshold. Move the threshold and both change. So people also grade the model's ranking: sort every payment by its probability, highest first, and ask how near the top the frauds landed. Two numbers do this, and they ask different questions.

A hand-drawn sketch headed two questions about the same sorted list, titled ROC counts shares; PR counts what is above the line. On the left, under the words sorted, highest first: a box, above the line: flagged; an arrow across, marked the threshold; and a box, below: passed. On the right, one box: ROC-AUC asks, above the line, what share of all fraud, and what share of all legit. A second box: PR-AUC asks, of the rows above the line, what share are fraud. Beneath: add more legit rows and the ROC shares stay the same; the PR share falls.

ROC-AUC looks at two shares at every threshold: the share of all frauds above the line, and the share of all legit payments above the line. Its area works out to a simple chance. It is the chance that a random fraud is ranked above a random legit payment. A ranking by coin toss scores 0.5, however rare fraud is.

PR-AUC looks at precision at every threshold: of the payments above the line, the share that are fraud. Here it means scikit-learn's average_precision_score. That sums, over each threshold, the recall gained times the precision there. scikit-learn's page says this "is different from computing the area under the precision-recall curve with the trapezoidal rule", which "can be too optimistic". A ranking by coin toss scores about the fraud share, here 0.0017.

The difference matters. ROC-AUC uses shares of each class on its own, so it does not care how many legit payments there are. PR-AUC counts every false alarm above the line against the frauds above it, so the number of legit payments matters a lot. The next slide measures exactly that.

The Same Model, With Fraud Made Rarer

To see how each number reacts to rarity alone, I kept everything else fixed. For each seed, the model was trained once. Then I scored its same probabilities on four test sets. Each kept all the test frauds and a different number of legit payments, so fraud made up 50, 10 or 1 percent, then the natural 0.17 percent.

A line chart headed ibl-demo.json, the same probabilities scored four ways, mean of 30 seeds, titled ROC-AUC stays put; PR-AUC falls as fraud gets rarer. Across four test sets, fraud at 50%, 10%, 1% and 0.17%, on a scale from 0.70 to 1.00. The ROC-AUC line stays flat near 0.975. The PR-AUC line falls from about 0.98 to about 0.76. A dashed line for precision at 0.5 falls from 1.00 to about 0.87. Beneath: ROC-AUC 0.976, 0.975, 0.975, 0.975; PR-AUC 0.981, 0.934, 0.858, 0.755; recall at 0.5 was 0.623 every time.

ROC-AUC barely moved: 0.976 with half the test set fraud, 0.975 at the natural 0.17 percent. PR-AUC fell from 0.981 to 0.934, then 0.858, then 0.755. Precision at 0.5 fell from 1.000 to 0.870. Recall at 0.5 stayed at 0.623 every time.

Two of these were certain before the run, and I want to say so plainly. Recall could not move, because the frauds and their probabilities were the same each time. ROC-AUC should not move, because it is built from shares of each class. What the run adds is the size of the fall in PR-AUC on real data, and the check that ROC-AUC really held. After the results, I counted seed by seed. ROC-AUC moved at most 0.007 across the four test sets of any one seed, and PR-AUC fell in all 30 seeds.

An isometric drawing of four blocks, heights to scale, headed ibl-demo.json, rows in each test set, titled the same 148 fraud rows, more and more legit ones. From left: 296 rows, PR-AUC 0.981, a block so flat it is almost a tile; 1,480 rows, PR-AUC 0.934; 14,800 rows, PR-AUC 0.858; 85,443 rows, PR-AUC 0.755, a tall block. Beneath: every block holds the same 148 fraud rows with the same probabilities; only the legit rows were added.

The heights are the number of payments in each test set, to scale. Every block holds the same 148 frauds. Going from left to right only adds legit payments, about 85,000 of them. Some of those get high probabilities, and each one that lands above the line is a false alarm. PR-AUC counts them. ROC-AUC only sees them as a share of all legit payments, which stays about the same.

Which Number to Read

So which one is right? Both are honest about what they measure. They answer different questions, and the mistake is to read one as if it answered the other.

A two-column table headed what each number counts, titled ROC-AUC and PR-AUC, side by side. Left, ROC-AUC: the share of all fraud above each line, against the share of all legit above it; the chance a random fraud outranks a random legit row; 0.976 at 50% fraud, 0.975 at 0.17%; a coin toss scores 0.5 at any rarity. Right, PR-AUC: of the rows above each line, the share that is fraud; every legit row above the line counts against it; 0.981 at 50% fraud, 0.755 at 0.17%; a coin toss scores the fraud share, 0.0017 here. Beneath, left: compare models on test sets whose fraud share differs. Right: read it with its fraud share.

ROC-AUC tells you how well the model ranks, and it is the same whether you test it on rare or common fraud. That makes it good for comparing models on test sets whose fraud share differs. It is bad at telling you what working life will be like at a low rarity.

PR-AUC tells you how clean the top of the list will be at this rarity. If your team can only check 100 payments a day, that is the question you care about. But PR-AUC only means something next to its fraud share. 0.755 at 0.17 percent fraud is far above the coin-toss score of 0.0017. The same 0.755 at 50 percent fraud would be poor.

A 2024 paper by McDermott and others, at the NeurIPS conference, argues that PR-AUC "is not generally superior in cases of class imbalance". It also shows PR-AUC can "unduly favor model improvements in subpopulations with more frequent positive labels": groups of people where the rare class is more common. That can widen unfair gaps between groups. So I read both numbers, and then precision and recall at the threshold I will really use.

Thirty Seeds, Not One

Every number so far is a mean over 30 seeds. Here is why one seed would not do.

A dot chart headed ibl-demo.json, one dot per seed, titled one split is one draw. For seeds 0 to 29, two rows of dots on a scale from 0.65 to 1.00. PR-AUC with 50% fraud sits in a tight band near 0.98. PR-AUC at the natural 0.17% spreads from about 0.69 to about 0.82. Beneath: natural PR-AUC 0.689 to 0.816 across seeds, mean 0.755; ROC-AUC 0.958 to 0.989.

At the natural rarity, PR-AUC ran from 0.689 on its worst seed to 0.816 on its best. Same data, same model, same code: only the split changed. If I had run one seed and got 0.69, I might have blamed the model. If I had got 0.82, I might have shipped it with too much confidence.

ROC-AUC also moved, from 0.958 to 0.989. Neither spread is a fault in the model. Each test set holds only 148 frauds, and which 148 land there changes the score. With few rare rows, the spread across seeds is part of the answer.

Stratified Splits: What They Fix and What They Do Not

A plain random split can give the test set more or fewer rare rows than its share, just by luck. A stratified split fixes the count. In scikit-learn you pass stratify=y to train_test_split, and the documentation says the data is then "split in a stratified fashion, using this as the class labels".

To see what that buys, I made the problem small on purpose. For each of 100 seeds, I drew 100 frauds and 49,900 legit payments, 0.2 percent fraud. Then I split those same 50,000 rows 80 to 20 twice: once plain, once stratified.

A dot chart headed ibl-demo.json, 100 seeds, 100 fraud in 50,000 rows, titled stratifying fixed the count, not the spread. Horizontal axis: fraud rows in the test set, 8 to 32. Vertical axis: test PR-AUC, 0.2 to 1.0. Dots for the plain split scatter from 10 to 30 fraud rows. Dots for the stratified split sit in one column at 20. Both clouds spread over most of the vertical range. Beneath: plain, PR-AUC sd 0.109, 0.292 to 0.945; stratified, sd 0.106, 0.411 to 0.973.

The plain split gave the test set anywhere from 10 to 30 of the 100 frauds. The stratified split gave it exactly 20 every time. That part worked as promised.

What surprised me is the spread of the score. PR-AUC had a standard deviation of 0.109 with plain splits and 0.106 with stratified ones. The standard deviation (sd) is a measure of how far results typically sit from their mean. So stratifying barely narrowed the spread. Recall's spread fell more, from 0.146 to 0.124.

A hand-sketched column of three boxes joined by arrows, headed why 20 fraud rows is not many, titled each fraud row is a big step. 20 fraud rows in the test set, even when stratified. One fraud row caught or missed moves recall by 1 in 20, 5 points. Recall's sd across 100 stratified seeds: 0.124. Beneath: stratify always, and still count your rare test rows before you trust a score.

The reason is the small count, not the split. With 20 frauds in the test set, each fraud caught or missed moves recall by 1 in 20, which is 5 points. Which 20 frauds land in the test set changes from seed to seed, stratified or not. So stratify always, because it is free and keeps the count steady. But the split alone does not cure a jumpy score. More rare rows narrow it; more seeds and folds show how wide it is and steady the average. Folds are the parts that cross-validation cuts the data into, explained on the next slide.

The Threshold Is a Dial

A probability is not a decision. The threshold is what turns 0.34 into flag or pass. scikit-learn's documentation says a positive class is predicted "when the conditional probability P(y|X) is greater than 0.5". That 0.5 is a default, not a law.

A flowchart headed the playground, seed 0's real test set, titled the same model, three places for the line. A question, where is the line on the model's probability?, branches three ways. 0.5: caught 89 of 148, 14 false alarms. 0.1: caught 113, 26 false alarms. 0.01: caught 124, 123 false alarms. Beneath: accuracy at those lines, 99.915%, 99.929% and 99.828%; it barely tells them apart.

Using the playground further down, I moved the line on seed 0's real test set. These counts came after the results. At 0.5 the model caught 89 frauds with 14 false alarms. At 0.1 it caught 113 with 26. At 0.01 it caught 124 with 123. Accuracy at those three lines was 99.915, 99.929 and 99.828 percent, so it would barely help you choose.

Which line is right depends on cost. If a missed fraud costs a hundred times a false alarm, catching 35 more frauds for 109 more alarms is a good trade. Real fraud systems work this way. Stripe's documentation says its Radar product gives each payment a risk score from 0 to 99, and by default "a score of 75 or above indicates high risk". It also says high-risk payments "are blocked by default", and some plans let a business change that block threshold.

How to choose the line well is the subject of lesson 12. Choose it on a validation set, which is a part of the data kept aside from training for making choices like this one, never on the test set. When the validation set holds only a handful of rare rows, choose it by cross-validation instead. Cross-validation cuts the training data into parts called folds, often 5. It trains on all but one fold, checks on the one left out, and repeats until every fold has had a turn. scikit-learn's TunedThresholdClassifierCV does this. In lesson 12 that did best when rare rows were few.

What Class Weights Do to the Probability

Rebalancing changes what the model learns from, so that the rare class counts for more. The simplest way is class weights. In scikit-learn it is one argument, class_weight="balanced", which the documentation defines as "n_samples / (n_classes * np.bincount(y))". In words: total rows divided by (2 times the rows in that class).

Lesson 12 measured whether rebalancing beats a well-placed threshold. Here I measured something it described but did not count: what weighting does to the probability itself. On each of the 30 splits, I trained the same model twice, plain and weighted.

A bar chart headed ibl-demo.json, mean of 30 seeds, titled the weighted model's probabilities were 44 times too high. Three bars on a scale from 0 to 0.08: the mean probability on the test set. Plain, about 0.0017. Weighted, about 0.076. Corrected, about 0.0024. A dashed line marks the real fraud share, 0.0017. Beneath: plain 0.0017, weighted 0.0761, corrected 0.0024; ranking about the same: ROC-AUC 0.975 and 0.977, PR-AUC 0.755 and 0.729.

The plain model's average probability on the test set was 0.0017, the same as the real fraud share. The weighted model's was 0.0761, about 44 times too high. Its ranking was about the same: ROC-AUC 0.977 against 0.975. PR-AUC was a little lower, 0.729 against 0.755. After the results I counted seed by seed: PR-AUC was lower with weights in 29 of 30 seeds.

The run also prints a Brier score for each model. It is the average squared gap between each probability and what really happened, 1 or 0, so lower is better. It was 0.00070 for the plain model and 0.02298 for the weighted one.

That matches lesson 12, which found rebalancing barely changed the ranking. It also matches the paper this dataset comes from. Dal Pozzolo and others showed in 2015 that, in theory, undersampling, one kind of rebalancing, shifts a model's probabilities without changing their order. They also noted that a model trained on the smaller, undersampled data gives estimates "subjected to higher variance".

Two panels headed ibl-demo.json, at the default threshold 0.5, mean of 30 seeds, titled class weights moved the line, not the model. Plain at 0.5: 13.9 alarms, 92.2 fraud caught. Weighted at 0.5: 2,013 alarms, 134.0 fraud caught. Beneath: false alarms for each fraud caught, plain 0.15, weighted 15.0, counted after the results.

At 0.5, the weighted model caught more frauds, 134.0 against 92.2 on average. It also raised 2,013.0 false alarms against 13.9. That is about 15 false alarms for every fraud it caught, against 0.15 for the plain model. Weighting did not make the model see fraud better. It pushed every probability up, so far more rows crossed the same line. A lower threshold on the plain model flags more rows too, and leaves its probabilities honest.

Putting the Probability Back

If anything downstream reads the probability as a real chance, such as a risk score shown to an analyst, a weighted model gives it the wrong number. You can either not rebalance, or correct the number afterwards.

A sequence diagram with four columns, training rows, class weights, the model and the correction, headed seed 0's training rows; steps 3 and 5 are means of 30 seeds, titled weights in, probabilities out, then put back. Step 1, training rows to class weights: 344 fraud, 199,020 legit. Step 2, class weights to the model: each fraud row counts as much as 578.5 legit rows. Step 3, the model to the correction: q, 0.0761, mean of 30 seeds. Step 4, the correction to itself: p equals b times q over (b times q minus q plus 1), b equals 344 over 199,020. Step 5, the correction to training rows: 0.0024, mean of 30 seeds; real 0.0017. Beneath: the formula is from a paper on undersampling; using it for class weights is my step.

Dal Pozzolo and others give the correction for undersampling, which is throwing away common rows before training. Say q is the model's probability, and b is the share of common rows that were kept. Then the true probability is p = b × q ÷ (b × q − q + 1). Class weights act like keeping a share b of the common rows, with b equal to the rare rows divided by the common rows. Using the formula for weights is my own step, not the paper's.

Worked by hand on seed 0's training rows: 344 frauds and 199,020 legit, so b = 344 ÷ 199,020 = 0.001728. A weighted probability of 0.5 becomes 0.000864 ÷ (0.000864 − 0.5 + 1) = 0.000864 ÷ 0.500864 = 0.0017. So a weighted "50 percent" meant about 0.17 percent.

After the correction, the mean probability was 0.0024, about 1.4 times the real share instead of 44 times. It is closer, not exact. One possible reason is that the model's built-in penalty, called regularisation, acts differently on weighted rows. The correction cannot change the ranking, so its PR-AUC stayed at 0.729. If you need probabilities, the simplest safe choice is still the plain model with a well-placed threshold.

The Rebalancing Tools, and What Each One Does

Class weights are one of several tools. Here is what each one does, from its own documentation, and where it fits after the results above and in lesson 12.

Four cards headed what each tool does, from its own documentation, titled four ways to make the rare class count more. scikit-learn, with its logo: class_weight balanced; each class weighted total rows over (2 times its rows). XGBoost: scale_pos_weight, default 1; a typical value is negatives over positives. imbalanced-learn: SMOTE; a new rare row on the line between a rare row and one of its 5 nearest rare neighbours. torchvision, a PyTorch library, with the PyTorch logo: sigmoid_focal_loss (Lin and others, 2017); turns down the loss on easy, well-classified rows; core PyTorch has none. Beneath: XGBoost and imbalanced-learn have no logo here, because the logo set has none.

XGBoost is a popular library for boosted trees, which are many small decision trees built one after another, each fixing the last one's mistakes. It has scale_pos_weight. Its documentation says the default is 1, and "a typical value to consider: sum(negative instances) / sum(positive instances)". It does the same job as class weights.

Copying rare rows until the classes are even is called random oversampling. It adds no new information. Copies made before the split also land on both sides of it, which lesson 13 measured as a large leak.

SMOTE, the Synthetic Minority Over-sampling Technique, comes from a 2002 paper by Chawla and others. It makes a new rare row on the straight line between a real rare row and one of its nearest rare neighbours. The imbalanced-learn library looks at 5 neighbours by default. Its documentation warns that "SMOTE might connect inliers and outliers". An outlier is a row far from the others of its class, and an inlier is a typical one. For data with category columns there is SMOTENC, which gives each new row "the most frequent category of the nearest neighbors". I did not run SMOTE here; lesson 12 did.

Undersampling throws away common rows. It makes training faster on huge data, and it shifts probabilities exactly as the correction slide described. Focal loss, from a 2017 paper by Lin and others, is used in deep learning. It "down-weights the loss assigned to well-classified examples", and was built for object detectors where most candidate boxes are background. Core PyTorch has no focal loss, but its torchvision library has , with defaults alpha 0.25 and gamma 2.

The Order That Held Up

The earlier version of this lesson said: start with class weights, add SMOTE if needed, and tune the threshold last. Lesson 12 measured that order and found it backwards. A plain model with its threshold tuned beat every rebalanced model left at 0.5 on its digits data. And choosing the threshold by cross-validation did best when rare rows were few.

This lesson adds two measured reasons to agree. Class weights left the ranking about the same here, with PR-AUC lower in 29 of 30 seeds. And they raised the probabilities about 44 times, which breaks anything that reads them.

So the order I now use is this. Report honest numbers: PR-AUC with its fraud share, ROC-AUC, and precision and recall at a stated threshold. Tune the plain model's threshold on validation data or by cross-validation. Then try rebalancing, give it its own tuned threshold, and keep it only if it beats the plain model. If you keep it and anything reads the probability, correct it.

There are cases where rebalancing may still help. Lesson 12 tested logistic regression only, and so did I. Tree models choose their splits by counting rows, so weighting can change what they learn. Test it on your own model rather than trusting a rule.

Messy Data: Copies Across the Split

Imbalance is half of this lesson's title. The other half is messy data. Real data has gaps, typos, mixed codes and copies. This dataset has no missing values, but it does have copies, and copies matter most at the split.

A hand-drawn sketch headed ibl-report.json, with the Time column, counted after the results, titled the same row, more than once. A box: 9,144 extra copies (14,293 rows in 5,149 groups); 19 fraud. An arrow to a box: 1,081 of them also share the same second. Below, a wide box: in each 70/30 split, 3,545.8 test rows on average had an exact copy in training; 7.8 of them fraud. Beneath: in no group of copies did the rows disagree on fraud; most copies were hours apart, median 3.0 hours.

9,144 rows are extra copies of another row in all 29 columns, and 19 of those are fraud. Counted as groups, 14,293 rows sit in 5,149 groups of two or more equal rows. After the results, I read the time column too. Only 1,081 of the copies also share the same second. Inside a group of copies, the rows were a median of 3.0 hours apart, and in no group did they disagree on fraud.

So some of these may be one payment logged twice, and some may be two real payments that look the same. From this data I cannot tell which. What I can measure is the effect on the split. On average, 3,545.8 test rows had an exact copy in training, and 7.8 of them were fraud. A model may have memorised those. So a few of its "test" frauds were not new to it.

The fix depends on what the copies are. If they are the same event logged twice, drop them before the split. If they may be real repeats, keep them but put all copies of a row on the same side of the split. scikit-learn's GroupShuffleSplit does that if you give each group of copies one id, but it does not stratify. When you want both, StratifiedGroupKFold keeps groups apart while it "attempts to return stratified folds". And for data that arrives over time, like fraud, test on later rows than you train on.

Gaps, Typos and Leaks

This dataset had no gaps, so this slide is general advice, checked against the documentation rather than measured. Lesson 5, on data validation, measured which checks catch a bad batch.

Missing values. Most models cannot take an empty cell, so you fill it, a step called imputation. For a number column, the training median is a common fill. The median is the middle value, so one huge typo cannot drag it the way it drags a mean. For a category column, make "Unknown" its own category with SimpleImputer(strategy="constant", fill_value="Unknown"), rather than the most common value. The fact that a value is missing can carry signal. scikit-learn's add_indicator=True adds a column that says so, which the documentation says "allows a predictive estimator to account for missingness despite imputation".

A column that is mostly empty is mostly guesswork once filled. Dropping it above about half empty is a rule of thumb, not a law.

Try It Yourself

This script is the lab. It downloads the card payments, then runs all five parts. They are the trap, the same model at four rarities, the two kinds of split, class weights, and the copies. It needs no GPU. On my laptop it ran in well under a minute, and it prints no timings.

A real screenshot of VS Code with imbalance_demo.py open at the top of the file. The docstring gives what it needs, scikit-learn and pandas, the download of about 70 MB, and how to run it. Then the design, written 2026-10-01 before the first run: the data, OpenML 1597, 284,807 card payments, 492 fraud; the model, StandardScaler then logistic regression, fit on the training rows only; 30 seeds; and the five parts, the trap, rarity with the same model, the split over 100 seeds, rebalancing and the probability with the correction formula, from a paper on undersampling that I applied to weights, and the copies; then a note that this wording changed after a review. Below the docstring, the start of the imports. The rest of it and the code are further down.

Before you run this lab. You need Python 3 and two libraries: pip install scikit-learn pandas. scikit-learn does the download, the split, the model and the scores, and brings NumPy with it. pandas is what scikit-learn uses to read the downloaded table. The first run downloads the data from OpenML (about 70 MB), so it needs an internet connection once. After that, scikit-learn keeps a copy in a folder in your home directory (scikit_learn_data).

I ran it with scikit-learn 1.9.1 on a Mac. These libraries run on Windows and Linux too, but I have not checked the numbers there. Another version of scikit-learn may give slightly different probabilities, so the first line printed is the version. Give it a file name, python imbalance_demo.py out.json, and it also saves every number. That is how results/ibl-demo.json was made. Two runs gave byte-for-byte the same file.

"""Imbalanced data: what accuracy hides, and which number to read.

Lesson 8 of 'Data Engineering for ML', made small. It needs Python 3
with scikit-learn and pandas (pip install scikit-learn pandas). The first
run downloads the credit card fraud data from OpenML (about 70 MB) and
keeps a copy. After that it runs in well under a minute on a laptop.
    python imbalance_demo.py            # print the results
    python imbalance_demo.py out.json   # and save every number

Design, written 2026-10-01 before the first run. Before writing it I
loaded the data once to see its columns and count exact copies, and
fitted two models on one split to check the plan could run.
  Data: OpenML 1597, 284,807 card payments by European cardholders over
  two days in September 2013; 492 are fraud (0.17%). 29 columns: V1 to
  V28 (already turned into PCA components by the publisher) and Amount.
  Model: StandardScaler then LogisticRegression(max_iter=1000), both fit
  on the training rows only. 30 seeds, 0 to 29, for parts 1, 2, 4, 5.
  1 The trap. Stratified 70/30 split. "Always legit" against the model
    at the default threshold 0.5: accuracy, fraud caught, false alarms.
  2 Rarity, same model. From each test set keep all its fraud and draw
    legit rows (numpy, the seed) so fraud is 50%, 10% or 1%, then all
    of them (natural). Score the SAME probabilities each time: ROC-AUC,
    PR-AUC (average precision), and precision and recall at 0.5.
  3 The split. 100 seeds. Draw 100 fraud and 49,900 legit rows (0.2%),
    split 80/20 twice on the same rows: plain, and stratify=y. Report
    the test fraud count and the spread of PR-AUC and recall.
  4 Rebalancing and the probability. On part 1's split, the same model
    with class_weight="balanced". Mean probability on the test set
    against the real fraud share, rows flagged at 0.5, ROC-AUC, PR-AUC,
    Brier score. Then the weighted model's probabilities corrected with
    p = b*q / (b*q - q + 1), b = training fraud rows / legit rows (Dal
    Pozzolo et al. 2015, for undersampling; I applied it to weights).
  5 Copies. Rows identical to another row in all 29 columns, and per
    seed the test rows with an exact copy among the training rows.
  Every result is a mean with its lowest and highest seed. No timings.
Changed after a review, wording only: the note on the correction.

Author: Roni Das
Created: 2026-10-01
"""
import json
import sys

import numpy as np
import sklearn
from sklearn.datasets import fetch_openml
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import (average_precision_score, brier_score_loss,
                             roc_auc_score)
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

SEEDS = range(30)
SPLIT_SEEDS = range(100)
SHARES = [0.5, 0.1, 0.01]          # fraud share of the test set; then all


def model(weights=None):
    return make_pipeline(StandardScaler(), LogisticRegression(
        max_iter=1000, class_weight=weights))


def at_half(y, p):
    """Counts at the default threshold 0.5."""
    flag = p >= 0.5
    caught = int((flag & (y == 1)).sum())
    alarms = int((flag & (y == 0)).sum())
    return {"caught": caught, "alarms": alarms,
            "fraud": int(y.sum()), "rows": len(y),
            "accuracy": float((flag == (y == 1)).mean()),
            "recall": caught / max(1, int(y.sum())),
            "precision": caught / max(1, caught + alarms)}


def scores(y, p):
    return {"roc_auc": float(roc_auc_score(y, p)),
            "pr_auc": float(average_precision_score(y, p)),
            **at_half(y, p)}


def spread(xs):
    xs = np.asarray(xs, float)
    return {"mean": float(xs.mean()), "min": float(xs.min()),
            "max": float(xs.max()), "sd": float(xs.std())}


# ── the data ──
d = fetch_openml(data_id=1597, as_frame=True, parser="auto")
X = d.data.to_numpy(float)
y = (d.target.astype(str) == "1").to_numpy().astype(int)
# row_key: the same number for rows equal in all 29 columns
_, first, row_key = np.unique(X, axis=0, return_index=True,
                              return_inverse=True)
row_key = row_key.ravel()
print(f"scikit-learn {sklearn.__version__}")
print(f"rows {len(y):,}; fraud {y.sum()} ({y.mean():.3%})")
out = {"version": sklearn.__version__, "rows": len(y),
       "fraud": int(y.sum()), "copies": len(y) - len(first),
       "fraud_copies": int(y.sum() - y[first].sum())}
print(f"exact copies of another row: {out['copies']:,}"
      f" ({out['fraud_copies']} fraud)")

runs = []
for seed in SEEDS:
    tr, te = train_test_split(np.arange(len(y)), test_size=0.3,
                              stratify=y, random_state=seed)
    plain = model().fit(X[tr], y[tr])
    weighted = model("balanced").fit(X[tr], y[tr])
    p, q = plain.predict_proba(X[te])[:, 1], weighted.predict_proba(X[te])[:, 1]
    yt = y[te]
    # 2: the same probabilities, on test sets of four fraud shares
    rng = np.random.default_rng(seed)
    fraud, legit = np.flatnonzero(yt == 1), np.flatnonzero(yt == 0)
    rare = {}
    for s in SHARES:
        keep = rng.choice(legit, round(len(fraud) * (1 - s) / s), replace=False)
        idx = np.concatenate([fraud, keep])
        rare[str(s)] = scores(yt[idx], p[idx])
    rare["natural"] = scores(yt, p)
    # 4: the weighted model's probabilities, corrected for the weights
    b = y[tr].sum() / (len(tr) - y[tr].sum())
    fixed = b * q / (b * q - q + 1)
    prob = {}
    for name, v in (("plain", p), ("weighted", q), ("corrected", fixed)):
        prob[name] = {"mean_p": float(v.mean()),
                      "brier": float(brier_score_loss(yt, v)), **scores(yt, v)}
    # 5: test rows with an exact copy among the training rows
    twin = np.isin(row_key[te], row_key[tr])
    runs.append({"seed": seed, "train_fraud": int(y[tr].sum()),
                 "base_accuracy": float(1 - yt.mean()), "rare": rare,
                 "prob": prob, "twins": int(twin.sum()),
                 "fraud_twins": int((twin & (yt == 1)).sum())})
out["runs"] = runs

# 3: plain and stratified splits of the same 50,000 rows
splits = {"plain": [], "stratified": []}
for seed in SPLIT_SEEDS:
    rng = np.random.default_rng(1000 + seed)
    idx = np.concatenate([
        rng.choice(np.flatnonzero(y == 1), 100, replace=False),
        rng.choice(np.flatnonzero(y == 0), 49_900, replace=False)])
    for kind in splits:
        tr, te = train_test_split(idx, test_size=0.2, random_state=seed,
                                  stratify=y[idx] if kind != "plain" else None)
        m = model().fit(X[tr], y[tr])
        r = scores(y[te], m.predict_proba(X[te])[:, 1])
        splits[kind].append({"fraud": r["fraud"], "pr_auc": r["pr_auc"],
                             "recall": r["recall"]})
out["splits"] = splits

# ── what it found ──
R0 = runs[0]["rare"]["natural"]
print("\n1 THE TRAP (seed 0, test set)")
print(f"  test rows {R0['rows']:,}; fraud {R0['fraud']}")
print(f"  always legit: accuracy {runs[0]['base_accuracy']:.3%}, caught 0")
print(f"  model at 0.5: accuracy {R0['accuracy']:.3%}")
print(f"    caught {R0['caught']} of {R0['fraud']}; "
      f"false alarms {R0['alarms']}")
acc = spread([r["rare"]["natural"]["accuracy"] for r in runs])
base = spread([r["base_accuracy"] for r in runs])
print(f"  30 seeds: model {acc['mean']:.3%}, always legit {base['mean']:.3%}")

print("\n2 SAME MODEL, RARER FRAUD (mean of 30 seeds)")
print(f"  {'fraud':>8}{'ROC-AUC':>9}{'PR-AUC':>8}{'prec':>7}{'recall':>8}")
summary = {}
for k in [str(s) for s in SHARES] + ["natural"]:
    summary[k] = {m: spread([r["rare"][k][m] for r in runs])
                  for m in ("roc_auc", "pr_auc", "precision", "recall",
                            "alarms", "rows")}
    share = "0.17%" if k == "natural" else f"{float(k):.0%}"
    s = summary[k]
    print(f"  {share:>8}{s['roc_auc']['mean']:9.3f}{s['pr_auc']['mean']:8.3f}"
          f"{s['precision']['mean']:7.3f}{s['recall']['mean']:8.3f}")
for k in ("0.5", "natural"):
    a, b = summary[k]["roc_auc"], summary[k]["pr_auc"]
    print(f"  {k:>8}: ROC {a['min']:.3f}-{a['max']:.3f},"
          f" PR {b['min']:.3f}-{b['max']:.3f}")
out["rare_summary"] = summary

print("\n3 PLAIN OR STRATIFIED SPLIT (100 seeds, 100 fraud)")
split_summary = {}
for kind, rows in splits.items():
    f = spread([r["fraud"] for r in rows])
    a = spread([r["pr_auc"] for r in rows])
    c = spread([r["recall"] for r in rows])
    split_summary[kind] = {"fraud": f, "pr_auc": a, "recall": c}
    print(f"  {kind}: test fraud {f['min']:.0f} to {f['max']:.0f}")
    print(f"    PR-AUC {a['mean']:.3f}, sd {a['sd']:.3f},"
          f" {a['min']:.3f} to {a['max']:.3f}")
    print(f"    recall sd {c['sd']:.3f}")
out["split_summary"] = split_summary

print("\n4 CLASS WEIGHTS AND THE PROBABILITY (30 seeds)")
print(f"  real fraud share of each test set {1 - base['mean']:.4f}")
prob_summary = {}
for name in ("plain", "weighted", "corrected"):
    s = {m: spread([r["prob"][name][m] for r in runs])
         for m in ("mean_p", "brier", "roc_auc", "pr_auc", "caught",
                   "alarms", "precision", "recall")}
    prob_summary[name] = s
    print(f"  {name}: mean p {s['mean_p']['mean']:.4f},"
          f" Brier {s['brier']['mean']:.5f}")
    print(f"    at 0.5: caught {s['caught']['mean']:.1f},"
          f" alarms {s['alarms']['mean']:.1f}")
    print(f"    ROC-AUC {s['roc_auc']['mean']:.3f},"
          f" PR-AUC {s['pr_auc']['mean']:.3f}")
out["prob_summary"] = prob_summary

print("\n5 COPIES ACROSS THE SPLIT (30 seeds)")
t, ft = spread([r["twins"] for r in runs]), spread([r["fraud_twins"] for r in runs])
print(f"  test rows with a copy in training: {t['mean']:.1f}")
print(f"    ({t['min']:.0f} to {t['max']:.0f}); fraud {ft['mean']:.1f}"
      f" ({ft['min']:.0f} to {ft['max']:.0f})")
out["twin_summary"] = {"rows": t, "fraud": ft}

if len(sys.argv) > 1:
    json.dump(out, open(sys.argv[1], "w"), indent=1)

The Lab Report

A real terminal recording headed python imbalance_report.py, titled every score recomputed with my own code. It opens: credit card fraud, OpenML 1597, 284,807 payments, 492 fraud; every score recomputed with my own code; 2,980 checks agree. Then five numbered sections: the trap for seed 0; the same probabilities at four fraud shares, ROC-AUC 0.976 to 0.975, PR-AUC 0.981 to 0.755; plain or stratified split; class weights and the probability; and the copies with the Time column: 9,144 extra copies in 5,149 groups, 1,081 sharing the same second.

The report lives in scripts/labs/dataeng/imbalance_report.py. It does not trust the demo's code for anything it can redo another way. It reads the downloaded data file itself, keeps the time column, and checks that the other 29 columns match. It fits the same models on the same splits, because the models must be the same. Then it computes every score with its own code: ROC-AUC from ranks, PR-AUC from its formula, the Brier score, the counts at 0.5, and the copies with pandas.

It covers all 30 seeds, all four test sets, all three versions of the probability and all 200 split runs. Every number must come back within one part in a billion. 2,980 checks agreed. It stops at the first one that does not.

Its json mode writes results/ibl-report.json, which the figures read. Its demo mode checks the demo's printed run line by line. Its box mode writes the playground on the next slide and checks it.

What came when: the demo's design, in its docstring, came before any run, and nothing in the demo changed after its first run. Four things came after I saw the results, and the report's docstring lists them. They are the copies with the time column, two sets of seed-by-seed counts, and the false alarms per fraud caught. The thresholds 0.1 and 0.01 on the dial slide came from the playground, also after.

Four logo cards headed the tools, with their logos, titled what ran where. scikit-learn: the download, the splits, the model and its scores. NumPy: the seeds and the random draws. pandas: the report's own read of the data file, and the copies. Python: the playground, no libraries at all.

Move the Line Yourself

This box holds seed 0's real test set, as the plain model scored it. It has each fraud payment's probability, and how many legit payments scored at or above each threshold on a fixed list. It has no model and no payment data. It runs in your browser.

As it is, the box prints the trap. Always legit scores 99.827 percent and catches 0. The model at 0.5 catches 89 of 148 with 14 false alarms. The report checked those counts at every threshold on the list against the real data.

Then try THRESHOLD = 0.1 and THRESHOLD = 0.01, and watch accuracy barely move while the catches and alarms change a lot. Then set LEGIT_KEPT = 0.01, which keeps about 1 legit payment in 100, so fraud becomes about 15 percent of the test set. Recall stays the same and precision jumps: the rarity slide, on one seed. With LEGIT_KEPT below 1, the false alarms are the expected share of the real count, not a new draw.

The Code, Part by Part

The data. fetch_openml(data_id=1597) downloads the payments once and reads the local copy after that. X holds the 29 columns as numbers, and y is 1 for fraud and 0 for legit. np.unique gives every row a number that is the same for exact copies, called row_key.

model. One function builds the pipeline: StandardScaler then LogisticRegression(max_iter=1000), with class_weight either off or "balanced". A pipeline is a chain of steps that is fitted as one, so the scaler only ever learns from the rows the model trains on.

at_half, scores, spread. at_half counts catches, false alarms, accuracy, precision and recall at 0.5. scores adds ROC-AUC and PR-AUC. spread turns a list of 30 or 100 results into a mean, a lowest, a highest and a standard deviation.

How to Handle a Rare Class, Step by Step

A hand-sketched column of six boxes joined by arrows, headed handling a rare class, step by step, titled count, report, split, set the line, then rebalance. 1, count the rare rows in each split before you trust any score. 2, report PR-AUC with its fraud share, ROC-AUC, and precision and recall at a stated line. 3, split with stratify, and keep exact copies on one side of the split. 4, tune the plain model's threshold, by cross-validation if rare rows are few. 5, rebalance only if it beats that, with its own threshold. 6, if anything reads the probability, correct it, or do not rebalance. Beneath: here, class weights raised the mean probability 44 times and left PR-AUC lower in 29 of 30 seeds.

Count the rare rows in each split. Before you trust a score, know how many rare rows it was measured on. Here 148 frauds per test set still left PR-AUC ranging from 0.689 to 0.816 across seeds.

Report the right numbers. PR-AUC with its fraud share, ROC-AUC, and precision and recall at the threshold you will really use. Never accuracy alone.

Split with stratify, and mind the copies. Stratifying is free. Keep exact copies on one side of the split, or drop them if they are the same event twice; StratifiedGroupKFold does both jobs at once. For data that arrives over time, like fraud, test on later rows than you train on.

Tune the plain model's threshold. On validation data, or by cross-validation when rare rows are few. Set it by the cost of each kind of mistake.

Rebalance only if it earns it. Give the rebalanced model its own tuned threshold and keep it only if it beats step 4.

Keep the probabilities honest. If a person or a system reads the number as a chance, correct a rebalanced model's output, or do not rebalance.

When to Rebalance, and When Not To

A two-column table headed grounded in this lesson's numbers and lesson 12, titled rebalance, or leave the data alone? Left, think about rebalancing when: a tree model, whose splits count rows, and you have tested it; the plain model with a tuned line still misses too much; the data is so large that dropping common rows saves real time; nothing reads the probability, or you will correct it. Right, leave it alone when: you need honest probabilities, here weights made them 44 times too high; a tuned threshold already does the job, as in lesson 12; you have few rare rows, where every change is hard to measure; you have not yet tuned the plain model's threshold. Beneath, left: then give it its own threshold. Right: the dial is cheaper.

Think about rebalancing when you use a tree model and have tested the effect, because trees choose splits by counting rows. Also when the plain model with a tuned threshold still misses too much, or when the data is so large that dropping common rows saves real time.

Leave it alone when you need honest probabilities, or when a tuned threshold already does the job. Also leave it when you have so few rare rows that any change is hard to measure. And always leave it alone until you have tuned the plain model's threshold, so you know the score it has to beat.

What These Numbers Can and Cannot Tell You

A two-column page headed read before you trust these numbers, titled what these runs are, and what they are not. They are: one public fraud dataset, two days in 2013; one simple model, logistic regression; 30 seeds, 100 for the split; class weights, measured; copies, counted; the playground, after the results. They are not: not every bank or every year; not a tree model or a neural network; not a law, a spread; not SMOTE, which lesson 12 ran; not proof the copies are errors; not a new draw of the data.

One dataset, two days. Fraud patterns change from bank to bank and year to year. The size of the PR-AUC fall depends on how well the model separates the classes, so another dataset will give other numbers. The direction, ROC-AUC steady and PR-AUC falling, follows from how the two are built.

Random splits, not time. The payments cover two days, and I split them at random. A real fraud system is trained on the past and used on the future, so test on later rows than you train on. I did not measure a split by time, so I cannot say how much these scores would change.

One kind of model. Logistic regression only, with one setting for its built-in penalty. Trees and neural networks may react to class weights differently.

What came after the results. The copies with the time column, the seed-by-seed counts, the false alarms per catch, and the thresholds tried in the playground all came after I saw the results.

The correction was borrowed. The formula comes from a paper on undersampling. Applying it to class weights is my step, and here it got the mean probability within 1.4 times of the real share, not exactly onto it.

What to Do Next

A hand-drawn list headed before the next model ships, titled five questions for your own rare class. Counted?: how many rare rows are in your test set, and in each fold? Reported?: is PR-AUC shown with its rare share, next to ROC-AUC? Line?: who chose the threshold, on what data, and by what cost? Weighted?: was the model rebalanced, and does anything read its probability? Copies?: can the same row land on both sides of your split? Beneath: here, 99.827% accuracy caught no fraud at all.

Take the model your team runs today and ask the first question: how many rare rows are in the test set behind its score? If the honest answer is "a few dozen", treat its score as a range, not a number.

Then check whether it was rebalanced, and whether anything downstream reads its probability as a real chance. Here, class weights made that number 44 times too high.

A closing card headed to keep, titled accuracy counts the free answers. In large type: 99.827%. Beneath: the accuracy of calling every payment legit; it caught none of 148 frauds. Then: the same model, PR-AUC 0.981 at 50% fraud and 0.755 at 0.17%. Last: class weights, mean probability 0.0761 against a real 0.0017.

The card keeps the lesson's three numbers. Calling every payment legit scored 99.827 percent and caught nothing. The same model's PR-AUC was 0.981 or 0.755, depending only on how rare fraud was in the test set. And class weights moved the mean probability from the real 0.0017 to 0.0761.

The next lesson in this chapter is about synthetic data: when made-up rows can help, and how to check that they do.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

On the lesson's test set, calling every payment legit scored 99.827% accuracy. What else was true of it?

Q2

The same probabilities were scored on test sets with 50% and with 0.17% fraud. What happened?

Q3

With 100 frauds in 50,000 rows, what did a stratified split change, compared with a plain one?

Q4

Class weights raised the model's mean probability from 0.0017 to 0.0761. What follows?

And for data that arrives over time, like fraud, test on later rows than you train on.

torchvision.ops.sigmoid_focal_loss

Leaks. A leak is information the model gets in training that it would not have at prediction time. The test set's numbers are one kind. If you fit a scaler or an imputer on all rows before splitting, it learns a little from the test rows. That is wrong, and scikit-learn's guide on common pitfalls says to fit such steps "only" on training data. But lesson 13 measured that rescaling "barely leaked at all". The large leaks were steps that read the label, such as choosing columns or copying rare rows before the split.

The other kind is a column that holds the answer. A made-up example: a column days_until_chargeback in a fraud table only has a value after the fraud is confirmed. Offline it looks like magic; live it is always empty. For every column, ask whether its value exists at the moment you predict.

This is a real run in VS Code's terminal (python imbalance_demo.py).

A real screenshot of VS Code's terminal after running python imbalance_demo.py. It prints scikit-learn 1.9.1; rows 284,807, fraud 492, 0.173%; 9,144 exact copies, 19 fraud. Part 1: test rows 85,443, fraud 148; always legit 99.827%, caught 0; the model at 0.5, 99.915%, caught 89 of 148, false alarms 14. Part 2: a table of ROC-AUC, PR-AUC, precision and recall at fraud shares of 50%, 10%, 1% and 0.17%. Part 3: plain split test fraud 10 to 30, stratified 20 to 20, with PR-AUC means and spreads. Part 4: plain, weighted and corrected, mean probability 0.0017, 0.0761 and 0.0024. Part 5: 3545.8 test rows with a copy in training.

When I ran it, it printed the version, the counts and all five parts. All of it matches the stored ibl-demo.json, and the longest printed line was 50 characters. The report's demo mode checks those lines against the file.

To try something I have not run, change model("balanced") in the main loop to model({0: 1, 1: 10}), a much gentler weight. I cannot tell you what it prints, because I have not run it. The question to ask is how far the mean probability moves, and whether the ranking does.

scikit-learn does almost everything in the demo. NumPy makes the random draws for each seed. pandas reads the data file a second time in the report, so the copies are counted by different code. The playground needs nothing but Python itself.

The main loop. For each of 30 seeds: a stratified 70/30 split, the plain and weighted models, and the test probabilities. The rarity part draws legit rows with a NumPy random generator seeded the same way. The correction applies b * q / (b * q - q + 1). np.isin finds test rows whose row_key is also in training.

The split part. For each of 100 seeds: draw 100 frauds and 49,900 legit rows, split them plain and stratified, and score both.

The printout. Everything is printed from the stored results, and with a file name on the command line json.dump saves every number.