Data Engineering For Ml

Data Labeling and Annotation: The Expensive Bottleneck of Supervised ML

0 of 31 complete

0%

Contents

Back|Data Engineering For MlData Labeling and Annotation: The Expensive Bottleneck of Supervised ML
1/31
89 min left
  1. Home
  2. AI Engineering: Data, RAG and Agents
  3. Data Engineering for ML
  4. Data Labeling and Annotation: The Expensive Bottleneck of Supervised ML
Prerequisites
Data Versioning: Reproducing the Exact Data That Trained a Modelrequired
Related Topics
Fine-Tuning vs RAG vs Prompting: Choosing Your ApproachLLM and GenAI OpsParameter-Efficient Fine-Tuning: LoRA and QLoRALLM and GenAI OpsEvaluating LLMs in Production: Grading Answers That Have No Right AnswerLLM and GenAI OpsPrompt Management and Versioning: Treat Prompts as Production CodeLLM and GenAI OpsVector Databases and Approximate Nearest Neighbor SearchLLM and GenAI Ops
1 of 31
Previous lessonData Versioning: Reproducing the Exact Data That Trained a ModelNext lessonHandling Imbalanced and Messy Data: Why Your 99% Accuracy Is a Lie

System Design

  • Foundation
  • Intermediate
  • Advanced
  • Capstone

AI Engineering

  • Foundation
  • Data, RAG and Agents
  • Evaluation, LLM Ops and Security

systemdesign.academy

  • Home
  • Glossary
  • Interview prep
  • Reviews
  • About
  • Privacy
  • Terms

The New Clerk and the Morning Post

Imagine your first morning at a new office job. The post has arrived: 3,450 letters. Your task is to sort them into two trays, keep and bin. Your manager is busy. She can sit with you for a while and look at 100 letters, and no more.

Which 100 do you show her? If you show her the letters you are already sure about, a bank statement or a pizza flyer, her answers teach you nothing new. If you show her the letters that make you stop and frown, every answer she gives you changes how you sort the next pile.

You also arrive with a few rules of thumb. A letter shouting about free money is usually junk. A letter that uses your boss's first name usually is not. Rules like these sort a lot of letters without asking anyone at all. They are sometimes wrong, and you know it.

An illustration of a woman at a desk sorting white cards into a small box, beside a laptop, a stack of paper and a mug, next to text. Headed a new clerk and the morning post, titled which letters do you ask about? Beside her: a new clerk must sort 3,450 letters into keep and bin; her manager has time to check 100. Beneath: which 100 should she ask about, the letters she is sure of, or the ones she cannot decide? She also knows a few rules of thumb: a letter shouting about free money is usually junk; a letter that uses her boss's first name usually is not. Last: this lesson tries both ideas on 4,601 real emails.

The picture puts the whole lesson on one page. The clerk has far more letters than her manager has time for, so every question she asks has to count. And she already knows things that can sort some letters for free.

A computer that learns from examples has exactly this problem. It needs examples with the right answer attached, and a person has to attach most of those answers. People are slow and cost money, so the answers are usually the most expensive part of the whole project.

In this lesson I take 4,601 real emails, hide their answers, and try four ways of getting them back. I pay for answers picked at random. I pay for the ones the computer is least sure about. I let ten rules of thumb answer for free. Or I pay for all of them. Then I measure how good the computer gets with each one, and I repeat the whole thing 30 times so that luck cannot pick the winner.

Where This Lesson Starts

This is the seventh lesson of the chapter on data engineering for machine learning. The last one, data versioning, was about keeping every version of a dataset so that you can go back to it. This one is about the part of the dataset that costs the most to make: the answers.

Two other lessons measure things this one does not. Wrong labels, measured, lesson 11 of this chapter, makes some answers wrong on purpose and measures what that does to four kinds of model. Inter-annotator agreement, in the evals course, measures how often two graders give the same answer and how much of that is luck. I link to both where they matter and do not repeat their labs.

A left-to-right flowchart headed the same emails, four ways to get their labels, titled pay for some labels, pay for none, or pay for all. A cylinder on the left, 3,450 emails; their labels are hidden until bought, leads to four boxes: buy labels at random; buy the ones the model is least sure of; ten rules, buy none; buy all 3,450. All four lead to one box on the right: the same model, scored on 1,151 emails it never saw. Beneath: thirty times, each with a new random split; the test labels are never bought, they only score.

The flowchart is the plan for the lab. On the left is a pool of 3,450 emails whose answers are hidden until I pay for them. Four arrows leave it: buy answers at random, buy the ones the model is least sure of, let ten rules answer and buy none, or buy all 3,450. All four feed the same model, and that model is always scored on the same 1,151 emails it never saw. Only the way of getting the answers changes, so any difference in the score comes from that.

Here is the order of the lesson. First the words. Then what a label is, why some cost far more than others, and why even trained people disagree. Then who does the labelling, what it costs, and how teams check it. Then the two ideas this lesson measures: rules instead of people, and asking only about the hard cases. Then the lab, its results, and what I found when I looked closer. Last, model pre-labelling, labels that change their meaning, and what real teams have published.

Ten Words for This Lesson

A hand-drawn list headed ten words for this lesson, titled what a label is, and three ways to get one. Label: the right answer attached to an example, spam or not spam. Annotator: a person who adds labels, also called a labeller. Pool: the examples you have, but have not labelled yet. Budget: how many labels you can pay for. Random: buy labels for examples picked at random. Active learning: let the model choose which examples to label next. Uncertainty sampling: the model picks the ones it is least sure of. Labelling rule: a few lines of code that vote a label, or stay silent. Weak supervision: training on labels that rules made, not people. Coverage: the share of examples a rule votes on. Beneath: accuracy, here, the share of test emails the model gets right.

A label is the right answer attached to an example. Here the examples are emails and the label is spam or not spam. Supervised learning is the kind of machine learning that learns from labelled examples, and it is the kind this lesson is about. An annotator, or labeller, is the person who adds the labels.

The pool is every example you have but have not labelled yet. The budget is how many labels you can pay for. Random labelling means buying labels for examples picked at random from the pool, which is what most teams do without thinking about it.

Active learning means letting the model choose which examples a person labels next. The simplest way is uncertainty sampling: the model scores every unlabelled example, and a person labels the ones it is least sure about. For a spam model that gives each email a chance of being spam between 0 and 1, the least sure are the emails nearest 0.5.

A labelling rule, which Snorkel's papers call a labelling function, is a few lines of code that look at one example and vote a label, or stay silent. Weak supervision means training a model on labels that rules made instead of people. A rule's coverage is the share of examples it votes on at all.

Throughout the lab, accuracy is the share of test emails the model gets right. A model that always said not spam would score 0.606 here, because 39.4 percent of the emails are spam.

What a Label Is, and Why Some Cost Far More

A model learns by being shown many examples where the right answer is already known, so that it can find the pattern from input to answer. The input is the raw data. The label is the answer attached to it. Without labels, the same raw data is an unsolved problem, which is why labelling comes before almost every supervised project.

The shape of a label depends on the task, and the shape is what drives the cost:

TaskThe raw inputThe label a person attaches
SentimentA product reviewpositive, negative or neutral
Image classificationA photo"cat", "dog" or "neither"
Object detectionA street photoa box around every car and person
SegmentationA street photothe right class for every pixel
Named entitiesA sentencespans marked as person, company, date
SpeechAn audio clipthe exact words spoken

People Disagree, Even Trained Ones

Labelling is also hard for a reason that has nothing to do with time: people disagree. Show three people the same slightly rude review and you can get three answers. The question is how often, and the honest answer is that it depends on the task.

Three panels headed how often trained people gave the same label; each from its own paper, titled even experts disagree. Lung scans: 34.7%; all four radiologists agreed on 928 of the 2,669 lesions any one of them called a nodule of 3 mm or more. Street photos: 96%; of pixels got the same label from two people, on 30 photos. Chat answers: 72.6%; how often OpenAI's labellers agreed with each other when ranking chat answers. Beneath: LIDC/IDRI, 2011; Cityscapes, 2016; InstructGPT, 2022; how much people agree depends on the task.

The three panels come from three published datasets. In the LIDC/IDRI lung scan collection, four experienced chest radiologists marked every scan in two reads. First each read it blind, on their own. Then each saw the other three's marks, without names, and gave a final opinion. After that second read, 2,669 lesions were called a nodule of 3 mm or more by at least one of them. All four agreed on 928 of those, which is 34.7 percent.

In Cityscapes, two people labelled the same 30 street photos pixel by pixel, and 96 percent of the pixels got the same label. In OpenAI's InstructGPT work, labellers ranked a model's chat answers from best to worst. The paper reports that they agreed with each other 72.6 percent of the time.

Those three numbers are not a ranking of people. Radiologists are not worse at their job than street-photo labellers. The lung task is harder to define: where a shadow stops being a nodule is a judgement call, and a pixel of road is not. So the first thing to measure on any new labelling job is how often your own people agree, before you pay for thousands of labels.

Raw agreement can flatter you. If two people both say "not spam" to 90 percent of the emails, they will agree often just by luck. Cohen's kappa corrects for that luck when there are two labellers, and Fleiss' kappa does the same when there are more. The evals lesson on inter-annotator agreement compares two graders on the same answers and computes Cohen's kappa by hand, so I do not repeat it here.

What a wrong label does to a model depends on the kind of mistake. Lesson 11, wrong labels, measured, found that some models shrugged off random mistakes. A random forest lost only 3.2 points of accuracy with 40 percent of its labels wrong at random. Mistakes with a pattern, where look-alike digits were swapped, did far more damage from 20 percent up.

Who Does the Labelling, and What It Costs

When a team says "we need to label the data", it is choosing among a few kinds of people, and the choice drives cost, speed and quality.

Your own experts. A radiologist labels scans, a lawyer labels contract clauses, a fraud analyst labels payments. The labels are careful and they cost the most, because an expert's hour is expensive and experts have other work. Teams use them when the task needs real expertise or the data is too private to send outside. Experts also build the small, careful set that everything else is checked against.

A crowd. The work is cut into small tasks and posted on a platform such as Amazon Mechanical Turk, where many people each do a few. It is cheap, fast and large, but you do not know who is on the other end or how carefully they read. The crowd on its own is not reliable. The checks you build around it can be, and the next slide is about those checks.

A labelling company. Scale AI describes its business as "high-quality training data, annotations, and RLHF to power the world's most advanced AI models". RLHF is reinforcement learning from human feedback: people compare a model's answers, and the model is trained toward the ones they prefer. Appen says it has "1M+ vetted contributors worldwide". Labelbox sells labelling software and also an expert network it calls Alignerr. These companies supply the people, the tools and their own checks, and you pay them for not having to run all that yourself.

Crowd platforms publish their fees, so this part of the bill can be computed exactly. The ledger below does it for the lab's 3,450 emails.

A two-column ledger headed Mechanical Turk's fee rules, on an example pay I chose: 5 cents an email, titled what asking more people costs. Left, 3 workers an email: 3,450 emails x 3 = 10,350 answers; paid to workers, $517.50; Amazon's fee, 20%, $103.50; total, $621.00. Right, 10 workers an email: 3,450 x 10 = 34,500 answers; paid to workers, $1,725.00; 20%, plus 20% more at 10 or more, $690.00; total, $2,415.00. Beneath: 5 cents is my example, not a market rate; the fee rules are Amazon's own.

Mechanical Turk's pricing page says Amazon takes a "20% fee on the reward and bonus amount (if any) you pay Workers". It also says "HITs with 10 or more assignments will be charged an additional 20% fee". A HIT is one task, and an assignment is one worker doing it. Asking only for Masters adds 5 percent more. Masters are workers Amazon's own systems picked as high performers across many tasks. The pay per email is your own choice, and I used 5 cents only as an example.

Consensus, Hidden Test Items, and the Rulebook

The idea that makes a crowd usable is to trust the system, not any one person. It is the same idea that makes a reliable service out of machines that sometimes fail. Three checks do most of the work.

means asking several people about the same item and taking the majority. If all three say spam, keep the label. If two say spam and one says not, keep the majority, and send it to an expert when the item matters. In 2008, Snow and colleagues tested Mechanical Turk against expert labels on five language tasks. On one of them, judging the emotion in news headlines, an average of 4 non-expert labels per item matched the quality of one expert. They paid US$2.00 for 7,000 of those labels, about three hundredths of a cent each, a 2008 research rate.

Do not copy that rate. Set the pay per item so that a typical worker earns at least a fair hourly wage. Anthropic's 2022 paper, cited later, says its team kept in touch with its crowdworkers "to ensure that they were being compensated fairly". It also says it aimed to pay them "significantly above the minimum wage in California".

Hidden test items, often called gold questions, are items where you already know the right answer. They are mixed into the work, so the labeller cannot tell which ones they are. Someone who gets them wrong is guessing, has misread the rules, or is a script. You drop their answers before they reach the dataset. These items cost almost nothing, and they catch what consensus would quietly average in.

A flowchart headed consensus and hidden test items, for one email, titled trust the system, not any one person. One email, sent to 3 people who cannot see each other's answers, leads to a question: did the person get the hidden test emails right? No: drop that person's answers. Yes: how many of the 3 agree? 3: keep the label. 2: keep the majority; an expert checks the ones that matter; that leads to add the case to the rulebook as a worked example. Beneath: Snow et al., 2008, on one task, an average of 4 non-expert labels per item matched one expert; they paid US$2.00 for 7,000 labels.

The flowchart shows both checks on one email. The email goes to three people who cannot see each other's answers. Anyone who failed the hidden test items is dropped first. If all three remaining answers agree, the label is kept. If only two agree, the majority is kept and an expert checks the cases that matter. Every such split then goes into the rulebook as a worked example, which is the third check.

The rulebook, or annotation guidelines, is the document that says what each label means and how to handle the hard cases. Teams like to skip it and start labelling, because labelling feels like progress. Then they find out, after paying for ten thousand labels, that half the team read "spam" one way and half another. The hard cases are everywhere. Is a newsletter you signed up for years ago spam? Is a sarcastic review positive or negative? Is a car that is 80 percent hidden behind a lorry worth a box?

Weak Supervision: Rules Instead of Hand Labels

Hand labelling has a ceiling: a person has to look at every item. Weak supervision removes that ceiling. You write a handful of rules and let them label the whole pool. Each rule may be wrong sometimes, and may stay silent when it has nothing to say. Then you train the model on those labels.

The idea got its best-known tool from Stanford. Snorkel was published in 2017 by Ratner, Bach, Ehrenberg, Fries, Wu and Ré. Their paper reports a user study: "subject matter experts build models 2.8x faster and increase predictive performance an average 45.5% versus seven hours of hand labeling". That is their study on their tasks, not a promise for yours.

A two-column ledger headed the lab's ten rules, written before the first run, titled five rules vote spam, five vote not spam. Left, vote spam when the email has: the word remove; the word free; 000, as in a large sum; the character $; a run of 40 or more capital letters. Right, vote not spam when it has: george, a first name; hp or hpl; 650, an area code; the word meeting; edu, as in a university address. Beneath: an email a rule does not fire on gets no vote from it; the dataset's page names george and 650.

These are the ten rules the lab uses, and I wrote them into the demo's design before its first run. Five vote spam. They fire when an email has "remove", "free", "000" as in a large sum, the character $, or a run of 40 or more capital letters. Five vote not spam when it has "george", "hp" or "hpl", the area code "650", the word "meeting", or "edu" as in a university address. An email a rule does not fire on gets no vote from that rule.

The not-spam rules need a warning. The dataset's own page says its normal emails "came from filed work and personal e-mails". That, it says, is why "george" and the area code "650" mark normal mail. So those rules know whose inbox this is.

The same page says that a filter for everyone would have to "blind such non-spam indicators or get a very wide collection of non-spam". My rules would not carry over to your inbox. They are the email version of "letters that use my boss's first name are real". I also cannot unlearn what I have read about this dataset before, so these are not the rules a newcomer would write.

How do the votes become one label? The simplest way, and the one the lab uses, is a majority vote. Count the spam votes and the not-spam votes, and take the larger. A tie, or no vote at all, means no label, and that email is left out of training.

Snorkel offers something smarter, its LabelModel. Its documentation, which calls the rules labeling functions, or LFs, says it learns "the labeling functions' conditional probabilities of outputting the true (unobserved) label", with no hand labels. It also says it "uses a conditionally independent label model, in which the LFs are assumed to be conditionally independent given Y".

Active Learning: Ask About the Ones It Cannot Decide

If people must label, the question becomes which items are worth their time. Labelling ten thousand easy emails teaches a model almost nothing, because it already gets them right. Active learning lets the model pick the examples that should teach it the most, like the clerk showing her manager only the letters she cannot decide.

Burr Settles's survey of the field calls uncertainty sampling "perhaps the simplest and most commonly used" way, and credits it to Lewis and Gale in 1994. For a model that gives each item a chance between 0 and 1, it "simply queries the instance whose posterior probability of being positive is nearest 0.5". A posterior probability is just the model's chance after it has seen the item.

With three or more classes there are three common versions. One takes the item whose top class is least likely. One takes the smallest gap between the top two classes. One takes the most spread-out chances. For two classes, the survey notes, all three pick the same item.

A sequence diagram with three columns, the model, the pool and a person, headed one round of uncertainty sampling, as the lab ran it, titled ask only about the emails it cannot decide. Step 1, the model to the pool: score every unlabelled email. Step 2, the pool to the model: a chance of spam for each. Step 3, the model to itself: keep the 10 nearest 0.5. Step 4, the model to a person: label these 10, please. Step 5, the person to the model: ten labels. Step 6, the model to itself: fit again on every label so far. Beneath: the lab starts from the same 20 random labels as random labelling, then runs this round until it has 500 labels.

The sequence diagram is one round of the loop as the lab runs it. The model scores every unlabelled email in the pool, and the pool hands back a chance of spam for each. The model keeps the 10 whose chance is nearest 0.5 and asks a person to label those 10. The person sends back ten labels, and the model is fitted again on every label bought so far. Then the round starts again.

How much does this save? The claims range widely. Lewis and Gale wrote that uncertainty sampling "reduced by as much as 500-fold the amount of training data" to be labelled by hand. That was on their text task. Settles's survey also reports the opposite result. Schein and Ungar "showed that active learning can sometimes require more labeled instances than passive learning", using logistic regression, the model this lab uses. Passive learning means labels bought without the model choosing.

What the Lab Ran

An editorial page in four labelled zones, headed what the lab ran: labeling_demo.py, designed before it ran, titled one inbox, one model, four ways to label. The data: UCI Spambase, 4,601 emails given to the UCI repository in 1999; 1,813 are spam; each email is 57 numbers, how often 48 words and 6 characters appear, and 3 counts of capital letters. The split: 30 repeats; each splits the emails at random into a pool of 3,450 and a test set of 1,151, keeping the share of spam the same in both. The model: log(1 + x) of each number, scaled, then logistic regression, the same for every way of labelling. The four ways: random, buy labels in a random order; uncertain, the same first 20, then 10 at a time nearest 0.5; rules, ten rules, no labels bought; all, every pool label. Beneath: fixed before the run; the paired counts, the histogram, the typical-email test and leaving rules out came after the results.

The data. UCI Spambase is 4,601 emails given to the UCI machine learning repository in 1999, and 1,813 of them, 39.4 percent, are spam. The raw text is not included. Each email is 57 numbers. 48 of them say what percentage of its words is each of 48 chosen words. 6 do the same for 6 chosen characters, and 3 count its capital letters. A rule such as "has the word free" becomes "the free number is above zero".

The split. Each of the 30 repeats splits the emails at random into a pool of 3,450 and a test set of 1,151.

Each repeat has its own random-number seed, s from 0 to 29 in the code, which fixes its split and its random order. That seed is a number, not a set of labels. The share of spam is kept the same in both. The pool plays the unlabelled emails: a strategy sees a pool label only after it pays for it. The test labels are never bought. They only score the model at the end, and every strategy on a repeat is scored on the same 1,151 emails.

The model. Every strategy uses the same model. Each number is replaced by log(1 + x), which pulls in the very large counts. Then the numbers are scaled to a similar size and fed to logistic regression. Logistic regression draws one straight boundary through the 57 numbers and turns the distance from it into a chance of spam. I used scikit-learn's defaults except max_iter, which I raised to 2000 so that it always finishes.

The four ways. Random buys labels in a random order of the pool. Uncertain buys the same first 20 labels as random, its seed set: the first labels bought at random, before the model chooses. Then it buys 10 at a time: the 10 unlabelled emails whose chance of spam is nearest 0.5. It fits the model again after each 10. Rules buys nothing and trains on the ten rules' majority vote. All buys every pool label, which is the most this pool can teach this model. Random and uncertain are measured at every budget from 20 to 500 labels.

Accuracy Against Labels Bought

A line chart headed lbg-demo.json: the mean of 30 repeats at every budget from 20 to 500, titled test accuracy against labels bought. Labels bought runs from 0 to 500 across; test accuracy from 0.80 to 0.96 up. Two lines, random solid and uncertain dashed, both start at 0.831 at 20 labels. Uncertain sits just below random from 30 to 50 labels, crosses above near 60, rises steeply to about 0.93 by 200 and flattens onto the upper dashed line, all 3,450 labels, near 0.938, by about 400. Random rises more slowly to about 0.922 at 500. A lower dashed line, ten rules, no labels, sits near 0.891; random crosses it near 100 labels. Beneath: at 20 labels both are 0.831; uncertain's mean is ahead from 60 labels on, and at 500 it reaches 0.938, the same as all 3,450 labels to three decimals; random reaches 0.922.

The two lines are the mean of 30 repeats at every budget. At 20 labels both are 0.831, because both strategies hold the same 20 random emails. Then they separate. Uncertain's mean falls a little behind at 30, 40 and 50 labels, and from 60 labels on it stays ahead. By 500 labels it reaches 0.938, the same to three decimals as the model given all 3,450 labels. Random, at 500, reaches 0.922.

The two dashed lines are the fixed marks. The top one is all 3,450 labels at 0.938. The lower one is the ten rules with no hand labels at all, 0.891. Random labelling needs about 100 bought labels before its mean passes the rules. Uncertain passes them sooner, but only after its slow start.

A table headed lbg-demo.json and lbg-report.json: the mean of 30, lowest to highest, titled the same numbers, with their spread. Columns: labels, random, uncertain, ahead. 20: 0.831 (0.706-0.903), 0.831 (0.706-0.903), 0 of 30. 50: 0.871 (0.786-0.917), 0.867 (0.774-0.910), 9 of 30. 100: 0.892 (0.847-0.924), 0.905 (0.871-0.932), 17 of 30. 200: 0.901 (0.864-0.925), 0.927 (0.910-0.950), 29 of 30. 300: 0.911 (0.888-0.935), 0.935 (0.917-0.954), 30 of 30. 500: 0.922 (0.897-0.949), 0.938 (0.923-0.954), 29 of 30. Then: all 3,450, 0.938 (0.927-0.954); rules, 0 bought, 0.891 (0.878-0.906). Beneath: ahead, the repeats where uncertain beat random on the same split, counted after the results.

The table shows the spread behind the means, the lowest and highest of the 30 repeats. At 100 labels random ranged from 0.847 to 0.924, and uncertain from 0.871 to 0.932. At 300 labels, uncertain's worst repeat, 0.917, was above random's mean, 0.911. At 500, uncertain ranged from 0.923 to 0.954, which is almost exactly the range of the model given every label, 0.927 to 0.954.

So on this data, with this model, aiming the labels bought the accuracy of the full pool for 500 of its 3,450 labels, about one in seven. Buying at random did not get there by 500. The next slides ask whether that could be luck and where the slow start came from. Then they show a cost of the aimed labels that the accuracy line hides.

Thirty Repeats, Not One

A single random split can tell almost any story. At 20 labels, one repeat scored 0.706 and another 0.903, on the same kind of emails with the same model. If I had run the lab once, the headline would have been whatever that one split said. So I ran it 30 times, and after seeing the means I also counted, repeat by repeat, how often the aimed labels won.

A line chart headed uncertain against random on the same 30 splits; counted after the results, titled behind at 50 labels; ahead almost every time from 200. Labels bought runs from 0 to 500 across; repeats uncertain is ahead, from 0 to 30, up. A dashed line at 15 is labelled half the repeats. The line starts at 0 at 20 labels, jumps to 16 at 30, falls to 9 at 50, climbs past 15 from about 60 and zigzags up to 29 by 200, then stays between 28 and 30 to 500. Beneath: at 20 labels the two are the same 20 emails, so neither is ahead; at 50, uncertain was behind in 21 of 30; at 100, ahead in 17; at 200, 29; at 300, all 30.

The line counts, at each budget, the repeats in which uncertain scored higher than random on the same split. At 20 labels it is zero, because the two are identical. At 50 labels uncertain was behind in 21 of 30 repeats, with a mean gap of -0.004. At 100 it was ahead in 17 and behind in 13, close to a coin toss, even though its mean was 0.013 higher. At 200 it was ahead in 29 of 30, at 300 in all 30, and at 500 in 29.

The slow start has a simple possible reason, which I did not test. With 20 labels the model is poor, so its idea of "the emails nearest 0.5" is poor too. It spends its next labels on emails that only a poor model finds confusing. Random labels at least cover the pool evenly. Once the model is decent, its doubts point at the right emails.

The practical lesson is to never trust one run of a comparison like this. A difference of 0.013 at 100 labels was a mean over 30 repeats that split 17 to 13. On one repeat it could have pointed either way.

How Many Labels to Get Close to All of Them

A second way to read the curves is to fix a target and count the labels needed to reach it. I set the target before the run: the accuracy of the model given all 3,450 labels on the same repeat, minus 0.01. That is one point of accuracy below the best this pool can do.

Two panels headed labels needed to get within 0.01 of all 3,450 labels, on the same repeat, titled uncertain always got there; random mostly did not by 500. Uncertain: 30 of 30; repeats got there, a mean of 188 labels, from 110 to 300. Random: 12 of 30; got there by 500, those 12 took 150 to 480, a mean of 326. Beneath: on the 12 repeats where both got there, random needed 1.97 times as many labels, from 0.83 to 3.54; those are random's luckier repeats.

Uncertain reached the target in all 30 repeats, with a mean of 188 labels and a range from 110 to 300. Random reached it by 500 labels in only 12 of the 30 repeats. In the other 18 it had not got there when the budget ran out, so its true count on those is somewhere above 500. On the 12 where it did, it took from 150 to 480 labels, a mean of 326.

The panel also gives a ratio, and it needs care. On the 12 repeats where both reached the target, random needed 1.97 times as many labels, from 0.83 to 3.54. Those 12 are random's luckier repeats, the ones where it got there early.

The other 18 can be bounded. There random needed at least 510 labels, so its ratio was at least 510 divided by uncertain's count on that repeat. After the review of this lesson, the report computed those 18 bounds: 17 of them are above 1.97. Counting each of the 18 at its bound, the ratio over all 30 repeats is at least 2.46. So "about twice as many" understates it.

Compare that with the published claims. Lewis and Gale reported up to 500 times fewer labels on their text task. Schein and Ungar found cases where active learning needed more. Here it was at least about twice, on one inbox with one model. Your data may land anywhere in that range, which is the reason to measure it on your own.

What Uncertainty Sampling Looked At

To see why aiming works, I looked inside one round. After the results, I took repeat 0 and looked at the model when it had 100 labels. I chose repeat 0 because it is the first, not because of what it showed. It had 3,350 emails left to choose from.

A bar chart headed repeat 0, the model on its first 100 labels; chosen by a rule after the results, titled how sure the model was about the 3,350 emails left. Twenty bars, one for each step of 0.05 in the model's chance of spam, from 0 to 1, on a scale of emails from 0 to 1.8 thousand. The first bar, below 0.05, stands near 1,744; the last, above 0.95, near 1,034; the eighteen bars between are all under about 110, most of them under 30. Beneath: 1,744 emails sat below 0.05 and 1,034 above 0.95; only 70 were within 0.1 of 0.5; the 10 bought next scored 0.491 to 0.505; 6 of them were spam.

The bars count those 3,350 emails by the model's chance of spam, in steps of 0.05. Almost all of them are piled at the two ends: 1,744 below 0.05 and 1,034 above 0.95. In total 2,946 of the 3,350 scored below 0.1 or above 0.9. Only 70 were within 0.1 of 0.5. With only 100 labels, this model was already sure about most of the inbox, and at that point it had 0.920 accuracy on the test set.

The 10 emails it bought next all scored between 0.491 and 0.505. Six of them turned out to be spam. So the model was right to be unsure: its coin-flip emails were close to half spam. Each of those labels settles a real question for the model. A label on one of the 2,778 emails in the two end bars mostly confirms what it already believed.

This also shows where the cost hides. The emails near 0.5 are the confusing ones, and confusing emails are likely to take a person longer to label. The lab prices every label the same, which the next slides come back to.

The Labels It Bought Are Not a Fair Sample

Active learning buys a biased set of labels on purpose. That is how it works, and it has two costs that the accuracy line does not show.

An isometric drawing of three upright blocks, heights to scale, headed lbg-report.json: the share of spam among the labels bought, at 100 labels; heights to scale, titled the labels it bought were not a fair sample. The pool, 0.394, and random's 100, 0.388, are about the same height; uncertain's 100, 0.486, is clearly taller. Beneath: mean of 30 repeats; at 500 labels, random 0.398, uncertain 0.472; emails near 0.5 are about half spam, so that is what it bought.

The first cost is the mix. The three blocks are the share of spam in the pool, 0.394, in random's first 100 labels, 0.388, and in uncertain's first 100 labels, 0.486. Random's labels look like the pool, as a random sample should. Uncertain's do not, because emails near 0.5 are about half spam. At 500 labels the shares were 0.398 and 0.472. Count spam in the aimed labels to estimate how much spam you get, and you would be 9 points too high at 100 labels, 8 at 500.

Two panels headed the model fitted on random's 500 labels, on each repeat; measured after the results, titled the emails it chose were the hard ones. The test set, picked at random: 0.922; accuracy, from 0.897 to 0.949. The emails uncertain bought: 0.698; accuracy, from 0.663 to 0.731, about 411 a repeat, none in random's 500. Beneath: lower on every one of the 30 repeats; scored on the emails active learning picked, the same model looks 22 points worse.

The second cost is harder. I took the model fitted on random's first 500 labels and scored it twice on each repeat. Once on the random test set, and once on the emails uncertain had bought that random had not, about 411 a repeat. On the test set it scored 0.922. On the emails active learning picked, it scored 0.698. It was lower on every one of the 30 repeats. I designed this check after the results, to test a worry, and the worry held.

The lesson is plain. Never use actively chosen labels as your test set, and never report accuracy measured on them. They are the hard emails by design, so they make any model look worse than it is. If you tune a model on them, you tune it for the edge of the problem. Settles's survey says the same in general terms: "the labeled instances are a biased distribution, not drawn i.i.d. from the underlying natural density", where i.i.d. means drawn independently from the same distribution, like fair random draws. Keep a random, carefully labelled test set that no strategy chooses.

Ten Rules, No Hand Labels

The rules strategy bought no labels at all, and it scored 0.891 on average. That is about what random labelling reached with 100 bought labels. To see how, I looked at each rule on all 4,601 emails, against the true labels.

A hand-drawn table headed every rule on all 4,601 emails, checked against the true labels, titled how often each rule voted, and how often it was right. Columns: rule, votes, right, for. A shaded block of five spam rules: remove, 807, 94.7%; free, 1,241, 79.7%; 000, 679, 88.7%; $, 1,400, 79.2%; caps, 1,220, 73.1%. A second shaded block of five not-spam rules: george, 780, 99.0%; hp, 1,143, 95.6%; 650, 463, 93.5%; meeting, 341, 94.1%; edu, 517, 86.8%. A last row, set apart: the vote, 3,443, 92.4%, either. Beneath: the rules for not spam were right more often, george 99.0%, hp 95.6%; the weakest was the capital-letters rule, 73.1%; a team using rules would not have these true labels.

The sketch lists every rule with the number of emails it voted on and the share of those it got right. The not-spam rules were right more often: "george" 99.0 percent of the 780 emails it voted on, "hp" 95.6 percent. The spam rules were weaker, with "free" at 79.7 percent and "$" at 79.2 percent, and the weakest was the capital-letters rule at 73.1 percent. A team using real rules would not have these true labels. It would check each rule on a small hand-labelled sample instead.

A hand-sketched bar split into four parts, headed what the vote did with all 4,601 emails, titled three in four got a label; one in thirteen of those was wrong. From left to right the parts are long, narrow, narrow and medium, with a key beneath in the same order: labelled, and right, 3,181; labelled, but wrong, 262; a tie, no label, 296; no rule voted, no label, 862. Beneath: the bar runs left to right in the same order; 3,443 labelled, 74.8%, 92.4% of them right; the 1,158 ties and silent emails are left out of training.

The bar shows what the majority vote did with all 4,601 emails. It gave a label to 3,443 of them, 74.8 percent, and 3,181 of those, 92.4 percent, were right. 262 got the wrong label. 296 were ties, where as many rules voted spam as not spam, and 862 got no vote at all. On each repeat the model trained on about 2,580 rule-labelled pool emails, with about 8 percent of their labels wrong, and still reached 0.891.

A bar chart headed the rules model fitted again without each rule, mean of 30; measured after the results, titled leaving out any spam rule raised the score; leaving out any not-spam rule lowered it. Ten bars, one for each rule left out, on a scale of change in test accuracy from -0.06 to +0.03, in order: free near +0.021, caps near +0.015, $ near +0.010, 000 near +0.007, remove near zero, 650 near -0.008, meeting near -0.015, george near -0.024, edu near -0.027 and hp near -0.058. Beneath: all ten, 0.891; without free, 0.912; without caps, 0.907; without hp, 0.833; found with test labels, so dropping rules on this basis would be tuning on the test set.

Letting a Model Suggest the Label

The newest way to cut labelling cost is to let a large language model make the first pass. It suggests a label for every item, and a person only confirms or corrects it. Checking a suggestion is usually faster than making a label from nothing, when most suggestions are right.

Tools support this directly. Label Studio's documentation says that "if you have predictions generated for your dataset from a model ... you can import the predictions with your dataset into Label Studio for review and correction". The suggestion appears on the labelling screen, and the person edits it instead of starting blank.

How good are such suggestions? One study gives a number. Gilardi, Alizadeh and Kubli published one in 2023 in PNAS, the Proceedings of the National Academy of Sciences.

It used "four samples of tweets and news articles (n = 6,183)". They found that "the zero-shot accuracy of ChatGPT exceeds that of crowd workers by about 25 percentage points on average". Accuracy there means agreement with the authors' own trained annotators, so ChatGPT beat the crowd, not the experts. The cost was "less than $0.003" a label, about thirty times cheaper than Mechanical Turk. Zero-shot means the model was given instructions but no labelled examples. That is one study, on text, with one model version, so treat it as one data point and not a law.

A flowchart headed a model suggests, a person checks, titled pre-labelling, with the check that keeps it honest. Every unlabelled item leads to: a model suggests a label, and how sure it is; then a person confirms it, or corrects it; then a question: on hidden test items, does the person catch wrong suggestions? Yes: keep the labels. No: the person is only clicking accept, retrain them, or drop their labels. Beneath: Gilardi et al., PNAS, 2023, on 6,183 tweets and news articles, ChatGPT's zero-shot accuracy beat crowd workers by about 25 points on average, at under $0.003 a label; one study, on text.

The flowchart adds the check that keeps this honest. Every unlabelled item gets a suggested label and a confidence. A person confirms or corrects it. Then comes the question that matters: on hidden test items where the suggestion is known to be wrong, does the person catch it? If yes, keep their labels. If not, they are only clicking accept, and their labels are the model's labels with a person's name on them.

Respect this risk. A model has its own blind spots. If you accept its labels without a real check, you copy those blind spots into your training data. There they are much harder to find than one person's slip. A confident wrong suggestion is especially easy to wave through. So keep hidden test items running under the checking step too, exactly as for a crowd.

When the Right Answer Changes

Labels are not bought once. Sometimes the meaning of the right answer changes. A new kind of fraud appears that the old "fraud" label never covered. A content policy changes what counts as abusive. A product adds a category that did not exist when you labelled. The emails stay the same, but the correct answer for some of them is now different.

People sometimes call this label drift. The usual name for it is concept drift: the link from input to right answer has moved. Your model was trained on an answer key for a world that has changed, and it gets quietly worse even though its weights never changed. The chapter's lesson on drift alarms measures when a drift alarm means real harm and when it does not.

The fix is relabelling: send a fresh sample of recent data back through the labelling process, often with an updated rulebook, and retrain on it. Mature teams schedule this the way they schedule retraining, because retraining on stale labels only relearns the old world with more confidence.

This is where the cheaper ways earn their keep a second time. A set labelled by rules can be relabelled by changing a rule and running it again. A set labelled by people has to be paid for again. If you kept a random, carefully labelled test set, you can also tell why the model got worse: the world changed, or the new labels are poor.

How Real Teams Describe It

A hand-sketched timeline of nine boxes down a line, headed dates and numbers from each paper or company page, titled thirty years of trying to buy fewer labels. 1994: Lewis and Gale, uncertainty sampling, up to 500 times fewer labels, on their text task. 2008: Snow et al., 4 Mechanical Turk labels matched one expert, on one task. 2011: LIDC/IDRI, all four radiologists agreed on 34.7% of lesions. 2016: Cityscapes, over 1.5 h to label one street photo finely. 2017: Snorkel, from Stanford, rules instead of hand labels. 2018: Snorkel DryBell at Google, rules matched tens of thousands of hand labels. 2022: InstructGPT, about 40 contractors label for OpenAI. 2023: Gilardi et al., ChatGPT against crowd workers, on tweets. 2025: Scale AI, a labelling company, valued at over $29 billion. Beneath: each claim is the paper's own, on its own data; Schein and Ungar, 2007, also found cases where active learning needed more labels than random.

The timeline puts the sources of this lesson in order, each with the number its own paper reports. In 1994 Lewis and Gale introduced uncertainty sampling. In 2008 Snow and colleagues found that 4 Mechanical Turk labels matched one expert on one task. The LIDC/IDRI paper in 2011 reported that all four radiologists agreed on 34.7 percent of lesions. Cityscapes in 2016 measured over 1.5 hours to label one street photo finely.

Snorkel came from Stanford in 2017, and Snorkel DryBell from Google in 2018. InstructGPT's roughly 40 contractors followed in 2022, and Gilardi's study of ChatGPT against crowd workers in 2023. In June 2025 Scale AI, a labelling company, announced an investment from Meta that "values Scale at over $29 billion".

Every one of those claims is about its own data. The line at the bottom of the figure is there for balance: Schein and Ungar found cases where active learning needed more labels than random.

Four cards with product logos, headed in each company's own paper or page, checked 2026-09-30, titled who labels for the large models. Google, with its logo, Snorkel DryBell, 2018: rules built from resources across the company; classifiers as good as ones from tens of thousands of hand labels. OpenAI, with its logo, InstructGPT, 2022: about 40 contractors, hired on Upwork and through Scale AI; they agreed with each other 72.6% of the time. Anthropic, with its logo, a helpful and harmless assistant, 2022: US-based Mechanical Turk workers with the Masters qualification, and crowdworkers hired on Upwork. Amazon Mechanical Turk, with the Amazon logo: a 20% fee on what you pay workers, and 20% more on tasks with 10 or more workers. Beneath: Scale AI, Snorkel and Label Studio have no card with a logo, because the logo set has none for them.

The cards are four companies in their own words. Google's Snorkel DryBell paper says its rules, built from resources across the company, gave classifiers as good as ones trained on tens of thousands of hand labels. OpenAI's InstructGPT paper says "we hired a team of about 40 contractors on Upwork and through ScaleAI". It adds that "training labelers agree with each-other 72.6 ± 1.5% of the time". Anthropic's 2022 paper on a helpful and harmless assistant says it "invited master-qualified US-based MTurk workers" and "also hired crowdworkers on Upwork". Amazon's own page gives Mechanical Turk's fees.

Try It Yourself

This script is the lab. It downloads Spambase and runs the 30 repeats of random and uncertain labelling, from 20 to 500 labels. It fits the rules model and the full-pool model. Then it prints the accuracy table, the labels needed to reach the target, the spam share of the labels bought, and each rule's coverage and accuracy. It needs no GPU. The first time I ran it, it took about 9 seconds after the download. The time depends on the machine and how busy it is, so the script does not print it.

A real screenshot of VS Code with labeling_demo.py open at the top of the file, showing the first part of its docstring: what the script is, that it needs Python 3 with scikit-learn and pandas, and how to run it; then the design written on 2026-09-30 before the first run. The data, UCI Spambase, 4,601 emails given to the UCI repository in 1999, each one row of 57 numbers. Thirty repeats, each a random split into a pool and a test set of a quarter. One model for every strategy, log(1 + x), scaled, then logistic regression. The four ways to get labels: random, uncertain, rules with the ten rules spelled out, and all. The budgets, 20 to 500, and the start of what is reported. The rest of the design and all the code are further down. Beneath: copy it from the box on the slide.

Before you run this lab. You need Python 3 and two libraries: pip install scikit-learn pandas. scikit-learn holds the model and the download, and brings NumPy with it. pandas holds the table the rules read. The first run downloads Spambase from OpenML (about 130 KB), so it needs an internet connection once. After that, scikit-learn keeps a copy in a folder in your home directory (scikit_learn_data).

I ran it with scikit-learn 1.9.1 and pandas 3.0.6 on a Mac. scikit-learn runs the same way on Windows and Linux, but I have not checked the numbers there. Another version may give different decimals, so the first line printed is the versions. Give it a file name, python labeling_demo.py out.json, and it also saves every number at full precision. That is how results/lbg-demo.json was made.

"""Data labelling: what a budget of hand labels buys, spent three ways.

Lesson 7 of 'Data Engineering for ML', made small. It needs Python 3 with
scikit-learn and pandas (pip install scikit-learn pandas). The first run
downloads UCI Spambase from OpenML (about 130 KB) and keeps a copy.
    python labeling_demo.py            # print the results
    python labeling_demo.py out.json   # and save every number

Design, written 2026-09-30 before the first run:
  Data: UCI Spambase, 4,601 emails given to the UCI repository in 1999.
  Each email is one row of 57 numbers: how often 48 words and 6
  characters appear in it, and three counts of its capital letters.
  class 1 is spam.
  Thirty repeats, seed s = 0 to 29. Each splits the emails at random,
  keeping the spam share, into a pool and a test set of a quarter
  (train_test_split, test_size 0.25, stratify, random_state s). The
  pool plays the unlabelled emails: a strategy sees a pool label only
  after it pays for it. Test labels are never bought; they only score.
  One model for every strategy: log(1 + x) of each number, scaled,
  then logistic regression (scikit-learn defaults, max_iter 2000).
  Four ways to get labels:
    random     buy labels in a random order of the pool
               (numpy default_rng(s).permutation).
    uncertain  the first 20 labels are random's first 20. Then, 10 at
               a time, buy the 10 unlabelled pool emails whose
               predicted chance of spam is nearest 0.5 (uncertainty
               sampling; ties go to the earlier pool row), and refit
               the model after each 10.
    rules      no hand labels. Ten labelling rules, written here
               before the first run, each vote spam, not spam, or
               nothing. Spam if the email has 'remove', 'free', '000'
               or '$' (one rule each), or a run of 40 or more capital
               letters. Not spam if it has 'george', 'hp' or 'hpl'
               (one rule), '650', 'meeting' or 'edu' (one rule each).
               "Has" means its count is above zero. Each pool email
               gets the majority of the votes cast on it; an email
               with no vote, or a tie, gets no label and is left out.
               The model is fitted on the pool emails that got a
               label, with the rules' labels.
    all        every pool email, with its true label: the most this
               pool can teach this model.
  Budgets: 20, 30, 40, ..., 500 hand labels, for random and uncertain.
  Reported, as the mean, lowest and highest over the 30 repeats:
    test accuracy (the share of test emails it gets right) at each
    budget for random and uncertain, and for rules and all;
    the labels each needs to first reach all's accuracy minus 0.01 on
    the same repeat, and how many repeats never reach it by 500; on
    repeats where both reach it, random's labels over uncertain's;
    the spam share of the labels random and uncertain bought, at 100
    and at 500 labels.
  Over all 4,601 emails: each rule's coverage (the share of emails it
  votes on) and its accuracy on those, checked against the true
  labels, which a team using rules would not have; the vote's
  coverage and accuracy. Last, the smallest budget at which random's
  mean accuracy passes the rules' mean accuracy.
  After the first run, one change, to words only: the data line said
  the emails were "collected at Hewlett-Packard". The dataset's page
  does not say that; it cites a Hewlett-Packard internal report.

Author: Roni Das
Created: 2026-09-30
"""
import json
import sys

import numpy as np
import pandas as pd
import sklearn
from sklearn.datasets import fetch_openml
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline, make_pipeline
from sklearn.preprocessing import FunctionTransformer, StandardScaler

REPEATS = 30
FIRST, STEP, MOST = 20, 10, 500
BUDGETS = list(range(FIRST, MOST + 1, STEP))
SHOWN = [20, 50, 100, 200, 300, 500]

raw = fetch_openml(data_id=44, as_frame=True, parser="auto").frame
X = raw.drop(columns="class")
y = (raw["class"].astype(str) == "1").astype(int)


def rule_votes(rows: pd.DataFrame) -> pd.DataFrame:
    """One column per rule: 1 votes spam, 0 votes not spam, -1 no vote."""
    def has(col: str) -> pd.Series:
        return rows[col] > 0
    spam = {"remove": has("word_freq_remove"), "free": has("word_freq_free"),
            "000": has("word_freq_000"), "$": has("char_freq_%24"),
            "caps": rows["capital_run_length_longest"] >= 40}
    ham = {"george": has("word_freq_george"),
           "hp": has("word_freq_hp") | has("word_freq_hpl"),
           "650": has("word_freq_650"), "meeting": has("word_freq_meeting"),
           "edu": has("word_freq_edu")}
    votes = {k: np.where(v, 1, -1) for k, v in spam.items()}
    votes |= {k: np.where(v, 0, -1) for k, v in ham.items()}
    return pd.DataFrame(votes, index=rows.index)


def majority(votes: pd.DataFrame) -> np.ndarray:
    """The majority of the votes cast; -1 for no vote or a tie."""
    n_spam = (votes == 1).sum(axis=1).to_numpy()
    n_ham = (votes == 0).sum(axis=1).to_numpy()
    return np.where(n_spam > n_ham, 1, np.where(n_ham > n_spam, 0, -1))


def model() -> Pipeline:
    return make_pipeline(FunctionTransformer(np.log1p), StandardScaler(),
                         LogisticRegression(max_iter=2000))


def one_repeat(s: int) -> dict:
    Xp, Xt, yp, yt = train_test_split(X, y, test_size=0.25, stratify=y,
                                      random_state=s)
    guess = majority(rule_votes(Xp))
    Xp, Xt, yp, yt = (a.to_numpy() for a in (Xp, Xt, yp, yt))

    def acc(idx: np.ndarray | list) -> float:
        return float(model().fit(Xp[idx], yp[idx]).score(Xt, yt))

    order = np.random.default_rng(s).permutation(len(yp))
    rand = {n: acc(order[:n]) for n in BUDGETS}
    bought = [int(i) for i in order[:FIRST]]
    if len(set(yp[bought])) < 2:
        sys.exit(f"repeat {s}: the first 20 labels hold one class only")
    unc = {}
    while True:
        m = model().fit(Xp[bought], yp[bought])
        unc[len(bought)] = float(m.score(Xt, yt))
        if len(bought) == MOST:
            break
        left = np.setdiff1d(np.arange(len(yp)), bought)
        p = m.predict_proba(Xp[left])[:, 1]
        near = np.argsort(np.abs(p - 0.5), kind="stable")[:STEP]
        bought += [int(i) for i in left[near]]
    keep = guess >= 0
    rules = float(model().fit(Xp[keep], guess[keep]).score(Xt, yt))
    return {"pool": len(yp), "test": len(yt), "all": acc(np.arange(len(yp))),
            "rules": rules, "rules_fitted_on": int(keep.sum()),
            "random": rand, "uncertain": unc, "bought": bought,
            "spam_bought": {k: {n: float(yp[idx[:n]].mean()) for n in (100, MOST)}
                            for k, idx in (("random", order), ("uncertain", bought))}}


def reach(run: dict, kind: str) -> int | None:
    goal = run["all"] - 0.01
    return next((n for n in BUDGETS if run[kind][n] >= goal), None)


def spread(v: list) -> tuple:
    return float(np.mean(v)), float(np.min(v)), float(np.max(v))


def show(v: list, d: int = 3) -> str:
    m, lo, hi = spread(v)
    return f"{m:.{d}f} ({lo:.{d}f}-{hi:.{d}f})"


print(f"scikit-learn {sklearn.__version__}, pandas {pd.__version__}")
print(f"Spambase: {len(y):,} emails, {y.sum():,} spam ({y.mean():.1%})")
runs = [one_repeat(s) for s in range(REPEATS)]
print(f"{REPEATS} repeats; pool {runs[0]['pool']:,}, test {runs[0]['test']:,}")
print("test accuracy: mean of 30 (lowest-highest)")
print(f"{'labels':>6}  {'random':<19}  uncertain")
for n in SHOWN:
    print(f"{n:>6}  {show([r['random'][n] for r in runs])}  "
          f"{show([r['uncertain'][n] for r in runs])}")
print(f"all {runs[0]['pool']:,}  {show([r['all'] for r in runs])}")
print(f"rules, 0    {show([r['rules'] for r in runs])}")

out = {"scikit_learn": sklearn.__version__, "pandas": pd.__version__,
       "emails": len(y), "spam": int(y.sum()), "repeats": REPEATS,
       "budgets": BUDGETS, "runs": runs, "reach": {}, "rules_all": {}}
print("labels to reach all minus 0.01:")
for kind in ("random", "uncertain"):
    got = [reach(r, kind) for r in runs]
    hit = [g for g in got if g is not None]
    out["reach"][kind] = got
    mean = f"{np.mean(hit):.0f} ({min(hit)}-{max(hit)})" if hit else "never"
    print(f"  {kind:<10} {mean}, missed {len(got) - len(hit)}")
both = [(a, b) for a, b in zip(out["reach"]["random"], out["reach"]["uncertain"])
        if a is not None and b is not None]
ratio = [a / b for a, b in both]
out["reach"]["ratio"] = ratio
if ratio:
    print(f"  random over uncertain: {show(ratio, 2)}, {len(both)} repeats")
print("spam share of the labels bought, at 100 and 500:")
for kind in ("random", "uncertain"):
    a = [r["spam_bought"][kind][100] for r in runs]
    b = [r["spam_bought"][kind][MOST] for r in runs]
    print(f"  {kind:<10} {np.mean(a):.3f}  {np.mean(b):.3f}")

votes = rule_votes(X)
print(f"rules on all {len(y):,} emails: voted  right")
for k in votes.columns:
    cast = votes[k] >= 0
    right = float((votes[k][cast] == y[cast]).mean())
    out["rules_all"][k] = {"coverage": float(cast.mean()), "accuracy": right}
    print(f"  {k:<8} {cast.mean():6.1%} {right:6.1%}")
guess = majority(votes)
cast = guess >= 0
out["vote_all"] = {"coverage": float(cast.mean()),
                   "accuracy": float((guess[cast] == y.to_numpy()[cast]).mean())}
print(f"  vote     {cast.mean():6.1%} {out['vote_all']['accuracy']:6.1%}")
rules_mean = np.mean([r["rules"] for r in runs])
passes = next((n for n in BUDGETS
               if np.mean([r["random"][n] for r in runs]) > rules_mean), None)
out["random_passes_rules"] = passes
print(f"random passes the rules at {passes} labels (means)")

if len(sys.argv) > 1:  # a file name was given: save every number too
    json.dump(out, open(sys.argv[1], "w"), indent=1)

The Lab Report

A real terminal recording headed python labeling_report.py, titled every number in this lesson, rebuilt with plain loops. It opens: Spambase, 4,601 emails, 1,813 spam, 30 random splits; every number rebuilt with plain loops, 3,228 checks, all agree. Then six numbered sections. 1, test accuracy, mean of 30 with the lowest and highest, for random and uncertain at 20, 50, 100, 200, 300 and 500 labels, with how many repeats uncertain was ahead, 0, 9, 17, 29, 30 and 29 of 30; all 3,450 pool labels 0.938, the ten rules 0.891, fitted on 2,580 emails. 2, labels to reach all minus 0.01: random in 12 of 30, mean 326; uncertain in 30 of 30, mean 188; the ratio 1.97 on 12 repeats; the spam shares bought. 3, uncertain minus random on the same repeat, after the results: at 50 labels -0.004, ahead 9, behind 21; at 300 +0.024, ahead 30. 4, what uncertainty sampling saw on repeat 0 at 100 labels: 2,946 of 3,350 emails below 0.1 or above 0.9, 70 within 0.1 of 0.5, the next 10 at 0.49 to 0.50, 6 of them spam. 5, the model on random's 500 labels: test set 0.922, the emails uncertain bought 0.698, lower on every repeat. 6, each rule's votes and share right, and the test accuracy, coverage and accuracy without it, from remove to edu; one email by a rule fixed first, row 2, five spam votes and one not-spam vote, the vote says spam and it is spam; with all five spam rules out, the vote labels 2,122 emails, none of them spam, and the model could not be fitted on any of the 30 repeats; counting random's 18 misses as 510 labels, 17 of the 18 bounds exceed 1.97 and the ratio over all 30 is at least 2.46; all ten label 3,443 emails, 74.8%, 92.4% right, 862 get no vote and 296 a tie. Beneath: the lab's own report; it rebuilds every number another way and stops unless every stored one comes back.

The report lives in scripts/labs/dataeng/labeling_report.py. It reads the demo's stored files, results/lbg-demo.json and the printed run, and Spambase from scikit-learn's local copy. It does not trust the demo's arithmetic. The demo works with pandas tables and NumPy arrays. The report holds the 4,601 emails as plain Python lists and takes log(1 + x) with Python's own math.log1p. It runs the ten rules and the vote with loops, and picks uncertainty sampling's next 10 by sorting in plain Python. It counts every accuracy itself from the model's answers.

It uses the same split, the same random order and the same model as the demo, because those are the design. It stops unless every stored accuracy comes back to within 0.000000000001, every email uncertain bought is the same email, and every rule count is exact. It ran 3,228 checks and all of them agreed. It changes nothing in the demo's files.

Be the Label Model Yourself

This box has no model in it. It holds all 4,601 emails in the dataset's own order. For each one it stores whether the email has each of 23 entries, which are 21 words plus the characters $ and !. It also stores its longest run of capital letters, and whether it is spam. Each email is five characters in base 64, which just means counting with 64 different digits instead of 10, so that the box stays small. It runs in your browser.

The ten rules in it are the lab's rules, exactly as the demo ran them. The report checked that every rule's count, and the vote's, match the demo's stored numbers.

As it is, the box prints each rule's votes and the share it got right. Then it prints the vote: 3,443 emails labelled, 74.8 percent, and 92.4 percent of those right. Those are the numbers on the rules slide.

Now change the rules. Delete the 'free' line and run it again, and watch the vote's accuracy and coverage move. Add a rule of your own, such as 'money': (has('money'), SPAM), or 're': (has('re'), NOT_SPAM),, using any of the 23 entries in WORDS. Change the 40 in the capital-letters rule to 20 or to 60.

One warning. Every time you judge a rule here you are using the true labels, which a real team would not have. Ask yourself which of your changes you could have made without them.

The Vote and the Picks, Worked by Hand

Before you trust a labelling rule, you should be able to apply it with a pencil. So here is one real email, chosen by a rule I fixed before looking at any email. It is the first one, in the dataset's own order, on which the rules vote both ways. That is row 2, the third email.

Row 2 has the word "remove", the word "free", "000" and the character $, and its longest run of capital letters is 485 letters. Each of those five fires a spam rule, so it gets five spam votes. It also has "edu", which fires one not-spam rule. None of the other four not-spam rules fire. Five spam votes against one not-spam vote is a majority for spam, so the vote labels it spam. The true label is spam, so the vote was right.

Now the vote over all 4,601 emails. 862 emails got no vote at all. 296 got as many spam votes as not-spam votes, a tie. That leaves 4,601 minus 862 minus 296, which is 3,443 labelled emails. Their coverage is 3,443 divided by 4,601, which is 0.748. Of the 3,443, 3,181 got the right label, and 3,181 divided by 3,443 is 0.924.

Next, one pick by uncertainty sampling. On repeat 0, with 100 labels, the first email the model bought scored a chance of spam of 0.4993. Its distance from 0.5 is 0.5 minus 0.4993, which is 0.0007. No other unlabelled email was closer. The model sorts all 3,350 unlabelled emails by that distance and buys the 10 smallest.

Last, the bill on the cost slide. 3,450 emails times 3 workers is 10,350 answers. At 5 cents each that is $517.50. Amazon's 20 percent fee is 517.50 times 0.2, which is $103.50, so the total is $621.00. With 10 workers, 34,500 answers cost $1,725.00, the fee is 40 percent, $690.00, and the total is $2,415.00.

The Code, Part by Part

Loading. fetch_openml(data_id=44) downloads Spambase once and reads the local copy after that. X is the 57 numbers of each email, and y is 1 for spam and 0 for not spam.

rule_votes(rows). The ten rules. has(col) is true when a column is above zero. Each spam rule becomes a column that holds 1 where it fires and -1 where it is silent, and each not-spam rule holds 0 where it fires. majority(votes) counts the 1s and the 0s in each row and returns the larger, or -1 for a tie or no vote.

model(). The pipeline every strategy shares: FunctionTransformer(np.log1p) for log(1 + x), StandardScaler() to put the numbers on a similar scale, and LogisticRegression(max_iter=2000).

one_repeat(s). One split, train_test_split(..., test_size=0.25, stratify=y, random_state=s). order is a random order of the pool from default_rng(s). fits the model on the first labels of that order for every budget. The loop is uncertainty sampling. It fits on what has been bought and scores the unlabelled emails with . It sorts them by distance from 0.5, with a stable sort, and buys the first 10. Then fits on the pool emails the vote labelled, and on the whole pool.

How to Spend a Labelling Budget, Step by Step

A hand-sketched column of seven boxes joined by arrows, headed spending a labelling budget, step by step, titled rulebook, random test set, then aim the rest. 1, write the rulebook, with a worked example for each hard case. 2, have two people label a small random sample; count how often they agree. 3, set aside a random test set, labelled carefully, that no strategy chooses. 4, write rules for what you already know; check each on a small labelled sample. 5, start active learning from random labels, then buy the unsure ones in rounds. 6, compare it with random labelling on the same budget before you trust it. 7, relabel a fresh sample when the meaning of a label changes. Beneath: here, step 6 showed uncertain behind random at 50 labels in 21 of 30 repeats, and ahead at 200 in 29.

Write the rulebook first. Give every label a definition and every hard case a worked example. It is the cheapest thing on this list, and every later step depends on it.

Measure agreement on a small random sample. Have two people label the same items and count how often they agree. If they rarely agree, fix the rulebook before you buy more labels. The evals lesson shows how to correct that count for luck.

Set aside a random test set. Label it carefully, perhaps with an expert, and never let any strategy choose its items. Here, the emails active learning chose scored 0.698 against 0.922 on a random test set, so a chosen test set would have lied.

Write rules for what you already know. Check each one on a small hand-labelled sample, not on the test set. Here, ten rules with no labels scored 0.891. Leaving out any one spam rule raised the score, and leaving out any one not-spam rule lowered it.

Start active learning from random labels. Buy a small seed set first, the first labels bought at random, then buy the unsure ones in rounds, fitting the model again between rounds.

Compare it with random on the same budget before you trust it. Here, at 50 labels, the aimed labels were behind in 21 of 30 repeats. From 200 labels they were ahead in 29 or 30.

Relabel a fresh sample when a label's meaning changes. Schedule it, like retraining.

When to Aim the Labels, and When Not To

A two-column ledger headed grounded in this lesson's numbers and sources, titled aim the labels, or buy them at random? Left, use active learning when: labels are expensive and the pool is large; the model gives a usable chance for each item; you can label in rounds and fit again between them; you keep a separate random test set. Right, use random or rules when: you have very few labels, at 50 uncertain was behind in 21 of 30; you need a fair sample, to test or to count a share; you already know strong signs, ten rules gave 0.891 with no labels; the labels must also serve a different kind of model later. Beneath: aim the paid labels; keep a random sample anyway.

Use active learning when labels are expensive and the pool is large. Here it matched the full pool's accuracy with 500 of 3,450 labels. With cheap labels or a small pool, the saving may not be worth the rounds.

Use it when the model gives a usable chance for each item, and you can label in rounds. Uncertainty sampling needs a score to sort by, and it needs to fit the model again between rounds, so the labellers and the model have to take turns.

Use it only with a separate random test set. The labels it buys are the hard ones, so they cannot tell you how good the model is.

Buy at random when you have very few labels. Here, at 50 labels, random was ahead in 21 of 30 repeats. A seed set of random labels helps, but expect a slow start anyway: this lab began with a seed set of 20 and was still behind at 50.

Buy at random when you need a fair sample, to estimate how common something is or to test a model. Uncertain's labels were 48.6 percent spam where the pool was 39.4.

Use rules when you already know strong signs. Ten rules gave 0.891 with no labels here. Check each rule on a small labelled sample, and expect some to hurt.

Be careful when the labels must serve a different model later. Settles's survey warns that a set built by active learning is "inherently tied to the model that was used to generate it". A set bought at random has no such tie.

Common Mistakes When Buying Labels

Measuring the model on the labels active learning bought. They are the hard cases on purpose. Here a model scored 0.698 on them and 0.922 on a random test set.

Judging a strategy on one run. At 20 labels, two repeats of the same lab scored 0.706 and 0.903. At 100 labels, a mean advantage of 0.013 split 17 repeats to 13.

Starting active learning with almost nothing. A model with 20 labels is a poor judge of what is confusing. Start from a seed set of random labels and expect a slow first stretch.

Choosing rules by their test-set score. That turns the test set into training data. Here, dropping "free" raised the test score from 0.891 to 0.912, and I did not drop it for that reason.

Using rules that only work on one inbox. "george" was right 99.0 percent of the time because this dataset's normal emails were its donors' filed work and personal email. The dataset's own page says so.

Accepting model suggestions without a check. Keep hidden test items under the checking step, or a person clicking accept becomes the model with a person's name.

Labelling into a spreadsheet. A labelling tool keeps who labelled what, when, and how long it took. Label Studio's export has all three.

What I Corrected on This Page

This page replaced an older version of the same lesson. I checked every claim in the old version against its source, and these are the ones that did not hold. The full list, with each source and quote, is in scripts/labs/dataeng/results/lbg-factcheck.json.

  • "15 percent wrong labels cannot be saved by any tuning." Lesson 11 measured otherwise for random mistakes, so I took it out.

  • "Snorkel's label model learns how the rules are correlated." Its current documentation says it treats them as independent given the true label. Fixed here and in the diagram on the weak-supervision slide.

  • "Teams routinely need 2 to 10 times fewer labels with active learning." I found no source for that range, so I cut it and measured one case instead.

  • A table of prices per label for experts, vendors and crowds. I found no source for any of the ranges, so I cut it and showed Mechanical Turk's real fee rules instead.

  • "5x to 50x cheaper", "a transcript takes four times the clip" and "30 seconds to label, 10 to verify". No source found; cut.

  • A story about Tesla moving from large labelling teams to auto-labelling. I found only press reports of a 2021 talk, not the company's own words, so I cut it.

  • "Armies of human annotators" behind chat models. The InstructGPT paper names about 40 contractors, so I cut the phrase.

What These Numbers Can and Cannot Tell You

A two-column page headed read before you trust these numbers, titled what these runs are, and what they are not. They are: one inbox, given to UCI in 1999; one model, logistic regression; 30 random splits of it; every label costs the same here; rules by someone who read its page; pairs, histogram, leaving out, after. They are not: not every task or kind of data; trees or text models may differ; not new mail from another year; hard emails take longer to label; not rules from a newcomer; chosen knowing the results.

One inbox. Everything here is one collection of 4,601 emails, given to UCI in 1999, whose normal mail came from its donors' own work and personal email. Other data, other tasks and a newer inbox could give very different numbers.

One model. Logistic regression on 57 counts. A tree model, or a model that reads the raw text, could gain more or less from aimed labels.

Thirty random splits of the same emails. The repeats test how much the result depends on which emails land in the pool and the test set. They do not test new emails from another year.

Every label costs the same here. In real work the emails near 0.5 are likely to take longer to label. Settles's survey notes that "simply reducing the number of labeled instances does not necessarily guarantee a reduction in overall labeling cost". If your tool records time per item, as Label Studio's lead_time does, measure it.

Rules by someone who has read about the dataset. My ten rules were written before the first run, but not by a newcomer, and three of them rely on the dataset's own note about its donors.

Some parts came after the results. The design, the four ways, the rules, the budgets and the target were fixed before the first run. The paired counts, the look inside repeat 0, the typical-email test and leaving rules out came after. I wrote them into the report's design before the report first ran. The worked email was added while I wrote the lesson.

What to Do Next

A hand-drawn list headed before you buy the next labels, titled five questions for your own labelling. Random test?: is there a test set no strategy chose, labelled with care? Agree?: how often do two of your labellers give the same label? Rules?: what do you already know that a rule could say? Compared?: has anyone compared aimed labels with random ones, on the same budget? Changed?: when did the meaning of a label last change? Beneath: here, ten rules gave 0.891 before a single label was bought.

Take one labelling job your team runs and ask the five questions on the card. The first is the most important. Is there a test set that no strategy chose, labelled with care? If not, every number you have about your model might be measured on the wrong emails.

Then do the measurement. Take a pool you already have labels for, hide them, and replay random labelling against uncertainty sampling on the same budget, many times. Write a few rules and check them on a small labelled sample. Here, the aimed labels needed 188 on average to get within 0.01 of the full pool. Random did not get there by 500 in 18 of 30 repeats, and ten rules gave 0.891 for nothing. Your numbers will differ, and now you have a way to find them.

A closing card headed to keep, titled aim the labels; keep a random sample. In large type: 188 labels. Beneath: on average, for uncertainty sampling to get within 0.01 of all 3,450; random got there by 500 in 12 of 30 repeats. Then: ten rules, no labels, 0.891; at 50 labels, uncertain was behind random in 21 of 30. Last: one inbox, one model, a way to measure, not a law.

The card keeps the lesson's numbers. Uncertainty sampling needed 188 labels on average to get within 0.01 of all 3,450, and did so in every repeat. Random labelling got there by 500 labels in 12 of 30. Ten rules and no labels scored 0.891. At 50 labels, aiming was behind random in 21 of 30 repeats.

The next lesson in this chapter is about handling imbalanced data, for when one label is far rarer than the other.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

In the lab, at 50 labels, how did uncertainty sampling do against random labelling on the same 30 splits?

Q2

Why does the lesson say never to measure a model on the labels active learning bought?

Q3

Leaving out the rule for 'free' raised the rules model from 0.891 to 0.912. What does the lesson say to do with that?

Q4

What does Snorkel's current LabelModel documentation say it assumes about the labelling rules?

The first row is one click. The rows further down can take minutes or hours for one example, because every box or every pixel is its own small decision. Two papers measured this carefully, and their numbers are on the scale below.

A vertical time scale with hand-sketched arrows, headed time for one label, from the papers; a log scale, so equal steps are equal ratios, titled the task sets the price of a label. Down the left: 1 s, 10 s, 1 min, 10 min and 1 h. Four arrows point at the scale: at 7 seconds, a box, by extreme clicking, 7 s, ICCV 2017; at about 35 seconds, a box, drawn the usual way, about 5 times as long; at 7 minutes, a street photo, labelled roughly, given under 7 min, Cityscapes; at 1.5 hours, the same kind of photo, labelled finely, over 1.5 h. Beneath: a spam or not-spam label is one click; the papers put a drawn box at seconds and a finely labelled street photo at more than an hour and a half.

Read the scale from the top. In 2017, Papadopoulos and colleagues drew a box around an object by clicking its four outermost points. They measured 7 seconds a box, "5x faster than the traditional way of drawing boxes". So the usual way takes about 35 seconds a box, and a street photo can hold dozens of boxes. The Cityscapes team labelled 5,000 street photos pixel by pixel in 2016. They wrote that "annotation and quality control required more than 1.5 h on average for a single image". For a rougher version of the same kind of label, each photo was given under 7 minutes.

So the whole story of labelling cost fits in one picture. The task picks the price, and the product picks the task. A spam label costs seconds, which is why the lab in this lesson can pretend to buy thousands of them. A medical scan or a street photo costs so much that every label you do not need is money you keep.

Asking 3 workers about each of 3,450 emails is 10,350 answers: $517.50 to the workers, $103.50 to Amazon, $621.00 in all. Asking 10 workers crosses the 10-assignment line, so the fee doubles to 40 percent and the total is $2,415.00. More opinions per email cost more than in proportion.

Whoever does the work, they should work in a real labelling tool, not a spreadsheet. Label Studio is an open-source one, under the Apache 2.0 licence, for audio, text, images, video and time series. When you export its labels, each one carries completed_by, the id of the person who made it, and its timestamps. It also carries lead_time, which its documentation describes as "time in seconds to label the task". That is an audit trail a spreadsheet does not keep, and lead_time tells you which items are slow, which matters later in this lesson.

The cheap way to write a good rulebook is a loop. Write a first version. Have two people label a small sample. Measure how often they agree. Read every item where they disagreed, and turn each one into a written rule with an example. Repeat until agreement stops rising, and only then scale up. My own rule of thumb is that if agreement barely rises after two rounds, the task itself needs redefining, not more labellers. That is a habit, not a measurement.

In plain words, it estimates how accurate each rule is from where the rules agree and disagree. It also assumes the rules do not copy each other once the true answer is fixed.

Snorkel's own spam tutorial then drops the rows no rule labelled. scikit-learn does not accept probabilities as labels, so the tutorial turns each one into a hard label with probs_to_preds. The lab does the same thing with its majority vote.

At Google, Snorkel DryBell built its rules from "existing knowledge resources from across an organization". In the paper's words, it "creates classifiers of comparable quality to ones trained with tens of thousands of hand-labeled examples". That paper covered three classification tasks.

The design, including the ten rules, was written into the demo's docstring before the first run. After the results I added four more measurements to the lab report. I counted the repeats where uncertain beat random, and looked at what the model saw on one repeat. I tested whether the bought emails were typical, and left each rule out in turn. The figures that show those say "after the results".

After the results I asked which rules mattered, by leaving each one out and fitting the rules model again on all 30 repeats. The pattern was clean. Leaving out any spam rule raised the score above the 0.891 of all ten. Without "free" it was 0.912, and without the capital-letters rule 0.907. Without "$" it was 0.901, without "000" 0.898, and without "remove" 0.892. Leaving out any not-spam rule lowered it, most of all "hp", to 0.833, then "edu", "george", "meeting" and "650".

It is tempting to drop "free" at once. Do not, at least not on this evidence. I found it using the test labels, so choosing rules this way would be tuning on the test set. The 0.912 would then no longer be an honest score. The fair way is to check each rule on a separate hand-labelled sample and decide there. The deeper point stands, though: a weak rule can do more harm than good, and the only way to know is to measure.

There is a second risk: where the data goes. Sending items to a third-party model's API, or to a crowd, shows them to someone outside your company. Private data, such as medical records or customer messages, stays inside. Give it to your own experts, as on the slide about who does the labelling, or to a model you run yourself.

The pattern across all of them is the one this lesson measured. The largest models in the world still stand on labels made by people and checked by other people. Rules and models are used to make every paid hour go further.

This is a real run in VS Code's terminal (python labeling_demo.py).

A real screenshot of VS Code's terminal after running python labeling_demo.py. It prints scikit-learn 1.9.1, pandas 3.0.6; Spambase, 4,601 emails, 1,813 spam, 39.4%; 30 repeats, pool 3,450, test 1,151; then test accuracy as the mean of 30 with the lowest and highest, for random and uncertain: at 20 labels, 0.831 (0.706-0.903) for both; 50, 0.871 and 0.867; 100, 0.892 and 0.905; 200, 0.901 and 0.927; 300, 0.911 and 0.935; 500, 0.922 and 0.938; all 3,450, 0.938 (0.927-0.954); rules, 0, 0.891 (0.878-0.906). Then labels to reach all minus 0.01: random 326 (150-480), missed 18; uncertain 188 (110-300), missed 0; random over uncertain 1.97 (0.83-3.54), 12 repeats. Then the spam share of the labels bought at 100 and 500: random 0.388 and 0.398, uncertain 0.486 and 0.472. Then each rule on all 4,601 emails, the share it voted on and the share right, from remove, 17.5% and 94.7%, to edu, 11.2% and 86.8%, and the vote, 74.8% and 92.4%. Last: random passes the rules at 100 labels, means.

When I ran it, it printed the versions, the 4,601 emails with 1,813 spam, and the pool and test sizes. Then the table: at 100 labels, random 0.892 and uncertain 0.905; at 500, random 0.922 and uncertain 0.938; all 3,450 labels 0.938; the rules 0.891. Then 326 and 188 labels to reach the target, with 18 and 0 repeats missed, the spam shares, and the ten rules. All of it matches the stored lbg-demo.json. The report's demo mode checks those lines against the file.

To try something I have not run, change STEP from 10 to 50, so that uncertainty sampling buys 50 emails between refits instead of 10. I cannot tell you what it prints, because I have not run it. The question to ask is whether buying in bigger rounds loses the advantage, since the model then chooses 50 emails with one stale idea of what is confusing.

Its json mode writes every number to results/lbg-report.json, which the figures read. The demo mode checks the demo's printed run line by line. The box mode writes the playground on the next slide and checks it against both files.

Here is what came before the run, in the demo's docstring: the data, the split, the model, the four ways, the ten rules, the budgets and the target.

Then there is what came after I saw the results, written into the report's docstring before the report first ran. That covers the paired counts, the look inside repeat 0, the typical-email test, and leaving each rule out. The worked email on the by-hand slide was added while I wrote this lesson, and the report's docstring says so.

One wording change came after the first run too. The demo's docstring first said the emails were collected at Hewlett-Packard, and the dataset's page does not say that, so I corrected it. No number changed.

Four brand cards headed the tools, with their logos, titled what ran where. scikit-learn: the model, the splits and the data download. pandas: the demo's table and its ten rules. NumPy: the random order and the uncertain picks. Python: the report's plain loops, and the box.

The split between the tools is deliberate. scikit-learn did the model, the splits and the download. pandas holds the demo's table and runs its rules. NumPy made the random order and ranked the uncertain picks. The report redoes the rules, the picks and the counting in plain Python. So the check that the numbers come back uses a different method from the one that made them. Plain Python also runs the playground, because it has to run in a browser.

rand
n
while
predict_proba
rules
all

reach(run, kind). The first budget at which a strategy's accuracy reaches the full-pool accuracy minus 0.01, or None if it never does by 500. spread and show turn 30 numbers into a mean, a lowest and a highest.

The rest. The prints, the rules on all 4,601 emails, and, with a file name on the command line, json.dump of every number.