Features And Feature Stores

Embeddings as Features: Did the Meaning of Product Names Help?

0 of 24 complete

0%

Contents

Back|Features And Feature StoresEmbeddings as Features: Did the Meaning of Product Names Help?
1/24
61 min left
Prerequisites
What a Feature Is: A Better Model or a Better Feature?requiredCategorical Features at Scale: One-Hot, Hashing or Target Encoding?required
Related Topics
Train/Serve Skew: One Input Computed Two WaysWhy Production BreaksETL vs ELT: What a Model Loses When the Raw Rows Are Thrown AwayData Engineering for MLVector DatabaseDatabase Types & StorageBefore You Fine-Tune: Try the Cheapest Baseline FirstFine-Tuning: The Narrow DecisionThe Silent Cut: How Much of Your Text an Embedding Model Really ReadsTokens and Embeddings
1 of 24

The Shelf Where Similar Things Sit

Let me start with a small picture.

Imagine a library where the librarian has read every book. She does not shelve the books by their number. She puts books that feel alike next to each other: cookery here, birds there, old maps in the corner. When you ask for something, she points at a shelf, not at one book.

Now suppose she wants to guess which readers will come back next month. She could look at how often each reader came, and how recently. Or she could look at which shelves each reader took books from. The shelves sound richer. They say something about taste. But will they help her guess who comes back? Taste and loyalty are not the same thing.

A flat illustration of a library. A woman at a wooden front desk smiles and points across the room. A young man with a backpack holds a slip of paper and looks at a tall bookshelf where the books are grouped in runs of similar colour. Below the scene: knowing which shelf you took books from is not the same as knowing whether you will come back.

This lesson asks the same question about a real shop. Every product has a name. A computer can place names that mean similar things close together, like the librarian's shelves. I measure whether that helps a model guess which customers will buy again.

Where This Lesson Starts

This lesson uses the same shop, the same customers and the same question as the rest of the chapter. If any of that is new, please read what a feature is first. It sets up the data and the task. It also built six simple facts about each customer, which I reuse here as the starting point. They are how recently they bought, how often, and how much they spent. The other three are how much they sent back, how long they have been a customer, and how many different products they bought.

Lesson 7, on categorical features at scale, gave the model the product CODES a customer bought. A code like 85123A is only a label. It says nothing about the product. The best result there came from filing codes into numbered drawers, and the gain was tiny.

This lesson gives the model the product NAMES instead, turned into numbers that carry meaning. If meaning matters, this is where it should show.

What an is, and how two of them are compared, is taught in the tokens and embeddings chapter: what an embedding is and cosine, dot and distance. I do not teach it again here. I only use it.

The Words You Need First

Please read this slide slowly if any word is new. Every slide after it uses these words.

A hand-drawn grid of eight cards, two per row, each with a word on a tab and its meaning below. Embedding: a list of numbers for a piece of text; similar texts get similar lists. Dimension: one place in that list; 128 dimensions means 128 numbers. TF-IDF: a number per word, high when the word is common in this text and rare in all texts. SVD: squeezes many word columns into a few, keeping the biggest patterns. Cosine: how much two lists point the same way, 1 the same, 0 unrelated. Mean vector: add a customer's product lists number by number and divide by how many. Control: a fake input built to look like the real one, minus the thing tested. Neighbour: the texts whose lists sit closest to this one's. Below: feature, cutoff, label, AP and seed mean what they meant in lessons 1 to 4.

. A list of numbers that stands for a piece of text. A good embedding model gives texts with similar meaning similar lists. Each list is also called a vector.

Dimension. One position in the list. A vector with 128 dimensions is a list of 128 numbers.

TF-IDF. An older way to turn text into numbers, with no neural model. TF means term frequency: how often a word appears in this text. IDF means inverse document frequency: a word that appears in few texts gets a bigger weight than a word that appears in many. Each word gets its own column.

SVD. Short for singular value decomposition. Here it squeezes thousands of word columns into a few, such as 64, keeping the patterns that explain the most. scikit-learn's page says that, used on TF-IDF like this, it is known as latent semantic analysis, or LSA.

Cosine. A measure of how much two vectors point the same way. 1 means the same direction, 0 means no relation.

Mean vector. For one customer, add up the vectors of every product they bought, number by number, and divide by how many products there were.

Control. A fake input, built to look like the real one in every way except the thing being tested. If the real one cannot beat the control, the thing being tested did not help.

Feature, cutoff, label, average precision (AP), ROC-AUC and seed mean what they meant in lessons 1 to 4.

From Many Product Names to One List

A row in this chapter is one customer at one cutoff date. But a customer buys many products, and each product name becomes a vector. So I need one list per row.

I used the simplest rule: the mean of the vectors of every distinct product the customer bought before the cutoff. Each product counts once, however often it was bought. Returns do not count. A customer whose only lines before the cutoff are returns gets empty values, which this model handles by itself. That happened on 960 train rows and 311 test rows.

A hand-drawn flow for customer 12349 at the cutoff 2011-07-01. A top box: 94 distinct products bought before the cutoff, assorted colour bird ornament, big doughnut fridge magnets, birdhouse garden marker and 91 more. An arrow to: each one, 128 numbers from the vector table. An arrow to: the mean, number by number, 128 numbers, starting -0.017, 0.108, -0.448. Below: the text nearest that mean is gumball monochrome coat rack, which they never bought. Label: did not buy in the 30 days after.

Here is one real customer, number 12349 at the cutoff 1 July 2011, the same customer lessons 1 and 7 used. They had bought 94 different products. The lab looked up the 128 numbers of each one and took their mean.

Then I asked which product name sits nearest to that mean. The answer was a gumball monochrome coat rack, a product this customer never bought. The second nearest, a bubblegum ring, they never bought either. A mean of 94 different things looks like none of them. Keep that in mind; it comes back later.

I also tried a second rule, declared before the run: a mean weighted by spend, so a product the customer spent more on counts for more. On 26 train rows and 10 test rows every weight was zero, so those rows fell back to the plain mean.

Two Ways to Turn a Name Into Numbers

The shop's data holds 5,257 different product descriptions, once spaces at the ends are trimmed. I turned every one of them into numbers twice.

Two panels about two ways to turn 5,257 descriptions into numbers. Neural model: 768 numbers; nomic-embed-text in Ollama, run once per text, 83 calls and 50,258 tokens, cut to 64 or 128. Word counts: 2,126 words; TF-IDF on the 4,488 texts seen before March 2011, squeezed by SVD to 64 or 128. Below: SVD kept 29.7 percent of the word variance at 64 and 43.6 percent at 128.

The neural model. I used nomic-embed-text, version 1.5, a small open model of about 0.1 billion parameters under the Apache 2.0 licence, running in Ollama on this laptop. It gives 768 numbers per text. Its model card asks for a short prefix before each text that says what the vectors are for. One prefix, "classification: ", is described as being for "texts into vectors that will be used as features for a classification model". That is exactly this lesson, so I used it. The whole job was 83 calls to Ollama, and Ollama reported 50,258 tokens read.

The word counts. TF-IDF with scikit-learn's default settings, then SVD down to 64 or 128 numbers. I fitted both only on the 4,488 descriptions that appear before 1 March 2011, the end of the training months. They found 2,126 different words. A later description whose words were all new gets a vector of zeros; that happened to 4 texts.

The SVD at 64 numbers kept 29.7 percent of the spread in the word table, and at 128 it kept 43.6 percent. The rest was thrown away.

Both ways were cut to the same two sizes, 64 and 128, so that each pair is a fair race at the same width.

Cutting 768 Numbers Down to 128

A model with 768 numbers per product would give each customer row 768 extra columns. That is a lot for a question this small. So I cut the neural vectors down to 64 and 128 numbers.

A hand-drawn flow of five boxes about the neural path, as the model card says to resize it. 1, "classification: " plus the description. 2, Ollama returns 768 numbers, length 1. 3, layer norm: subtract the mean, divide by the spread. 4, keep the first 128, or 64, and divide by the length. 5, round to 16-bit numbers and store, once. Below: nothing is embedded again on a rerun; the 5,257 vectors are read from the file.

You cannot cut most by simply dropping numbers. This model was trained so that you can: its first numbers carry the most meaning. The chapter on embeddings has a lesson on shorter embeddings that measures this idea. The card gives three steps. First a "layer norm", which subtracts the vector's own mean and divides by its spread. Then keep the first numbers. Then divide by the length, so the result has length 1. The card also reports that the model scores lower on its own tests as the vector gets shorter.

Two practical points. The lab embeds each distinct text exactly once and keeps the result in a file. Run it again and it reads the file; it never calls Ollama twice for the same text. And the stored numbers are rounded to 16-bit floats, half the size of the usual 32-bit ones. Every feature in this lesson is built from those rounded numbers, so the report can rebuild everything from the stored file alone.

The Shop, the Task and the Split

The data is UCI Online Retail II, a public dataset under a CC BY 4.0 licence. It holds every sale of a UK online shop from 1 December 2009 to 9 December 2011. The shop mainly sells gift items, and many of its customers are wholesalers, people who buy to sell again. The chapter's fixed cleaning drops lines with no customer id and keeps returns as flagged rows.

The task is the chapter's fixed one. On the first of each month, for every customer seen before that day, will they buy in the next 30 days? Each pair of a customer and a cutoff is one row, and the label is 1 if they bought.

The rows are split by time. Train months, March 2010 to March 2011, give 44,521 rows; the models learn from these. Valid months, April to June 2011, give 14,673 rows; the one choice in this lesson is made here. Test months, July to November 2011, give 26,851 rows; they only score, and they choose nothing.

Average precision (AP) asks: when the model's list of likely buyers is read from the top, how many of the names near the top really bought? A random list scores about the share of buyers, which averaged 0.196 over the five test months. Every AP in this lesson is worked out per test month and then averaged over the five.

How the Lab Was Built

I wrote the lab's design into the docstring of scripts/labs/features/embeddings_features.py before it first ran. Before that I had looked at the data's columns, counted the distinct descriptions, and sent two descriptions to Ollama once to see the shape of what came back. No model had been trained on this task with any text.

A flowchart. 5,257 texts, embedded once, feed a mean per customer, before the cutoff, then six, plus a block, then the default model with 20 seeds, which splits into valid: pick, and test: score once. Below: every vector table is fixed before any model trains; the test months choose nothing.

The model is the chapter's usual one: scikit-learn's HistGradientBoostingClassifier with its default settings. It builds many small decision trees, each one fixing the mistakes of the trees before it. Every contestant is lesson 1's six features plus one block of extra columns. Only the block changes.

A page of six labelled zones in two columns, the blocks added to lesson 1's six. Nothing added, 1: the six alone. Neural model, 3: mean of nomic vectors at 64 or 128, and 128 weighted by spend. Word counts, 3: mean of TF-IDF and SVD vectors at 64 or 128, and 128 weighted by spend. Random directions, 2: one fixed random list per description, averaged the same way, 64 or 128. Pure noise, 2: 64 or 128 random columns per row, tied to nothing. The block alone, 3: neural 128, word counts 128 and random 128, with no six. Below: random directions know which products a customer bought, but nothing about what the words mean.

The two kinds of fake block matter most. Random directions give every description a fixed list of random numbers, then take the customer's mean the same way. This control knows exactly which products each customer bought, as the real do. It knows nothing about the names. Pure noise is just random columns, the same for nobody. It tests whether a model gains from more columns alone.

The design wrote down six guesses so they could be wrong in public. Here they are, scored.

  1. No block beats the six by an interval clear of zero. Right for every real embedding. With the first random table, random directions at 64 did clear zero; over ten random tables, drawn after a review, they did not.

What the Vectors Put Together

Before any model, it is worth checking that the vectors do what they promise. For the three products bought on the most lines before March 2011, the lab listed the five nearest other descriptions in each vector table.

Two columns for assorted colour bird ornament, five nearest of 5,257 texts by cosine. Neural model at 128: painted bird assorted christmas 0.928, etched glass bird tree decoration 0.921, ivory hanging decoration bird 0.915, silver finch decoration 0.912, bird decoration red spot 0.908. Word counts at 128: assorted colour silk glasses case 0.794, assorted colour set 6 teaspoons 0.753, painted bird assorted christmas 0.727, assorted colour jumbo pen 0.706, assorted flower colour leis 0.704. Below the left column: birds and decorations, the meaning. Below the right: anything assorted colour, the words.

This one shows the difference best. The neural model found birds and decorations: a silver finch decoration, an etched glass bird. The word counts found anything that shares the words "assorted colour": a glasses case, teaspoons, a pen. Word counts match words, not meaning.

For the regency cakestand, both did well: the neural model's nearest was a sweetheart cakestand, 3 tier, at 0.934. For the white hanging heart t-light holder, both found other hanging heart holders. The random table found nothing in common at all; its nearest to the bird ornament began with a floral print tape.

Across all 5,257 texts, the 10 nearest neighbours in the neural table and in the word table shared 3.41 texts on average. The neural table and the random table shared 0.02. So the two real tables agree in part, and the random one agrees with nothing. It is a fair control.

A scatter chart, a rough map of the 150 most-bought descriptions, the neural vectors flattened to two directions, first direction across and second direction up, both from about -0.3 to 0.3. Small dots spread over the whole area. Four big dots are labelled below. Below: big dots, white hanging heart t-light holder, regency cakestand 3 tier, assorted colour bird ornament, 60 teatime fairy cake cases. Within 0.05 of the first sit 6 others, 3 of them metal signs, such as please one person metal sign. These two directions hold only 9.8 percent of the spread: a flat map puts things together that the 128 numbers keep apart.

A map is tempting, so I drew one, after the results, as a picture and not a finding. It is honest only about its own weakness. The first two directions hold just 9.8 percent of the spread. On this flat map, three metal signs sit right next to the white hanging heart. None of them is among its five nearest in the full 128 numbers.

The Lab's Report, Running

This is a real recording of the report script, emb_report.py, on the laptop where the lab ran. It trains nothing new except four checks. It reads the stored results and the stored vectors, and checks every number this lesson uses.

A terminal recording of emb_report.py. It prints a table with one line per set of inputs: columns, mean test AP over 20 seeds, and the 95 percent interval of its gap to the six. Then the gaps between meaning and random directions, the pick on the valid months and at seed 0, the two nearest texts to the bird ornament in each table, a table headed after the results about how crowded the vectors are, a section headed after the review with the ten-table intervals and the range of six plus random 64 across tables, a line about the demo and the box, and a last line saying every check against the stored run and the raw invoices is equal.

The report rebuilds every customer's mean vector from the raw invoices with its own, simpler code: a plain group-by mean per cutoff, not the lab's sparse table product. It compares those features with sums the lab stored, then retrains four models at seed 0 and checks their test scores to twelve decimal places.

It refits TF-IDF and SVD and checks that they give exactly the stored 16-bit vectors. If the raw 768-number file is present, it also redoes the model card's resize and checks those vectors. It recomputes the nearest neighbours, the overlaps, every mean of 20 seeds, every table mean of the ten-table check and every interval from its 1,000 stored redraws. It checks the student demo against the lab and runs the playground box. Last, it checks that every number this lesson quotes appears in the lesson text. If anything disagrees, it stops with an error.

What Each Block Scored

Here is every set of inputs with the six, scored on the five test months and averaged over 20 training seeds.

A strip chart with one row per set of inputs, from the six at the top to the six plus noise 128 at the bottom. Each row has 20 short ticks, one per seed, and a diamond at their mean, on an axis from about 0.536 to 0.554. A dashed vertical line marks the six. The diamonds all sit within a few thousandths of the line; random 64 sits furthest right and noise 128 furthest left. Below: dashed, the six, 0.5452. Picked on valid: words 128 with spend, 0.5460. Highest: random 64, 0.5486, with one random table; ten tables averaged 0.5466.

inputscolumnstest APROC-AUC
the six alone60.54520.798
+ neural 64700.54650.799
+ neural 1281340.54640.799
+ neural 128, spend1340.54720.799
+ words 64

How Sure Can I Be?

Gaps of a few thousandths need care. I checked two kinds of luck.

Luck in training. The model sets aside 10 percent of the train rows at random to decide when to stop adding trees. This is called early stopping, and the seed moves it. So every set of inputs was trained with 20 seeds, 0 to 19. The six alone ranged from 0.5385 to 0.5488 over the seeds. That spread is larger than any gain in the table.

Luck in which customers were tested. A bootstrap draws the test customers again at random, with repeats allowed, within each test month. I did that 1,000 times, with the same draws for every set of inputs. Each time I worked out the 20-seed mean AP of each one minus the six. The middle 95 percent of those gaps is the 95% interval.

Ten rows, one per block against the six, each with its 95 percent interval written out and drawn as a line with end marks and a dot at its middle, against a dashed line at zero. Neural 64, -0.0014 to +0.0039. Neural 128, -0.0017 to +0.0039. Neural 128 with spend, -0.0008 to +0.0047. Words 64, -0.0031 to +0.0024. Words 128, -0.0039 to +0.0020. Words 128 with spend, -0.0021 to +0.0035. Random 64, +0.0010 to +0.0059, drawn heavier, wholly right of zero. Random 128, -0.0016 to +0.0035. Noise 64, -0.0028 to +0.0016. Noise 128, -0.0044 to +0.0005. Below: dashed, no gap; random 64 used one table, and over ten tables it was -0.0006 to +0.0035.

Every real block crossed zero. The neural model at 128 numbers: -0.0017 to +0.0039. The pick, word counts at 128 with spend: -0.0021 to +0.0035. None of them can be told apart from the six.

Only one interval sat clear of zero: random directions at 64, +0.0010 to +0.0059, above the six in all 5 test months. An independent review then retrained it with five other random tables, and found mine the luckiest.

So, after the review, I drew ten new random tables, with the design written into the lab's docstring before it ran. Over all ten, random 64 against the six was -0.0006 to +0.0035, which crosses zero. The one clear interval came from one lucky draw of the random table, the same mistake as reading one shuffle as a finding. What survives is narrower: a block with no meaning at all did as well as the blocks with meaning.

Meaning Against Random Directions

The fair question is not "did the six plus a block beat the six". A block adds columns, and the columns alone could matter. The fair question is whether a block with meaning beats a block of the same width without it.

A grid of small squares, one per column the model was given. The six: 6 columns, a short row of 6 squares. The six plus a 128 block: 134 columns, about seven rows of 20 squares. Below: pure noise at the same width scored 0.5431, random directions 0.5461, the neural model 0.5464. A fair test of meaning compares blocks of the same width.

So the design put each real block against random directions of the same width. With the first random table alone, random 64 seemed to beat both real 64 blocks. Neural minus random was -0.0042 to -0.0001, and words minus random -0.0058 to -0.0016. That rested on one draw of the table.

After the review I drew ten new random tables and ten new noise draws. I retrained every random and noise contestant on each one with 20 seeds, and bootstrapped the mean over all 200 fits. These intervals were added after the results, and are labelled so:

real minus random, ten tables95% interval
neural 64-0.0018 to +0.0015
words 64-0.0036 to +0.0001
neural 128-0.0006 to +0.0027
words 128-0.0026 to +0.0006

Every interval crosses zero. No real beat random directions at either width, and random directions did not beat them either. Random directions did as well as either embedding; how far ahead they look depends on which random table you draw.

The Blocks on Their Own

A block that adds nothing to the six could still carry signal of its own. Maybe the six already know it. So the design also trained each 128 block with no six at all.

A bar chart of mean test AP over 20 seeds for four sets of inputs: the six, about 0.545; neural alone, about 0.425; words alone, about 0.426; random alone, about 0.436. A dashed line marks the share of buyers, about 0.196. Below: alone, neural 0.4248, words 0.4263, random 0.4360 with one table and 0.4329 over ten; neural minus random, ten tables, -0.0137 to -0.0023.

All three blocks carry real signal on their own: they score about 0.43, far above the 0.196 of a random list. A customer's products say something about whether they come back. But all three are far below the six at 0.5452, and the six already hold most of what the blocks know.

Here the random block came out on top, and the ten tables held it up. Over ten tables, the neural block alone minus random was -0.0137 to -0.0023 at 128, and -0.0163 to -0.0040 at 64. Word counts alone minus random was -0.0121 to -0.0015 at 128, and -0.0117 to +0.0011 at 64, which crosses zero. With the first single table, the two 128 gaps had been -0.0188 to -0.0041 and -0.0171 to -0.0030. All ten random tables scored above both real 128 blocks alone.

So the result is not just "the six already had it". On their own, a mean of the neural vectors told the model less than a mean of random ones. With the six, the two could not be told apart.

Why Would Random Do as Well? A Check After the Results

Everything on this slide was designed after I had seen the main results. Its design was written into the lab's docstring before it ran, and it trains no model.

My idea: product names that mean similar things get similar vectors. In this shop most products are gifts and home decorations, so many vectors point roughly the same way. A mean over 30 or more of them gets pulled toward one common direction, and customers end up looking alike. Random directions sit almost at right angles to each other, so a mean of them keeps more of which products went in.

A page in three parts, asked after the results, how close together are the lists. Neural 128: 0.690 between two products, 0.948 between two customers, with a long bar. Words 128: 0.023 between two products, 0.316 between two customers, with a shorter bar. Random 128: 0.000 between two products, 0.033 between two customers, with a dot. Below: mean cosine over all pairs; customers at 1 March 2011, 4,539 of them.

The numbers fit the idea for the neural model. Two of its product vectors had a mean cosine of 0.690, so they mostly point one way. Two customers' mean vectors at 1 March 2011 had a mean cosine of 0.948: nearly the same direction. For random directions those numbers were 0.000 and 0.033.

But the numbers do not fit for word counts. Its product vectors were barely alike, 0.023, and two customers only 0.316. Yet at 128, word counts alone also lost to random directions alone, by about as much as the neural model did. So crowding cannot be the whole story.

There is also a reason to doubt crowding matters at all. A tree model cuts each column at its own values, into at most 255 steps. Values packed close together are spread over those steps just the same. So the cause is not pinned down. Crowding describes the vectors; it does not show that crowding is what cost the score.

The Pick, and the Seed That Lied

Two smaller lessons came out of the run, both about how easy it is to fool yourself.

A table of the six real candidates with their valid AP, used to choose, and test AP, used to score, both means of 20 seeds. Neural 64, 0.5293 and 0.5465. Neural 128, 0.5281 and 0.5464. Neural 128 with spend, 0.5290 and 0.5472. Words 64, 0.5266 and 0.5449. Words 128, 0.5280 and 0.5444. Words 128 with spend, highlighted, 0.5296 and 0.5460. Below: seed 0 alone would have picked neural 64. The pick against the six: -0.0021 to +0.0035.

The pick. The valid months chose word counts at 128 with spend weights, at 0.5296. The next two were within 0.0007 of it. With seed 0 alone, the valid months would have chosen the neural model at 64. Either way, the picked block could not be told apart from the six on the test months.

A dot chart, one seed against twenty, for the six plus neural 128. A big number: seed 0, 0.5486. Below it: mean of 20 seeds, 0.5464, from 0.5422 to 0.5489. Twenty dots on an axis from 0.541 to 0.550, with a dashed line at the six; seed 0 is the highlighted dot, near the right end. Below: seed 0 alone put neural 128 +0.0036 above the six's own seed 0; twenty seeds put it +0.0013, and its range crossed zero.

The seed. Suppose I had trained the neural block at 128 with seed 0 only. I would have seen 0.5486 against 0.5450 for the six at seed 0. That is a gain of +0.0036, from the third best seed of the twenty. Over 20 seeds the gain shrank to +0.0013, and the interval crossed zero. One seed is one draw. The chapter has learned this the hard way more than once, and it would have fooled me here too.

What a Block Costs

Even if a block helped, it would not be free. Here is what this one costs.

An isometric drawing of four blocks, one per set of columns, each taller for more bytes per customer row as 32-bit numbers, not to scale. The six, 24 bytes, a thin slab. The six plus 64, 280 bytes. The six plus 128, 536 bytes. Lesson 7's bag, 4,008 bytes, a tall cube. Below: the table of 5,257 vectors at 128 numbers, 2.7 MB as 32-bit, 1.3 MB as stored, 16-bit.

Per customer. A 128 block adds 512 bytes per row as 32-bit numbers, against 24 for the six. The training table the model gets, in 64-bit numbers, grew from 2.1 MB to 47.7 MB.

The vector table. 5,257 vectors of 128 numbers are 2.7 MB as 32-bit numbers, or 1.3 MB as I stored them. That is small. It is also a one-off cost: 83 calls to Ollama, 50,258 tokens. I give no time for it. Ollama and the laptop were shared with other work while it ran, so any seconds would be unfair.

A sequence diagram with four lifelines: shop, embedder, feature store and model. Step 1, the shop sends the embedder a new description. Step 2, the embedder stores its 128 numbers once in the feature store. Step 3, the shop tells the feature store a customer bought it, so their mean is updated. Step 4, the model asks the feature store for 6 plus 128 numbers. A large 5,257 in the corner is labelled texts embedded once, offline. Below: step 2 must happen before step 3, or a new product has no vector; 4 texts had no word TF-IDF knew.

At serving time. The real cost is in the moving parts. A new product needs its vector before any customer's mean can include it, so an embedding job must run before the updates the customer. If the embedding model ever changes, every vector and every customer mean must be rebuilt, which is the versioning problem of an earlier lesson in a new form. For word counts there is a quieter trap: a product whose words were all unseen gets a vector of zeros. Four texts here had that.

The Lab's Code, Piece by Piece

The lab is one file, scripts/labs/features/embeddings_features.py. It imports lesson 1's feature code for the six, so they are built exactly as before. Its base contestant must reproduce lesson 1's score of 0.5450 at seed 0, to twelve decimal places, or the run stops. It did.

build_vectors embeds each distinct description once. It sends batches of 64 texts to Ollama's /api/embed, each with the "classification: " prefix, and keeps the 768 numbers in a file outside the repository. matryoshka does the card's resize: layer norm, keep the first 64 or 128, divide by the length. For word counts it fits TfidfVectorizer() and TruncatedSVD(n_components=k, random_state=0) on the texts from before March 2011 only. The four tables are saved, rounded to 16 bits, in results/emb-vectors.npz. If that file exists, nothing is embedded at all.

load_tables reads that file and adds the random directions: one fixed list of normal random numbers per description, from numpy.random.default_rng(7), each divided by its length. After the review, multi draws ten more tables from default_rng(1000 + k) and ten noise draws from default_rng(2000 + k). It retrains every random and noise contestant on each, and bootstraps the mean over all 200 fits per contestant.

Try It Yourself

The full lab trains 280 models, plus 1,780 more for the check after the review, and needs Ollama. I wrote a small demo that needs neither: it uses the word-count path only. It builds every customer's mean of TF-IDF vectors at 64 numbers, and of random directions, then trains three models at seed 0 and prints their test AP.

A page of four labelled zones in two columns, headed emb_demo.py, designed before it ran. The data: the chapter's shop data through task.py, and lesson 1's six. The vectors: TF-IDF and SVD at 64 numbers, and random directions. The models: three, seed 0; the six; plus words 64; plus random 64. The check: every test month must match the lab's seed 0. Below: it printed 0.5450, 0.5449 and 0.5458, as the lab did at seed 0.

I wrote the demo's design into its docstring after the lab had run and before the demo first ran.

A real screenshot of VS Code with emb_demo.py open at the top of the file, showing its docstring: what it needs, how to run it, and the design written before it first ran.

Before you run this lab. The full lab uses nomic-embed-text, running in Ollama on your own computer. The demo on this slide does not need it. If you want to run the full lab and have not set Ollama up yet, the lab setup guide shows how to install Ollama, download the model and check that everything works, on macOS, Windows or Linux. You can use a different model instead: the guide shows the one line to change, and your numbers will differ from the ones in this lesson.

For the demo you need Python 3 with pandas, pyarrow and scikit-learn: pip install pandas pyarrow scikit-learn openpyxl. First run python fetch_data.py from the scripts/labs/features folder. It downloads the shop data once, about 46 MB, and writes one cleaned file. Then run the demo. It needs no GPU. I ran it with scikit-learn 1.9.1 on a Mac. These libraries run on Windows and Linux too, but I have not checked the numbers there. Give it a file name, python emb_demo.py out.json, and it also saves every number. That is how was made.

Find a Product's Neighbours Yourself

This box holds the 400 real descriptions bought on the most lines before March 2011. It also holds TF-IDF with scikit-learn's default settings, written in plain Python, so it runs in your browser. The report checked that it gives exactly scikit-learn's numbers on these 400 texts.

Press Run to see the five descriptions nearest to the white hanging heart t-light holder. It also counts how many of the others share no word with it at all. For TF-IDF those get a cosine of exactly 0, however close their meaning: 321 of the 399. Then change MINE to any description in the list, such as "REGENCY CAKESTAND 3 TIER", and run again.

At the bottom it prints each set of inputs' real mean test AP over 20 seeds from the lab. So you can set what the vectors find next to what they did, or did not do, for the score.

Common Mistakes, and When Embeddings Help

No same-width control. Adding 128 columns can move a score with no meaning in them at all. Always compare against random directions and pure noise of the same width. Here, over ten random tables, random directions were never beaten by a real at either width. And draw several random tables, not one: the first one I drew was the luckiest of eleven.

Trusting one seed. At seed 0 the neural block looked like a gain of +0.0036. Over 20 seeds it was +0.0013, and the interval crossed zero.

Fitting the text model on the future. Fit TF-IDF and SVD on texts from before training ends, as here. A pretrained neural model learned from other text, not from your labels, so it reads no future label; but it can still change version, and every vector with it.

Embedding the same text again and again. Cache by text. Here 5,257 texts covered every line of two years of invoices.

When embeddings may help. When the text is the main thing you know about an item, for example a new product with no sales yet, where a code or a random direction has nothing to go on. When the model can use each item's vector directly, rather than a blurred mean. Or when the task is about the content itself, such as search or grouping. I did not test any of these here.

When to skip them. On evidence like this lesson's, when strong behaviour features already exist, as the six do here, and the question is about behaviour, such as whether someone comes back.

What This Lab Cannot Tell You

Two columns titled what this lab shows, and what it cannot. Shows: ten blocks on top of lesson 1's six, with one tree model, on one shop; that random directions did as well as two real embedders here; what the vectors put near each other. Cannot show: a bigger or better embedding model, or a model trained on this shop; other ways to combine a customer's products than a mean; why random directions did as well.

One shop, one task, one model. A tree model with default settings, on a 30-day buying question, with one small model and one word-count method. A larger model, a model trained on this shop's own data, or a linear model might behave differently. I did not test them.

One way to combine. I took a plain mean, and a spend-weighted one. A mean of 94 products looks like none of them, as customer 12349 showed. Other ways, such as the most recent products only, or one vector per product group, could keep more. I did not build them.

One check after the results. The crowding check is labelled on its own slide, and its design was written down before it ran. It describes the vectors; it does not explain the score.

No timings. The laptop and Ollama were shared, so I report counts and bytes only.

What to Do on Monday

A hand-drawn grid of six cards, titled six steps, before you add embeddings to a tabular model. 1, embed once: cache by text, never embed again on a rerun. 2, fit on the past: fit TF-IDF and SVD on texts from before training ends. 3, add a control: random directions at the same width; draw several tables. 4, add noise too: the same number of columns, tied to nothing. 5, use 20 seeds: one seed moved neural 128 by more than its gain. 6, count the bytes: per row, and the vector table. Below: if meaning cannot beat random directions, the meaning is not what helped.

These are the steps I would take the next time someone says "let's add ".

  1. Embed once, and cache by text. It is cheap to store and expensive to repeat.

  2. Fit anything fitted on the past only. TF-IDF and SVD here saw only texts from before March 2011.

  3. Build random-direction controls of the same width, and draw several tables. It is three lines of code. Here one table seemed to beat both real embedders at 64 numbers; ten tables showed that was a lucky draw.

  4. Add pure noise too, to see what extra columns alone do.

  5. Use 20 seeds and an interval. One seed showed a gain of +0.0036 that was not there.

  6. Count the bytes, per row and for the vector table, before you put it in a .

A closing card titled meaning did not beat a random control. A level balance beam with two pans: random 64, 0.5466, ten tables, and neural 64, 0.5465. Below: neural minus random, -0.0018 to +0.0015; how far apart depends on which random table you draw.

The one idea to keep: on this shop, the vectors grouped products by meaning very well, and the model did not need it. Which products a customer bought helped a little over pure-noise columns; what those products were called did not add anything I could measure. Random directions did as well as either embedding, and how far ahead they look depends on which random table you draw.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

Why did the lab add a block of random directions, one fixed random list per description?

Q2

The neural block at 128 numbers scored +0.0036 over the six at seed 0. What did 20 seeds show?

Q3

For the assorted colour bird ornament, how did the two real embedders' neighbours differ?

Q4

Why should a gain from 128 extra columns be compared with noise and random directions of the same width?

  • The neural model and word counts cannot be told apart at the same width. Wrong at 128: the neural model beat word counts, +0.0003 to +0.0038. At 64 the design held no such pair, so it was not tested.

  • Random directions land with the embeddings. Right, once ten random tables were drawn: no real block could be told apart from random directions at either width. The first single table had made it look as if random beat both real blocks at 64.

  • Pure noise does not beat the six. Right.

  • Alone, the neural block beats the random block, and every block alone scores far below the six. The first half was wrong, in the opposite direction: random beat the neural block alone, at both widths over ten tables. The second half was right: about 0.43 against 0.5452.

  • Weighting by spend changes nothing measurable. I cannot score this one. I did not put the spend arms against the plain arms in the design, so there is no interval for that pair. Both spend arms, like the plain ones, could not be told apart from the six.

  • 70
    0.5449
    0.798
    + words 1281340.54440.798
    + words 128, spend1340.54600.798
    + random 64700.54860.799
    + random 1281340.54610.798
    + noise 64700.54450.799
    + noise 1281340.54310.799

    Every row sits within a few thousandths of the six. The highest mean came from random directions at 64 numbers, a block that knows nothing about the names. That was one random table, and a later check found it the luckiest of eleven; the next slide explains.

    The design allowed one choice, made on the valid months. Among the six real blocks, pick the one with the best mean valid AP over 20 seeds. It picked word counts at 128 with spend weights. On the test months that scored 0.5460, against 0.5452 for the six.

    For comparison, lesson 7's bag of product codes, hashed into 1,024 drawers with 1,002 columns, scored 0.5442 with the same six and the same seeds. That came from lesson 7's own run, not this one.

    A dot chart of mean test AP over 20 seeds by block width, 64 and 128, with four series: neural, words, random and noise, and a dashed line for the six. At 64 the random dot sits highest, near 0.5486, neural near 0.5465, words and noise just under the line. At 128 neural and random sit just above the line, words and noise below it. Below: at 64, random 0.5486, neural 0.5465, words 0.5449, noise 0.5445; that random table was the luckiest of eleven drawn.

    This chart uses the first run's single random table, the one the lab drew before any result was seen. Read it next to the one below, which shows what the same contestant scored with ten other tables.

    A dot chart, asked after a review: six plus random 64, one dot per random table, mean of 20 seeds, on an axis from 0.544 to 0.550. Ten dots for ten new tables spread between about 0.5454 and 0.5475; one darker dot for the first run's table sits higher, at 0.5486. Dashed lines mark neural 64 and words 64. Below: ten new tables 0.5454 to 0.5475, mean 0.5466; the first run's table 0.5486; 6 of ten sit above neural 64. Neural 64 minus random 64, over ten tables: -0.0018 to +0.0015.

    The ten new tables put six plus random 64 between 0.5454 and 0.5475, with a mean of 0.5466. The first run's table, at 0.5486, sat above all ten. Six of the ten sat above neural 64, and all ten above words 64. At 128, only 1 of ten sat above neural 128, and 8 above words 128. So which side wins depends on the draw.

    Random directions against pure noise, over ten tables and ten noise draws: +0.0010 to +0.0036 at 64 and +0.0009 to +0.0038 at 128. So knowing WHICH products a customer bought helped a little over pure-noise columns. Against the six alone, though, random directions crossed zero at both widths. Neural beat word counts at 128, +0.0003 to +0.0038. But neither beat random directions, so that edge is not evidence that meaning helped.

    bag_matrices builds a sparse table for each cutoff, with one row per customer and one column per description. A cell holds 1 if they bought it before the cutoff. A second table holds their spend on it. means multiplies that table by a vector table and divides by the row count, which gives every customer's mean in one step. A customer with no purchase gets empty values.

    mean20_bootstrap comes from lesson 7's lab unchanged. It redraws customers within each test month, the same draws for every contestant, and checks its fast AP against scikit-learn's own on the first draws before trusting it.

    results/emb-demo.json
    """Do product-description embeddings help a tabular model? The lab, small.
    
    Lesson 11 of 'Features and Feature Stores'. It needs Python 3 with pandas,
    pyarrow and scikit-learn, and the shop data: run fetch_data.py once first
    (it needs openpyxl too, and downloads UCI Online Retail II, about 46 MB).
    It needs NO Ollama and NO GPU: it uses the TF-IDF path only.
    Then, from the folder above this one or from this folder:
        python examples/emb_demo.py            # print the table
        python examples/emb_demo.py out.json   # and save every number
    It prints no timings.
    
    Design, written 2026-10-01 after the lab (embeddings_features.py) had run
    and before this file first ran:
      Task and data: task.py, exactly as in the lab. The six features are
      lesson 1's, built by lesson 1's own code (what_a_feature_is.py).
      The texts: every distinct product description, spaces trimmed. TF-IDF
      (scikit-learn defaults) is fitted on the descriptions seen before
      1 March 2011, then TruncatedSVD keeps 64 numbers, then each vector is
      divided by its length and rounded to 16-bit floats, as the lab stores it.
      A customer's feature: the mean vector of the distinct descriptions they
      bought before the cutoff (returns not counted).
      Three models, the default HistGradientBoostingClassifier, seed 0: the
      six alone; plus the TF-IDF mean; plus the mean of RANDOM vectors, one
      fixed random direction per description (the lab's control).
      It also prints the five nearest descriptions to the shop's most-bought
      product. It must agree with the lab's seed-0 test AP on every test
      cutoff, to 1e-9. emb_report.py checks this.
    
    Author: Roni Das
    Created: 2026-10-01
    """
    import json
    import sys
    from pathlib import Path
    
    import numpy as np
    import scipy.sparse as sp
    from sklearn.decomposition import TruncatedSVD
    from sklearn.ensemble import HistGradientBoostingClassifier as HGB
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.metrics import average_precision_score
    
    sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
    import task  # noqa: E402
    from what_a_feature_is import HAND_COLS, joined  # noqa: E402
    
    DIM = 64
    
    
    def unit(x):
        n = np.linalg.norm(x, axis=1, keepdims=True)
        return np.divide(x, n, out=np.zeros_like(x), where=n > 0)
    
    
    def mean_vectors(ev, d, texts, table):
        """Each row's mean vector over the distinct texts bought before its cutoff."""
        col = {t: i for i, t in enumerate(texts)}
        buys = ev[~ev["is_return"]]
        rows, cols = [], []
        for t, g in d.groupby("cutoff").indices.items():
            past = buys[buys["ts"] < t]
            pairs = past[["customer_id", "text"]].drop_duplicates()
            pairs = pairs.sort_values(["customer_id", "text"])
            where = dict(zip(d["customer_id"].to_numpy()[g], g))
            pairs = pairs[pairs["customer_id"].isin(where)]
            rows += [where[c] for c in pairs["customer_id"]]
            cols += [col[x] for x in pairs["text"]]
        bag = sp.csr_matrix((np.ones(len(rows)), (rows, cols)),
                            shape=(len(d), len(texts)))
        n = np.asarray(bag.sum(axis=1)).ravel()
        out = np.full((len(d), table.shape[1]), np.nan)
        out[n > 0] = (bag[n > 0] @ table) / n[n > 0, None]
        return out
    
    
    def test_ap(x_tr, y_tr, x_te, d_te):
        m = HGB(random_state=0).fit(x_tr, y_tr)
        p = m.predict_proba(x_te)[:, 1]
        d = d_te.assign(p=p)
        return [average_precision_score(g["label"], g["p"])
                for _, g in d.groupby("cutoff")]
    
    
    ev = task.load_events()
    ev["text"] = ev["description"].str.strip()
    lab_tr, _, lab_te = task.splits(ev)
    tr = joined(ev, lab_tr, task.TRAIN_CUTOFFS)
    te = joined(ev, lab_te, task.TEST_CUTOFFS)
    texts = sorted(ev["text"].unique())
    early = sorted(ev.loc[ev["ts"] < "2011-03-01", "text"].unique())
    print(f"rows: train {len(tr):,}, test {len(te):,}")
    print(f"distinct descriptions: {len(texts):,} ({len(early):,} before March 2011)")
    
    vec = TfidfVectorizer().fit(early)
    svd = TruncatedSVD(n_components=DIM, random_state=0).fit(vec.transform(early))
    tfidf = unit(svd.transform(vec.transform(texts)))
    tfidf = tfidf.astype(np.float16).astype(np.float64)
    rand = unit(np.random.default_rng(7).standard_normal((len(texts), 128))[:, :DIM])
    print(f"TF-IDF words: {len(vec.vocabulary_):,}; SVD keeps "
          f"{svd.explained_variance_ratio_.sum():.1%} of their variance")
    
    top = ev.loc[(ev["ts"] < "2011-03-01") & ~ev["is_return"], "text"]
    top = top.value_counts().index[0]
    i = texts.index(top)
    cos = tfidf @ tfidf[i]
    cos[i] = -np.inf
    print(f"\nnearest to '{top}' (TF-IDF, {DIM} numbers):")
    for j in np.argsort(-cos, kind="mergesort")[:5]:
        print(f"  {cos[j]:.3f}  {texts[j]}")
    
    y = tr["label"].to_numpy()
    six_tr, six_te = (d[HAND_COLS].to_numpy(float) for d in (tr, te))
    arms = {"six alone": (six_tr, six_te)}
    for name, table in ((f"+ tfidf {DIM}", tfidf), (f"+ random {DIM}", rand)):
        arms[name] = (np.hstack([six_tr, mean_vectors(ev, tr, texts, table)]),
                      np.hstack([six_te, mean_vectors(ev, te, texts, table)]))
    
    print(f"\n{'inputs':14s} {'columns':>7s} {'test AP':>8s}")
    results = {}
    for name, (x_tr, x_te) in arms.items():
        aps = test_ap(x_tr, y, x_te, te)
        results[name] = {"columns": x_tr.shape[1], "test": list(map(float, aps))}
        print(f"{name:14s} {x_tr.shape[1]:7d} {np.mean(aps):8.4f}")
    if len(sys.argv) > 1:
        json.dump({"texts": len(texts), "early": len(early),
                   "vocabulary": len(vec.vocabulary_), "top": top,
                   "neighbours": [texts[j] for j in np.argsort(-cos, kind="mergesort")[:5]],
                   "arms": results}, open(sys.argv[1], "w"), indent=1)
    

    This is a real run in VS Code's terminal: python emb_demo.py, run inside the examples folder.

    A real screenshot of VS Code's terminal after running python emb_demo.py inside the examples folder. It prints the train and test row counts, the number of distinct descriptions and how many came before March 2011, the TF-IDF word count and the share of variance the SVD keeps, the five nearest descriptions to the white hanging heart t-light holder with their cosines, and then the columns and test AP of three sets of inputs: the six alone, plus tfidf 64, and plus random 64.

    When I ran it, every test month matched the lab's seed 0 within one billionth, and the report checks this from the stored files. At seed 0 it printed 0.5450 for the six, 0.5449 with words and 0.5458 with random directions. Over 20 seeds the lab found the same order for these three. Notice the nearest descriptions too: TF-IDF at 64 numbers puts a heart string memo holder among the hanging heart t-light holders, because it shares three of its words.

    Four cards with logos in two columns, titled what ran where. Ollama: nomic-embed-text, run once per text; the demo skips it. scikit-learn: TF-IDF, SVD and the model. pandas: the invoices and each customer's products. NumPy: the vector table, the means, the bootstrap. Below: all of it ran on a laptop CPU; no paid service, no GPU.

    Everything here ran on a laptop CPU with no paid service. I give no timings in this lesson, because the laptop and Ollama were shared with other work the whole time.