Features And Feature Stores

Feast Hands-On: The Same Join, Row for Row, Until a Row Has No Snapshot in Its Window

0 of 28 complete

0%

Contents

Back|Features And Feature StoresFeast Hands-On: The Same Join, Row for Row, Until a Row Has No Snapshot in Its Window
1/28
66 min left
Prerequisites
What a Feature Is: A Better Model or a Better Feature?requiredPoint-in-Time Joins: A Latest-Value Join Promised 0.714 and Delivered 0.301requiredMissing at Serving: No Fill Gives Back a Row That Is Gonerequired
Related Topics
Feature Stores: Killing the Train and Serve Skew BugCore Concepts
1 of 28

My Own Notebook, Then a Real Library

Let me start with a notebook and a library.

Imagine you help at a small club. For a year, you keep every member's record in your own notebook. Each night you write a new line for anyone who did something that day. The line says how many times they have visited so far, and how much they have paid so far. You date each line with the next morning, because that is when the line is finished.

When someone asks "how active was this member on 1 July?", you do not read the newest line. You turn back to the last line dated on or before 1 July. That way you never use what happened later.

A flat illustration of a woman in a large library, holding a stack of books and standing beside a book cart, looking thoughtfully along the tall shelves. Below the scene: for eleven lessons I kept every customer's record by hand; now I ask a real library the same question, what did this card say on 1 July?

Now a real library offers to keep the records for you. It has a proper catalogue, rules and staff. It says it can answer the same question for thousands of members at once. Before you hand over your notebook, you would want to check one thing. Does the library give the same answer as your notebook, for every member and every date? And are there any library rules you did not know about?

I do exactly this in this lesson. My notebook is the code I wrote in lessons 1 and 2. The library is a real called Feast.

Where This Lesson Starts

This is the last lesson of the chapter, so it leans on the earlier ones. If the shop, the customers or the question are new, please read what a feature is first. That lesson built six features from each customer's history: recency, frequency, money, return share, tenure and products. The question is always the same. At the start of a month, will this customer buy something in the next 30 days?

The most important earlier lesson here is point-in-time joins. It showed that joining each customer's latest values promised a score of 0.714 and delivered 0.301, and it wrote the correct join by hand with pandas' merge_asof. In this lesson I hand the same job to Feast and compare the two, row by row.

Three more lessons come back here. Feature freshness described how a copies values to a fast store. Online and offline consistency compared two programs that compute the same feature. Missing at serving read in Feast's source code what happens to a customer the store does not know. Here I run it.

The survey lesson feature stores explains what a feature store is and why teams use one. I do not repeat it. This lesson only measures.

The Words You Need First

Please read this slide slowly if any word is new. Every slide after it uses these words.

A hand-drawn grid of nine cards, three per row. Feature store: a system that keeps feature values for training and for serving. Feast: an open-source feature store; this lab runs version 0.66.0. Entity: the thing features describe; here, one customer. Feature view: a named group of features, read from one table. Snapshot row: one customer's values at the end of a day, stamped the next midnight. Historical read: get_historical_features, values as of each row's own time. Materialize: copy the newest values into the fast online store. Online read: get_online_features, the newest value, for a live request. TTL: time to live, how old a value may be and still count. Below: cutoff, label, AP and the point-in-time join mean what they meant in lessons 1 to 11.

A is a system that keeps feature values, both for building training data and for answering live requests. Feast is a popular open-source one. I use version 0.66.0, installed from the Python package index.

An entity is the thing the features describe. Here it is one customer, named by a customer number. A feature view is a named group of features that Feast reads from one table. A snapshot row is one customer's values at the end of one day. In my notebook story, it is one dated line.

The historical read is a Feast call named get_historical_features. You give it a list of rows, each with a customer and a time, and it returns the feature values as they were at each row's own time. To materialize means to copy the newest values into a small, fast store, called the online store. The online read, get_online_features, asks that fast store for the newest values, as a live service would. means time to live: how old a value may be and still count.

What I Ran, and Where

Everything in this lesson ran on one laptop, with no server and no cloud account.

Six rows, each with a logo where the logo set has one, titled one feature store, five tools underneath. Feast 0.66.0: the definitions, the point-in-time join, materialize and the online read. Parquet files: the two snapshot tables, 38,502 rows each, read through a FileSource. Dask, as Feast's file offline store: the class Feast loaded, DaskOfflineStore, does the historical join in Python. SQLite, the online store: the class Feast loaded, SqliteOnlineStore, newest value per customer. scikit-learn 1.9.1: lesson 1's model, trained on what Feast returned. pandas 2.3.3: merge_asof, the hand-written join it was compared with. Below: Feast has no logo here, the logo set has none.

Feast does not do every job itself. When you tell it to read plain files, it uses a library called Dask to do the historical join in Python. Feast calls this the "file" offline store. The offline store is where the full history lives. For the online store I used SQLite, a small database that lives in one file. My tables were Parquet files, a common file format for tables. The lab recorded which classes Feast actually loaded, so these are not guesses.

Feast needed its own Python environment. A virtual environment is a separate folder of Python packages, so two projects cannot break each other. Feast 0.66.0 requires pandas older than version 3, while the rest of this chapter used pandas 3.0.6. Feast's own tests run on Python 3.10, 3.11 and 3.12, so I used Python 3.12.14. The lab saved the exact list of every package and version into its results file.

Feast's own page for the Dask store warns that it "may not scale to production workloads". It is meant for trying Feast on one machine, which is what I did.

Two Reads From One Definition

Here is the path a feature value takes in this lab, with the Feast calls I made.

A flowchart. Raw invoice lines lead to a snapshot table in parquet, which leads to the feature view customer_daily. From the feature view, get_historical_features leads to training rows, as of each cutoff. Also from the feature view, materialize leads to the SQLite online store, and from there get_online_features leads to a live request, the newest value. Below: the historical read answers as of this time for every row; the online read only knows the newest value it was given; both read the same feature view, so the definition is written once.

I built the snapshot table from the raw invoice lines myself, with pandas, and wrote it to a Parquet file. Feast does not compute these running totals for me. I only told Feast where the file is, which column is the customer and which column is the time.

From there, two reads use the same feature view. The historical read builds training rows: for each customer and each cutoff, the values as they were at that cutoff. The online read serves a live request: for each customer, the newest value that was copied in.

This split is the main promise of a . You describe the features once, and both training and serving read from that one description. The survey lesson explains why that matters. My job here is to check that the historical read really is a correct point-in-time join.

The Table I Gave Feast

Feast needs a table with a time column. I built it the way a nightly batch job would, the same way as lesson 2.

For every customer, and every calendar day on which that customer has at least one invoice line, the table has one row. The row holds running values up to the end of that day. They are the time of the first line and of the last line, the number of purchase invoices, and the net money. It also holds the share of lines that were returns, and the number of different products bought. The row is stamped at the next midnight, because that is the moment the night's job has seen the whole day.

Recency and tenure change every day, even when nothing happens. So I did not store them. I stored the times of the first and last lines instead. After the join, I computed recency as the cutoff minus the last line, in days. Tenure is the cutoff minus the first line. Lesson 1 computes them the same way. The table has 38,502 rows for 5,942 customers.

These are the definitions I gave Feast. They live in scripts/labs/features/feast_repo/definitions.py, next to a short feature_store.yaml that names the SQLite online store and the file offline store.

"""Lesson 12 of 'Features and Feature Stores': what Feast is told about the data (Feast 0.66.0).

One entity (a customer) and three feature views over two parquet snapshot tables that feast_lab.py writes to
$FEAST_DATA (default ~/lab-data/features/feast-repo):

  customer_daily.parquet  lesson 1's six features, as running totals. One row per customer per calendar day
                          on which they have any invoice line, stamped at the NEXT midnight (the moment a nightly
                          job has seen the whole day). Two date columns are stored as whole seconds since
                          1970-01-01 so recency and tenure can be worked out at any cutoff after the join.
  customer_l2.parquet     lesson 2's feature table, unchanged (point_in_time.feature_table), same stamps.

ttl = timedelta(0) means "no limit": Feast looks back as far as the table goes. customer_daily_ttl90 is the same
table with a 90-day TTL, to see what TTL does.

Author: Roni Das
Created: 2026-10-01
"""
import os
from datetime import timedelta
from pathlib import Path

from feast import Entity, FeatureView, Field, FileSource
from feast.types import Float64, Int64
from feast.value_type import ValueType

DATA = Path(os.environ.get("FEAST_DATA", Path.home() / "lab-data" / "features" / "feast-repo"))

customer = Entity(name="customer", join_keys=["customer_id"], value_type=ValueType.INT64,
                  description="One shop customer (CustomerID in UCI Online Retail II)")

daily_source = FileSource(name="customer_daily_source", path=str(DATA / "customer_daily.parquet"),
                          timestamp_field="feature_ts")
l2_source = FileSource(name="customer_l2_source", path=str(DATA / "customer_l2.parquet"),
                       timestamp_field="feature_ts")

DAILY_FIELDS = [
    Field(name="snap_unix", dtype=Int64),          # the row's own stamp, so I can see WHICH row was joined
    Field(name="first_event_unix", dtype=Int64),   # first invoice line so far -> tenure_days
    Field(name="last_event_unix", dtype=Int64),    # last invoice line so far -> recency_days
    Field(name="frequency", dtype=Int64),          # distinct purchase invoices so far
    Field(name="money", dtype=Float64),            # net spend so far, returns included (negative)
    Field(name="return_share", dtype=Float64),     # return lines / all lines so far
    Field(name="products", dtype=Int64),           # distinct stock codes on purchase lines so far
]

customer_daily = FeatureView(name="customer_daily", entities=[customer], ttl=timedelta(0),
                             schema=DAILY_FIELDS, source=daily_source, online=True)

customer_daily_ttl90 = FeatureView(name="customer_daily_ttl90", entities=[customer], ttl=timedelta(days=90),
                                   schema=DAILY_FIELDS, source=daily_source, online=True)

customer_l2 = FeatureView(
    name="customer_l2", entities=[customer], ttl=timedelta(0), source=l2_source, online=False,
    schema=[Field(name="snap_unix", dtype=Int64), Field(name="frequency", dtype=Int64),
            Field(name="money", dtype=Float64), Field(name="return_share", dtype=Float64),
            Field(name="last_buy_unix", dtype=Float64)])   # empty (NaN) while a customer has only returned

OBJECTS = [customer, daily_source, l2_source, customer_daily, customer_daily_ttl90, customer_l2]

Which Row Is 'As Of 1 July'?

Here is one real customer, number 12395, and three of their 17 snapshot rows.

A hand-drawn sketch of three snapshot rows for customer 12395 and a cutoff box below them. Left: stamped 2011-05-13, frequency 9, £3,254. Middle: stamped 2011-07-01, frequency 10, £3,418. Right: stamped 2011-08-20, frequency 11, £3,584. Below them: cutoff 2011-07-01 00:00. Text: the middle row covers 30 June and is stamped at midnight, exactly at the cutoff; the right one is from the future; Feast returned the middle row, frequency 10, money £3,418.36. A row stamped exactly at the cutoff counts; lesson 2's merge_asof chose the same.

This customer bought something on 30 June 2011. The nightly job wrote a row for that day and stamped it at midnight, which is 00:00 on 1 July. The cutoff is also 00:00 on 1 July. So the row's stamp and the cutoff are exactly the same moment.

Should the join take that row? Yes. It covers 30 June, so it holds nothing from 1 July onward. It is the freshest row that is still legal. Lesson 2 thought about this carefully, and chose allow_exact_matches=True in merge_asof so that a row stamped exactly at the cutoff counts.

The row on the right, stamped 20 August, is from the future and must never be used for a 1 July cutoff. Feast returned the middle row: frequency 10 and money £3,418.36. Recency came out as 0.297 days, about seven hours, because the last line on 30 June was at 16:52.

What Feast's Code Says

Before running anything, I read how Feast's file store does the join. I cloned Feast's code at the tag for version 0.66.0, commit 1d5be950, and read the file dask.py. Every quote is in results/feast-factcheck.json, with the line numbers.

A sequence diagram with four lifelines: my code, FeatureStore, Dask store and parquet. Step 1, my code sends 86,045 rows to FeatureStore. Step 2, FeatureStore asks the Dask store to join customer_daily. Step 3, the Dask store reads every row of the parquet file. Step 4, a self step on the Dask store: stamp at or before cutoff. Step 5, another self step: newest one wins. Step 6, a dashed return to my code: one row per entity row. Below: with a TTL, step 4 also needs stamp at or after cutoff minus TTL; the same rule as merge_asof with allow_exact_matches=True.

The code does three things. First it joins every one of a customer's snapshot rows onto each of that customer's entity rows. An entity row is one row of the list I hand to Feast: a customer and a time. Then it keeps only the snapshot rows whose stamp is less than or equal to the entity row's time. The code says <=, and Feast's documentation says the time is the "upper bound (inclusive)". Then it sorts by stamp and keeps the newest one.

With a , there is one more condition. The stamp must also be at or after the entity time minus the TTL. A TTL of zero means no limit. Feast's own description says a TTL of 0 means the features live "forever". And Feast treats a time with no time zone as UTC, on both sides of the join, so my plain shop times did not shift.

One more detail mattered later. The step that keeps rows only lets through rows inside the time window, or rows where the customer had no snapshot at all. So a customer who has snapshots, but none inside the window, is left with nothing. I wrote in the lab's design, before it ran, what I expected that to do to rows outside a TTL. After an independent review, I also checked rows from before a customer's first snapshot. The slide on TTL shows both.

How the Lab Was Built

I wrote the lab's design into the docstring of scripts/labs/features/feast_lab.py before it first ran. I had read Feast's code first, as the last slide says, but no Feast call had run on this data.

A page in four labelled zones, headed feast_lab.py, designed before it ran. customer_l2: lesson 2's own table through Feast, compared row by row with lesson 2's merge_asof. customer_daily: lesson 1's six features as daily snapshots, compared with lesson 1's own code; then lesson 1's model trained on them. customer_daily_ttl90: the same table with a 90-day TTL, what is kept, what disappears. Online: materialize up to 2011-07-01; read every customer, plus one not yet a customer and one never. Below: 86,045 entity rows, every customer and cutoff row of the chapter's train, valid and test months.

The rows. I asked Feast for every row the chapter uses: 44,521 training rows, 14,673 validation rows and 26,851 test rows, 86,045 in all. Each row is a customer and a cutoff, the first of a month, as in every lesson.

Three feature views. The first, customer_l2, is lesson 2's own table. I compared Feast's answer with lesson 2's merge_asof on every row and every column. The second, customer_daily, holds lesson 1's six features. I compared it with lesson 1's own code, rebuilt in the same environment, and then trained lesson 1's model on Feast's columns. The third, customer_daily_ttl90, is the same table with a of 90 days.

Equal means equal. A row matches only if every value is identical, to the last bit, with two empty values counted as equal. I did not allow a tolerance, so any difference in how a number was added up would show.

The model is the chapter's own: scikit-learn's HistGradientBoostingClassifier with its default settings and seed 0, trained on the training months and scored by average precision (AP) on the five test months. Lesson 1 stored a test AP of 0.5450 for it.

The Lab's Report, Running

This is a real recording of the report script, feast_report.py, on the laptop where the lab ran, inside the Feast environment.

A terminal recording of feast_report.py. It prints Feast 0.66.0, 809,561 raw lines and 86,045 customer and cutoff rows; that it rebuilt six features from raw lines and a 38,502-row snapshot by a plain loop, whose money differs in the last bits on 28,561 rows with a plain +=, on 28,561 with a plain Series.cumsum and on 0 when compensated like groupby cumsum; that Feast's customer_daily equals the raw-line answer on 86,045 of 86,045 rows; that lesson 2's table through Feast equals lesson 2's merge_asof on 86,045 of 86,045 rows; that 1,194 rows have a snapshot stamped exactly at the cutoff, 768 of them training rows, and Feast took that row every time; that 41,232 rows are older than 90 days and Feast returned 44,813 rows, missing exactly those, while the 232 rows exactly 90 days old were kept, with a buy rate of 0.098 in the dropped rows and 0.317 in the rest; post-review, 5,942 rows dated before each customer's first snapshot with no TTL, of which Feast returned 0, while the docs example returned 4 of 5 with no TTL and 3 of 5 with a 2-hour TTL; test AP 0.5450 and ROC-AUC 0.7973 against lesson 1's stored 0.5450; that 5,106 of 5,106 customers online equal the raw-line answer, customers 12364 and 1 come back None, and the 90-day view online still answered for 3,082 customers; that online latency was skipped because the machine was not quiet; and that all 33 checks agree with the stored lab.

The report does not reuse the lab's code for the parts that matter. It computes lesson 1's six features straight from the raw invoice lines, for every customer and cutoff, with no snapshot table and no join. A correct point-in-time join must give this answer. Then it builds a new snapshot table with a plain Python loop, puts it into a brand new Feast store in a temporary folder, and asks Feast again.

Feast's answer from that new store must equal the raw-line answer on every row. The report also rechecks lesson 2's table, the counts, the model's score and the online answers. If any number differs from the stored run, it stops with an error. It ran 33 checks, including the two added after the review, and all of them agreed.

The Answer: The Same Join, Row for Row

Here is the main result. With no , the historical read of Feast's file store gave exactly the same answer as my hand-written joins.

Three panels, all 86,045 customer and cutoff rows, where equal means every value identical to the last bit. Lesson 2's table: 86,045 of 86,045 rows equal lesson 2's merge_asof. Lesson 1's six: 86,045 of 86,045 rows equal lesson 1's own code. Lesson 1's model: test AP 0.5450; lesson 1 stored 0.5450. Below: the model's test predictions were identical too, to the last bit, so the score could not differ; with no TTL, the file store's historical read and merge_asof were the same join.

Lesson 2's table through Feast equalled lesson 2's merge_asof on 86,045 of 86,045 rows. Every column matched to the last bit: the joined row's stamp, frequency, money, return share, the last purchase time, and the recency computed from it.

Lesson 1's six features through Feast equalled lesson 1's own code on 86,045 of 86,045 rows, on all six columns. Lesson 1's code never builds a snapshot table. It adds up every line before each cutoff, directly. So this checks the whole path at once: my snapshot table, Feast's join, and the recency and tenure I computed after it.

Lesson 1's model, trained on Feast's columns, scored a test AP of 0.5450 and a ROC-AUC of 0.7973. Lesson 1 stored 0.5450. This match was not luck. The model's test predictions were identical to the last bit, so the score could not differ. With the same inputs in the same order, the same model gives the same answer.

So the answer to this lesson's question is yes. On this data, with no TTL, the historical read of Feast's file store is the point-in-time join that lesson 2 wrote by hand. One condition holds on every row here: each customer has a snapshot at or before the cutoff. The TTL slide shows what happens when that fails. A does not add a different join. It adds two things: you no longer write the join yourself, and training and serving read one description.

The Edge Case Lesson 2 Chose by Hand

The most delicate part of a point-in-time join is a row stamped exactly at the cutoff. Here is how often it happened.

A sidebar figure of four numbers. 1,194 of 86,045 rows have a snapshot stamped exactly at the cutoff: the customer had an invoice line on the day before it. 768 of them are training rows, the same count lesson 2 found. 1,194 of them, Feast joined that row. 1,194 rows would change if the edge were excluded, and 139 would get no row at all. Below: Feast's upper bound is inclusive, at or before.

In 1,194 of the 86,045 rows, the customer had an invoice line on the last day before the cutoff, so their newest snapshot is stamped exactly at the cutoff. 768 of those are training rows. Lesson 2 counted the same 768, which the report checks. Feast took that row in all 1,194 cases.

If the edge had been excluded, as allow_exact_matches=False would do, all 1,194 rows would have changed. 139 of them would have had no row at all, because that day was the customer's first. So the difference between "before" and "at or before" matters here. It touches about one row in 72.

Feast's documentation says the entity time is an inclusive upper bound, and the file store's code says <=. My run agrees with both. That only matters because my rows mean "valid from midnight". If your stamps mean the time of the event itself, an inclusive bound could let an event at the cutoff leak in, and lesson 2's advice applies.

What the TTL Did

Now the library rule I did not know well enough: the . I set the third feature view to a TTL of 90 days and asked for the same 86,045 rows.

An isometric row of three blocks for rows asked for and rows returned by get_historical_features. Asked: 86,045, tall. No TTL: 86,045, equally tall. TTL 90 days: 44,813, about half the height. Below: 41,232 entity rows were not in the output at all, train 16,222, valid 8,910, test 16,100. Not empty rows: missing rows. The list came back shorter.

Without a TTL, Feast's file store returned all 86,045 rows. With a 90-day TTL, it returned 44,813. The other 41,232 entity rows were not returned with empty features. They were missing from the output altogether. My list of rows went in with 86,045 entries and came back with 44,813, with no error or warning printed. A program that does not count the rows would simply train on fewer customers and never know.

The dropped rows were exactly the ones the rule predicts: rows whose newest snapshot was more than 90 days older than the cutoff. The report computed that set from the raw lines and checked that it equals the missing set, row for row. The 232 rows whose newest snapshot was exactly 90 days old were kept, so the lower bound is inclusive too. Every row that came back was identical to the answer with no TTL.

A two-column figure titled what the docs example shows, and what the file store did. The docs: the docs example's result table keeps all 5 entity rows, 2 of them with NULL features; its text says rows outside the TTL couldn't be joined. My runs: on the file store, 4 of 5 rows with no TTL and 3 of 5 with a 2-hour TTL; in this shop, 41,232 rows missing under a 90-day TTL, and with no TTL, 5,942 of 5,942 rows dated before a first snapshot.

Feast's page on point-in-time joins shows a small example with one driver and five requests. Its text says a row outside the TTL "couldn't be joined". Its picture of the result keeps all five entity rows, two of them with NULL, meaning empty, features. The file store at version 0.66.0 does not do that. It drops the row. I expected this from reading its code, and wrote the guess down before the run. The left join brings in all of a customer's snapshot rows. The time filter then removes every one of them, so nothing is left to carry the entity row through.

The TTL is not the only cause. An independent review pointed out that the same filter also removes snapshot rows stamped after the entity time. I added a check to the lab after the review, and it is labelled there. The file store drops any entity row whose customer has snapshots but none in the window. The snapshots are either older than the TTL, or, with no TTL, only newer than the cutoff. The no-TTL run lost nothing only because every row in this chapter has history before its cutoff.

Why Half the Rows, and Which Half

A of 90 days sounds generous. Why did it drop almost half the rows?

A line chart titled almost half the rows were older than 90 days, of how old each row's newest snapshot was at its cutoff, all rows: the share of rows at most that old, from 0 to 100 percent, against age from 0 to 730 days, with a dotted line at 90 days. The line climbs quickly to about 52 percent at 90 days, about 92 percent at 365 days, and reaches 100 percent by 730. Below: at most 90 days old, 52.1 percent; at most 365 days, 91.8 percent; every row's newest snapshot was at most 730 days old. In a shop where people buy a few times a year, a 90-day TTL throws away most of what it knows.

The chart shows how old each row's newest snapshot was at its cutoff. I added it to the report after seeing the results, as a description of the data. Only 52.1 percent of rows had a snapshot at most 90 days old. In this shop, many customers buy a few times a year, or once and never again. Their newest row is old, but it is still true: it is still everything they have done.

A TTL is meant for values that go stale, like a sensor reading or a price. A running total over a customer's whole history does not go stale. A customer who last bought 200 days ago still bought 12 times in total. For features like these, a TTL deletes good information.

A hand-drawn bar sketch of the share of rows that bought in the next 30 days. Dropped by the TTL: 9.8 percent. Kept: 31.7 percent. Below: the rows a 90-day TTL drops are mostly quiet customers, who mostly did not buy; training without them changes what the model sees.

The dropped rows are also not random. Of the dropped rows, 9.8 percent bought in the next 30 days. Of the rows kept, 31.7 percent did. A training set built with this TTL would hold far fewer quiet customers, and its buy rate would look much higher than the shop's. I did not train a model on it, so I cannot say what score it would get. I can only say the training rows would be a different, biased set.

Reading Online

Next, the online side. I ran materialize from 1 December 2009 to 1 July 2011 into the SQLite online store, and then asked it about customers.

A list of four customers asked online after materialize up to 2011-07-01. 12395: frequency 10, money £3,418.36, row stamped 2011-07-01, the row for 30 June. 12346: frequency 12, money minus £64.68, row stamped 2011-01-19, old, but the newest it has. 12364: None for every feature; first invoice 2011-08-19, after the cutoff. 1: None for every feature; not a customer in the data at all. Below: all 5,106 customers seen before the cutoff, online equal to the historical answer at the same cutoff, on every value. An unknown customer gets None and no error; lesson 8 read this in the source; here it ran.

The online store keeps only the newest value for each customer. Feast's documentation says so, and the code's materialize step keeps the newest row in the time window. The window included its end, 00:00 on 1 July, so customer 12395 got the row for 30 June, the same row the historical read chose.

I asked for all 5,106 customers seen before 1 July 2011. Every one of them got exactly the same seven values online as the historical read gave at that cutoff. So online and offline agreed. Lesson 5 found that two separate programs did not quite agree; here it holds because one program wrote both sides.

Then I asked about two customers the store cannot know. Customer 12364 first appears in the data on 19 August 2011, after the end of the copied window. Customer 1 does not exist in the data. Both came back with None for every feature, and an event time of zero. There was no error and no warning. Lesson 8 read this behaviour in Feast's source. Now I have seen it run. Your serving code must check for None itself, and lesson 8 measured what each way of filling the gap costs.

The Online Read Ignored the TTL

Then I asked the 90-day view online, for the same 5,106 customers.

Three panels for the 90-day view, the same 5,106 customers at 2011-07-01. Historical read: 2,024 customers with a value. Online read: 5,106 customers with a value. The gap: 3,082, answered online, dropped in training. Below: Feast's SQLite page lists TTL at retrieval as not supported; the view's online answers were the no-TTL answers for all of them. One setting, two meanings: training dropped these customers, serving did not.

At the 1 July cutoff, the historical read with a 90-day had values for only 2,024 of these customers. The online read of the same view answered for all 5,106. For 3,082 customers, the store served a value online that the same view's historical read would have dropped from training. The online answers were identical to the view with no TTL.

Feast's own page for the SQLite online store has a table of features, and one row says "support for ttl (time to live) at retrieval: no". So this is documented. But a general page about feature retrieval says event times are used "to ensure that old feature values aren't served to models during online serving". For SQLite in version 0.66.0, that did not happen in my run.

So one setting meant two different things. In training, the TTL removed these customers. In serving, it did nothing. A model trained with this view would never see a customer whose newest row is older than 90 days. Then, at serving, it would be asked about thousands of them. This is exactly the kind of training and serving mismatch that lesson 5 measured, made here by a single setting.

Why Money Matched to the Last Bit

One result surprised me: money matched lesson 1 exactly, on every row. Money is a sum of many decimal amounts, and lesson 5 showed that two ways of adding the same amounts can differ in the last digits.

A sidebar figure, found after the results while writing the report, titled why money matched to the last bit. 28,561 of 38,502 snapshot rows: money from a plain Python += loop differs in the last bits. 28,561 differ with a plain Series.cumsum() per customer, checked after the review. 0 differ when the loop compensates for rounding, as pandas' groupby cumsum and groupby sum do, Kahan summation. 86,045 rows equal lesson 1, because both sides used those compensated groupby sums. Below: exact equality held because both sides added the same compensated way, not by luck.

I found the reason after the results, while writing the report, and the report labels it so. My first version of the report built the snapshot table with a plain Python loop that added each amount with +=. Its money differed from the lab's table in the last bits on 28,561 of 38,502 rows, and the report stopped.

The lab built its table with pandas' groupby running sum, a groupby cumsum, and lesson 1 adds with pandas' groupby sum. In pandas' source code, both use Kahan summation: a careful way of adding that keeps track of the small rounding error at each step and puts it back. When I changed the report's loop to add the same careful way, the count went to 0. Not every pandas sum does this. After the review I checked that a plain Series.cumsum(), run on each customer's amounts, does not compensate: it differed on the same 28,561 rows as the plain loop.

So the exact match was not luck, and it was not Feast. Feast only moved values that were already in the table. It held because both sides added in the same compensated way. If your training side and your table builder add in different ways, expect last-digit differences, and compare with a small tolerance, as lesson 5 advised.

Online Speed: Not Measured

The plan for this lesson included one more number: how long an online read takes on this laptop. I did not measure it, and here is why.

A timing is only honest when the machine is not busy with something else. Other work slows every step, by an amount you cannot know. So I wrote a rule into the lab before running it: measure only if the laptop's one-minute load average is below 2.0. The load average is roughly how many programs were waiting to run, on average, over the last minute. This laptop has 10 cores.

When I ran the timing step, the load average was 21.17. Other jobs were running on the same laptop. So the lab wrote "skipped: machine not quiet" into results/feast-latency.json and stopped. I do not quote the few seconds the historical reads took, for the same reason.

If you want a number, run the demo on a quiet machine and time it yourself. Even then, a time from one laptop with SQLite says little about a production online store on a server, which is a different system.

My Guesses Before the Run, Checked

I wrote my expectations into the lab's docstring before it ran. I had already read Feast's code, so these were informed guesses, not blind ones. Here they are, quoted, against the results.

  1. "(a) equals lesson 2's merge_asof on all 86,045 rows." Right: 86,045 of 86,045.

  2. "(b)'s join equals merge_asof on the snapshot exactly." Right, on every row and every column.

  3. "the snapshot's money can differ from lesson 1's in the last bits ... so the AP may differ from 0.5450 in the 4th decimal or not at all." Wrong about the money: it did not differ at all, for the reason on the last slide but one. The AP was exactly 0.5450.

  4. "I expect rows older than the to be DROPPED from the output, not returned empty." Right: 41,232 rows were missing. I only knew this from reading the code. The documentation's example shows the opposite. The review later showed the cause is wider than the TTL, as the TTL slide explains.

  5. "Online equals historical at 2011-07-01; the unknown customers get None for every feature." Right, for all 5,106 customers, and for both unknown ones.

  6. "(c) online returns values the TTL join would not." Right: 3,082 customers.

Five right out of six sounds good, but please read it carefully. I read the source code first, so I was largely predicting what the code says. These guesses show the code does what it says. They are not independent evidence about anything else.

Try It Yourself

The full lab asks Feast for all 86,045 rows, three times. I wrote a smaller demo that builds its own Feast store in a temporary folder and runs the main checks.

A page in three labelled zones, headed feast_demo.py, designed before it ran. The store: one entity, one parquet file, two feature views, no TTL and 90 days, SQLite, in a temporary folder. The checks: every test row against merge_asof; lesson 1's model at seed 0; rows the 90-day view returns. Online: materialize up to 2011-07-01; ask for 12395, 12346, 12364 and 1. Below: it had to match the lab, 26,851 of 26,851 test rows equal, test AP 0.5450, 10,751 rows from the 90-day view.

I wrote the demo's design into its docstring after the lab had run and before the demo first ran. It builds its snapshot table with its own short code, so it is also a second check of the lab.

A real screenshot of VS Code with feast_demo.py open at the top of the file, showing its docstring: what it needs, how to make the Feast environment and run it, and the design written before it first ran.

Before you run this lab. Feast needs its own environment. Feast's tested range is Python 3.10 to 3.12. On a Mac or Linux, make one and turn it on with python3.12 -m venv ~/venv-feast and then source ~/venv-feast/bin/activate. On Windows, the activate script is in the environment's Scripts folder instead. Then install the pinned versions: pip install feast==0.66.0 scikit-learn==1.9.1 openpyxl. openpyxl is only needed by fetch_data.py, to read the shop's Excel file. First run python fetch_data.py from the scripts/labs/features folder. It downloads the shop data once, about 46 MB, and writes one cleaned file.

The demo imports task.py from that same folder. It needs no GPU and no server. I ran it only on a Mac, with Python 3.12, so I have not checked it on Windows or Linux. Give it a file name, , and it also saves every number. That is how was made.

Run the Join Yourself

This box holds two real customers' snapshot tables, exactly as Feast stored them. Customer 12395 has a row stamped exactly at the cutoff. Customer 12346's newest row before July 2011 is from January. It does Feast's join rule in plain Python, so it runs in your browser with nothing installed.

Press Run. It shows which row the join picks at the cutoff, and what Feast itself returned. Then try three things. Set TTL_DAYS = 90 and CUSTOMER = 12346: the row disappears, as it did in Feast. Set INCLUSIVE = False: customer 12395 falls back to an older row. Change CUTOFF to another date and watch the chosen row move.

The report script writes this box from the lab's stored results. Then it runs the box for both customers, with and without the , and checks that each answer equals what Feast returned. The rule in the box is short, which is the point: a point-in-time join with a TTL is only two comparisons and "newest wins".

The Lab's Code, Piece by Piece

The lab is one file, scripts/labs/features/feast_lab.py, and it runs in the Feast environment. It imports lesson 1's and lesson 2's lab files instead of copying them, so the comparisons are against their real code.

daily_snapshot builds lesson 1's six features as running values. It sorts the lines by customer, keeping time order, and uses pandas' running sums and counts. Two counts need care: an invoice or a product counts only the first time it appears. Then it keeps the last line of each customer's day and stamps it at the next midnight. l2_table calls lesson 2's own feature_table and adds whole-second versions of its times.

feast_hist hands the entity rows to get_historical_features, with a row number, so the answer can be matched back. It records how many rows went in and how many came back. That count is how the missing rows were found. asof is the same join done with merge_asof, for the comparison, and compare counts exactly equal values, column by column.

main applies the three feature views, runs the three historical reads and the comparisons, trains the model, and then runs materialize and the online reads. latency is the timing step, which checks the load average first and skipped itself. The report, , has its own code for every check, as the report slide describes.

Building Training Rows With Feast, Step by Step

These are the steps I would follow to build training rows with Feast, or with any .

  1. Stamp each row with the moment it became true. For a nightly job, that is the midnight after the day it covers. Then an inclusive join is correct. Write down what your stamps mean.

  2. Make the entity table yourself, with a row number. One row per customer and cutoff, with the cutoff in a column called event_timestamp. Feast sorts the rows by time inside the join, so keep a row number to match them back. Do not send the same customer and time twice: the file store keeps one answer per pair.

  3. Count the rows that come back. If fewer come back than you sent, find out why before you train. Here a removed 41,232 rows with no error. Even with no TTL, a row dated before a customer's first snapshot disappears from the file store's output, so new customers need their own handling.

  4. Choose the TTL on purpose. A TTL of zero means no limit. Use a real TTL only for values that go stale. For running totals over a customer's whole history, a TTL deletes true information.

  5. Compare the store's answer with a join you wrote yourself, once. A small merge_asof on a sample is enough. Here the two matched on every row, and now I trust it.

  6. Read the same customer online and offline. After materializing, check that the online answer equals the historical answer at the same moment, and decide what your service does with None.

When to Use a Feature Store, and When Not To

Use a like Feast when several models or services read the same features. It also helps when the team that builds features is not the team that trains models. Then one description for both training and serving is worth a lot, as lesson 5 showed when two programs disagreed. It also helps when people keep writing the point-in-time join by hand and getting it wrong, as lesson 2 showed can happen.

A hand-written join is fine for a one-person project where one script builds the features and trains the model together. Here, Feast and merge_asof gave the same rows. The join itself was not the hard part. Getting the stamps right, which lesson 2 did, was the hard part, and Feast does not do that for you.

Do not expect a feature store to compute your features. In this lab, I built every running total myself and gave Feast a finished table. Feast joined and served it. Feast can also run Python code at read time, in on-demand feature views, which lesson 5 described, but I did not use them here.

Be careful with any setting that has a different effect online and offline. The in this lab is the example: it removed customers from training, and did nothing at serving.

What This Lab Cannot Tell You

Two columns titled what this lab shows, and what it cannot. Shows: Feast 0.66.0's file store joins exactly like merge_asof, when each row has a snapshot in its window; it drops entity rows with no snapshot in the window, TTL or not; its SQLite online read ignores the TTL. Cannot show: other offline stores, like DuckDB, BigQuery or Snowflake, which I did not run; online speed, because the laptop was busy, so I did not measure it; other Feast versions, where behaviour can change.

One version and one pair of stores. Everything here is Feast 0.66.0, with the file offline store and the SQLite online store. Other offline stores use different code, and a different version could behave differently. The dropped rows, with or without a , and the ignored TTL online are facts about this setup only.

One shop and one kind of feature. My features were running totals, stamped once a day. A store with many updates per second, or features that really do go stale, would test other parts of Feast. I did not test them.

No speed numbers. The laptop was busy, so I measured no timing, and quote none.

Guesses informed by the code. I read the source before the run, so my right guesses confirm the code, not something beyond it. Two parts were added after the results, and both are labelled: the reason money matched, and the chart of row ages. Two more checks were added after an independent review: rows from before a first snapshot with Feast's docs example, and which pandas sums compensate.

The Whole Chapter on One Page

This is the last lesson of the chapter, so here is all of it on one page. Every number in the two figures was read from that lesson's own results file by the figure script, and I copied them into the text below.

The chapter on one page, part one, lessons 1 to 6, each number from that lesson's results file. 1, what a feature is: six made features, test AP 0.5450; the last invoice line's raw columns, 0.3929. 2, point-in-time joins: a latest-value join promised 0.714 offline and gave 0.301 on later months; point-in-time 0.522. 3, feature freshness: 14 days stale +0.0011, 30 days -0.0037, neither interval clear of zero. 4, window aggregations: all windows 0.556, the six 0.545, the six plus windows 0.561. 5, online and offline consistency: the same choices in two programs, 17,670 test rows differed, in the last digits of money. 6, late events, simulated: heavy delays with late events dropped, -0.0038 with no wait, -0.0013 waiting 48 hours. Below: test AP on the five test months; seed 0 for lessons 1 and 4 and lesson 2's later months; means over 20 seeds for lesson 2's offline score and lessons 3 and 6.

Lesson 1 asked whether a better model or a better feature is worth more. Six simple features made from each customer's history scored a test AP of 0.5450, against 0.3929 for the raw columns of their last invoice line. Tuning the model on the raw columns added almost nothing. Lesson 2 showed that joining each customer's latest values promised 0.714 offline and gave 0.301 on later months, while the correct point-in-time join gave 0.522.

Lesson 3 served features up to 30 days stale. At 14 days the change was +0.0011 and at 30 days -0.0037, and neither could be told apart from luck. Lesson 4 tried windows of days: all of them together scored 0.556, the six features 0.545, and both together 0.561. Lesson 5 wrote each feature twice, in batch and event by event, and found 17,670 test rows that still differed, only in the last digits of money. Lesson 6 simulated late events. When heavy delays meant late events were dropped, AP fell by 0.0038 with no wait, and by 0.0013 when the pipeline waited 48 hours.

The chapter on one page, part two, lessons 7 to 12. 7, categorical features: hashing into 1,024 buckets +0.0024; target encoding with row folds -0.0064. 8, missing at serving: 20 percent of lookups lost whole, filled with NaN, -0.0529. 9, versioning and backfill: money retrained on a new rule -0.0012, frequency +0.0040; a drift alarm against the training months fired 5 of 5 test months with no definition change. 10, feature cost and selection: cut lists of 5 to 12 columns 0.5612, all 21 columns 0.5594, interval -0.0006 to +0.0045; lines read 2.80M against 6.66M. 11, embeddings as features: with the six, no real embedding beat random directions over ten random tables, neural with 64 numbers -0.0018 to +0.0015; random directions beat pure noise, +0.0010 to +0.0036. 12, Feast, this lesson: the file store, 86,045 of 86,045 rows equal; rows with no snapshot in the window dropped, 41,232 under a 90-day TTL. Below: lessons 7 to 11, test AP, its change, or a 95% interval of the change, from means over 20 seeds; same shop, task and model.

What to Do on Monday

A hand-drawn grid of six cards, titled five checks, before you trust any feature store's join. 1, stamp the end of day: a row valid from midnight after the day it covers. 2, compare row by row: your store's training rows against a join you wrote yourself. 3, count the rows back: rows with no snapshot in the window, or before the first one, vanish. 4, choose the TTL on purpose: 0 means no limit; 90 days dropped almost half the rows here. 5, read online and offline: the same customer, the same moment, both reads. The reason: a store does what its code says, which is not always what you assumed. Below: trust the join after you have compared it, not before.

If you already use a , here is one thing to do on Monday. Take one training set it built, and count its rows. Compare the count with the number of entity rows you sent. If they differ, find out which rows went missing and why, before your next training run.

If you are about to adopt one, write a small merge_asof join for one month of data and compare it with the store's historical read, row by row. In this lab it took a few lines, and it turned "I hope the store is right" into "I checked, on every row".

A closing card titled the file store did the join right, the settings are yours. Three numbers in large type: 86,045 of 86,045, rows where Feast equalled lesson 2's hand-written point-in-time join; 41,232, rows a 90-day TTL removed from the training set, with no error or warning; 3,082, customers the same view still answered online.

The one idea to keep: a good feature store does the point-in-time join correctly. Feast's file store did it on every row that had a snapshot in its window. What it cannot do is choose your settings. A stamp, a or a missing customer can still change your training data quietly. Compare once, count the rows, and read both sides.

Knowledge Check

Knowledge Check

4 questions - Score 80% to pass

Q1

With no TTL, how did Feast's historical read compare with lesson 2's merge_asof in this lab?

Q2

In the file store, what happened to the entity rows whose newest snapshot was older than the 90-day TTL?

Q3

The same 90-day view was read online for the 5,106 customers at 1 July 2011. What did the lab find?

Q4

Why did money match lesson 1 exactly on every row, even though it is a long sum of decimals?

The second table, customer_l2, is lesson 2's own feature table, built by lesson 2's own code and not changed at all. That lets me compare Feast with lesson 2 directly.

To check it, I asked, with no TTL, for one row per customer dated the day before their first snapshot. 5,942 of 5,942 were missing. I also ran Feast's own docs example on the file store. With no TTL it returned 4 of 5 rows: the request from before the driver's first row was missing. With a 2-hour TTL it returned 3 of 5. So, contrary to the docs example, the file store at 0.66.0 drops these rows instead of returning them empty.

I only ran the file store. Feast has other offline stores, for DuckDB, BigQuery, Snowflake and more, and they use different code. I did not test them, so I make no claim about them.

python feast_demo.py out.json
results/feast-demo.json
"""Does Feast do the point-in-time join we wrote by hand?

Lesson 12 of 'Features and Feature Stores'. It needs Python 3.10 to
3.12 with Feast 0.66.0 and scikit-learn, in their own environment:
    python3.12 -m venv ~/venv-feast
    source ~/venv-feast/bin/activate
    pip install feast==0.66.0 scikit-learn openpyxl
and the shop data: run fetch_data.py once first (it downloads UCI
Online Retail II, about 46 MB). Then, inside this folder:
    python feast_demo.py            # print the results
    python feast_demo.py out.json   # and save every number
It makes a throwaway Feast store in a temporary folder and prints
no timings.

Design, written 2026-10-01 after the lab (feast_lab.py) had run and
before this file first ran:
  Table: lesson 1's six features as running totals, one row per
  customer per day with any invoice line, stamped the next midnight.
  Feast: one entity (customer), one file source, two feature views
  over it, TTL 0 (no limit) and TTL 90 days. SQLite online store.
  It asks Feast for the train and test rows (customer, cutoff) and
  checks every test row against pandas merge_asof on the same table;
  trains lesson 1's model (seed 0) on Feast's columns; counts the test
  rows the 90-day view returns; materializes up to 2011-07-01 and
  reads four customers online: 12395 (a row stamped exactly at the
  cutoff), 12346 (newest row older than 90 days), 12364 (not a
  customer yet) and 1 (never a customer).
  It must agree with the lab: test AP 0.5450, 10,751 rows from the
  90-day view, None for 12364 and 1. feast_report.py demo checks this.

Author: Roni Das
Created: 2026-10-01
"""
import json
import sys
import tempfile
import warnings
from datetime import datetime, timedelta
from pathlib import Path

import feast
import numpy as np
import pandas as pd
from feast import Entity, FeatureStore, FeatureView, Field, FileSource
from feast.types import Float64, Int64
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.metrics import average_precision_score

warnings.filterwarnings("ignore")
sys.path.insert(0, str(Path(__file__).resolve().parent.parent))
import task  # noqa: E402

COLS = ["recency_days", "frequency", "money", "return_share",
        "tenure_days", "products"]
STORED = ["snap", "first", "last", "frequency", "money",
          "return_share", "products"]


def secs(s):
    return (s - pd.Timestamp("1970-01-01")) // pd.Timedelta(seconds=1)


def snapshot(ev):
    """Running totals per customer, one row per active day."""
    e = ev.sort_values("customer_id", kind="stable").copy()
    g = e.groupby("customer_id", sort=False)
    buy = ~e["is_return"]
    e["money"] = g["amount"].cumsum()
    e["return_share"] = g["is_return"].cumsum() / (g.cumcount() + 1)
    new_inv = buy & ~e.duplicated(["customer_id", "invoice"])
    new_prod = buy & ~e.assign(b=buy).duplicated(
        ["customer_id", "stock_code", "b"])
    e["frequency"] = new_inv.astype(int).groupby(e["customer_id"]).cumsum()
    e["products"] = new_prod.astype(int).groupby(e["customer_id"]).cumsum()
    e["first"] = secs(g["ts"].transform("min"))
    e["last"] = secs(e["ts"])
    e["feature_ts"] = e["ts"].dt.normalize() + pd.Timedelta(days=1)
    t = e.groupby(["customer_id", "feature_ts"]).tail(1).copy()
    t["snap"] = secs(t["feature_ts"])
    return t[["customer_id", "feature_ts"] + STORED]


def derive(d):
    t = secs(d["cutoff"])
    d["recency_days"] = (t - d["last"]) / 86400
    d["tenure_days"] = (t - d["first"]) / 86400
    return d


ev = task.load_events(check=False)
train, _, test = task.splits(ev)
rows = pd.concat([train, test], ignore_index=True)
table = snapshot(ev)

folder = Path(tempfile.mkdtemp())
table.to_parquet(folder / "daily.parquet", index=False)
(folder / "feature_store.yaml").write_text(
    f"project: demo\nprovider: local\nregistry: {folder}/registry.db\n"
    f"online_store:\n  type: sqlite\n  path: {folder}/online.db\n"
    "offline_store:\n  type: file\nentity_key_serialization_version: 3\n")
customer = Entity(name="customer", join_keys=["customer_id"])
source = FileSource(path=str(folder / "daily.parquet"),
                    timestamp_field="feature_ts")
schema = [Field(name=c, dtype=Float64 if c in ("money", "return_share")
                else Int64) for c in STORED]
views = {ttl: FeatureView(name=f"daily_ttl{ttl}", entities=[customer],
                          ttl=timedelta(days=ttl), schema=schema,
                          source=source) for ttl in (0, 90)}
store = FeatureStore(repo_path=str(folder))
store.apply([customer, source, *views.values()])


def historical(ttl):
    ent = rows.rename(columns={"cutoff": "event_timestamp"})
    out = store.get_historical_features(
        entity_df=ent, features=[f"daily_ttl{ttl}:{c}" for c in STORED]
    ).to_df().rename(columns={"event_timestamp": "cutoff"})
    out["cutoff"] = out["cutoff"].dt.tz_localize(None)
    return out


got = derive(historical(0))
got = rows.merge(got, on=["customer_id", "cutoff", "label"], how="left")
hand = derive(pd.merge_asof(
    test.sort_values("cutoff"), table.sort_values("feature_ts"),
    left_on="cutoff", right_on="feature_ts", by="customer_id"))
mine = test.merge(got, on=["customer_id", "cutoff", "label"])
mine = mine.merge(hand, on=["customer_id", "cutoff", "label"],
                  suffixes=("", "_hand"))
equal = np.ones(len(mine), dtype=bool)
for c in COLS:
    equal &= (mine[c] == mine[c + "_hand"]).to_numpy()
print(f"Feast {feast.__version__}: {len(table):,} snapshot rows")
print(f"test rows equal to merge_asof on all six columns: "
      f"{equal.sum():,} of {len(mine):,}")

is_test = got["cutoff"] >= test["cutoff"].min()
model = HistGradientBoostingClassifier(random_state=0).fit(
    got.loc[~is_test, COLS], got.loc[~is_test, "label"])
part = got[is_test].assign(p=model.predict_proba(got.loc[is_test, COLS])[:, 1])
ap = float(np.mean([average_precision_score(g["label"], g["p"])
                    for _, g in part.groupby("cutoff")]))
print(f"lesson 1's model on Feast's columns: test AP {ap:.4f}")

ttl90 = historical(90)
n90 = int((ttl90["cutoff"] >= test["cutoff"].min()).sum())
print(f"the 90-day view returned {n90:,} of {len(test):,} test rows")

store.materialize(start_date=datetime(2009, 12, 1),
                  end_date=datetime(2011, 7, 1), feature_views=["daily_ttl0"])
ask = [12395, 12346, 12364, 1]
answer = store.get_online_features(
    features=[f"daily_ttl0:{c}" for c in ("frequency", "money", "snap")],
    entity_rows=[{"customer_id": c} for c in ask]).to_dict()
online = {}
print("online, materialized up to 2011-07-01:")
for i, c in enumerate(ask):
    online[str(c)] = {k: answer[k][i] for k in ("frequency", "money", "snap")}
    snap = answer["snap"][i]
    when = "" if snap is None else f", row stamped {pd.Timestamp(snap, unit='s').date()}"
    print(f"  customer {c}: frequency {answer['frequency'][i]}, "
          f"money {answer['money'][i]}{when}")

if len(sys.argv) > 1:
    json.dump({"feast": feast.__version__, "test_rows": len(mine),
               "rows_equal_to_asof": int(equal.sum()), "test_ap": ap,
               "ttl90_test_rows_returned": n90, "online": online},
              open(sys.argv[1], "w"), indent=1)

This is a real run in VS Code's terminal: python feast_demo.py, run inside the examples folder, with the Feast environment turned on.

A real screenshot of VS Code's terminal after running python feast_demo.py inside the examples folder, with the Feast environment active. It prints Feast 0.66.0 and 38,502 snapshot rows; 26,851 of 26,851 test rows equal to merge_asof on all six columns; lesson 1's model on Feast's columns, test AP 0.5450; the 90-day view returned 10,751 of 26,851 test rows; Feast's own materialize message; then the online answers for customers 12395 and 12346, and None for 12364 and 1.

When I ran it, every number matched the lab. 26,851 of 26,851 test rows were equal to merge_asof, and the test AP was 0.5450 to every printed digit. The 90-day view returned 10,751 test rows, which is 26,851 minus the lab's 16,100 dropped test rows. The report script checks this from the stored files with python feast_report.py demo.

feast_report.py

Lesson 7 turned 4,646 product codes into numbers. Hashing into 1,024 buckets added 0.0024, and target encoding with folds of rows lost 0.0064, because a customer's other months leaked in. Lesson 8 lost whole lookups at serving. With 20 percent lost and filled with NaN, AP fell by 0.0529. Lesson 9 changed two definitions. Retrained on the new rules, money lost 0.0012 and frequency gained 0.0040. A drift alarm measured against the training months fired in all 5 test months with no definition change.

Lesson 10 priced each feature in event lines read. Cut lists of 5 to 12 columns, picked on the validation months for each seed, scored 0.5612 against 0.5594 for all 21 columns. The interval of the difference ran from -0.0006 to +0.0045, so the two cannot be told apart. The cut lists read about two fifths of the lines: 2.80 million against 6.66 million per cutoff.

Lesson 11 added product-name to lesson 1's six features, from a neural model and from word counts, at 64 and 128 numbers. None of them beat random directions, checked over ten different random tables: every interval of real minus random crossed zero. For the neural model at 64 numbers it ran from -0.0018 to +0.0015.

Random directions did beat columns of pure noise, by +0.0010 to +0.0036 at 64 numbers. So knowing which products a customer bought helped a little, but the meaning of their names added nothing I could measure. And this lesson found that Feast's file store does the same join as lesson 2, on all 86,045 rows. That holds as long as each row has a snapshot in its window. A 90-day quietly removed 41,232 of them.

If I had to keep one sentence from the whole chapter, it would be this, and it is my reading of the numbers above. Most of the score came from simple features built only from the past. The biggest losses came from how and when values were joined, not from the model.