This idea carries a full system design question on its own. Each walks through the full answer.
It is 2 a.m. Fraud losses are spiking. The team pulls up the model that scores every transaction, and someone asks a simple question: which version is this? The answer that comes back is the one you never want to hear. Nobody is sure.
Was it the model trained last Tuesday, or the hotfix from Thursday? Is it model_final.pkl in the shared bucket, or model_final_v2_really.pkl that someone uploaded during the last incident? What data trained it? Which git commit? Everyone starts grepping through object storage and scrolling old chat messages while money leaks out the door. The engineer who trained the current model is asleep, and the file that is supposedly live has a timestamp nobody can explain.
This is the situation a exists to end. A team ships model after model, each as a loose file saved somewhere, and after a few months nobody can answer the three questions that matter most in an incident:
When you cannot say with certainty which model is live, you do not have a production system. You have a liability that happens to make predictions.
The failure here is not a lack of talent. It is a lack of a system of record. Every other part of the stack already has one. Your code lives in git, where every change is a commit you can name, review, and revert. Your infrastructure lives in a state file. Only the model, the single most valuable artifact your team produces, gets tossed into a bucket with a hopeful filename. The registry closes that gap.
A model registry is the fix. It is the catalog of every model version your team has ever produced, with its metrics, its parameters, its data lineage, and one field that changes everything: its stage. The registry always knows which version is in Production, and it can put a known-good version back in seconds. By the end of this lesson you will understand exactly how it stores all of that, how a version earns its way to serving real users, and why the whole design turns a 2 a.m. scavenger hunt into a single command.
Think of a the way you would think of a git repository, but for trained models instead of source code. Git does not just store your latest code. It stores every commit, tags releases, and lets you check out any past state. A model registry does the same job for models, and it sits in one specific place in your platform.

The registry is not where the model runs. It sits between training and serving. Training writes new versions into it. Serving reads the Production version out of it. Everything on the platform that needs to know "what is the current model" asks the registry, and the registry gives one definite answer. The crucial detail is in the arrow from serving: it never names a version number. It asks for a stage, models:/churn/Production, and the registry resolves that stage to whatever version is live at that instant. That one layer of indirection is the source of nearly every benefit in this lesson.
Concretely, a registry is a service plus storage that keeps, for every registered model:
| What it stores | Why it matters in production |
|---|---|
| A named model with numbered versions (v1, v2, v3...) | You refer to "the churn model, Production stage," never a filename |
| The metrics of each version (AUC, precision, recall) | You can compare version 5 against version 8 in one query |
It is tempting to think a version is just a weights file with a number stuck on it. It is much more than that, and understanding the full record is what makes the rest of the lesson click. A registered version bundles the artifact pointer together with everything you need to compare it, trust it, reproduce it, and answer for it later. All of it is written once and frozen.

Read the fields in order. The identity is a stable name plus an immutable integer version: v8 never changes once written, and a retrain produces v9 rather than overwriting v8. The stage says where this version sits on the promotion path. The metrics are the holdout scores that let you compare v8 to v5 in one query instead of rerunning old models. The params and the random seed are the exact knobs, and with the seed pinned a rerun rebuilds this version rather than a lookalike. The lineage keys point to the frozen data, the code commit, and the run. The artifact URI says where the heavy bytes actually live. And the model card documents the owner, the intended use, and the known limitations.
That last piece, the model card, is worth calling out because it is easy to skip and dangerous to skip. A model card is a short document attached to the version that says what the model does, what data it was trained on, its known weaknesses, and any fairness or bias checks. It is the difference between a model you can defend to a regulator and one you merely hope is fine. High-stakes promotions, like a credit model, do not clear the approval gate until that card is filled in. The version number is the smallest part of the record. Everything bolted to it is what lets you compare, reproduce, and defend the model months after whoever trained it has forgotten the details.
Here is the single mental model that makes a registry click, and it is worth slowing down for. There are two kinds of thing in play, and they behave in opposite ways. Versions are immutable. Stages are movable.

Every version is a frozen, numbered record. Once v7 is written it never changes again, which is exactly why model versions are integers and not names you overwrite. They are immutable checkpoints. The stages, Production and Staging and Archived, are not versions at all. They are movable pointers, exactly one of each per model, and you slide them onto whichever version you want to serve, test, or keep as a fallback.
Once you see it this way, two operations that sound scary become trivial. Promotion is moving the Production pointer from v8 to v9. Rollback is moving that same pointer from v8 back to v7. In both cases nothing about the versions themselves changes, which is why both are instant and neither can ever lose data. The stage tag is the one mutable thing in the whole system, and because it is a pointer, moving it is cheap and reversible. Serving asks for the pointer by name, so the moment you move it, the next request follows it to the new target with no code change at all.
The single most useful idea in a registry is the stage. Every version lives in exactly one stage at a time, and the registry enforces how versions move between them. In MLflow the classic stages are None, Staging, Production, and Archived.
A version does not go straight from a training run to serving real users. It has to earn it.

Read the flow. A new version starts at None, freshly logged and idle. If it looks promising, it moves to Staging, where it runs against real-shaped traffic and gets scored against the current Production champion. Most versions die in Staging, and that is exactly what you want: the registry is a filter, not a rubber stamp. Only a version that clears the gate gets promoted to Production.
Here is the part that saves you at 2 a.m. When you promote a new version to Production, the previous one is automatically moved to Archived. Archived does not mean deleted. That old model sits there fully intact, which makes it the perfect rollback target. The red arrow in the diagram, from Archived back to Production, is the whole reason people sleep at night.
There is one rule the registry guarantees: for any given model name, there is exactly one version in Production at a time. That is what lets serving ask for models:/churn/Production and always get a single, unambiguous answer. No ambiguity, no voting, no guessing which of three candidate files is the real one.
The old way to ship a model was to copy a pickle file over the current one and redeploy. There was no record of who decided, no check that the new model was actually better, and no way to undo it cleanly. A registry replaces that with a gated, audited promotion built around a champion-challenger duel.

Walk the sequence. The training pipeline registers version 9 at stage None and moves it to Staging. An evaluation job then reads both version 9 and the current Production champion, version 8, and scores them on the same holdout set. Here version 9 wins by 2.1 percent AUC, so the result, along with the model card, goes to a human reviewer, who approves the transition to Production. The instant that transition lands, models:/churn/Production resolves to version 9, and the next serving deploy picks it up. Nobody edited a config file. Nobody copied a pickle. They flipped a stage, and every step is recorded: what the eval said, who approved, and when.
That approval step is where governance actually lives, and in a regulated domain it is stricter than "the number went up." The gate checks more than one thing at once.

A higher AUC is necessary but not sufficient. The candidate must beat the champion on the holdout, its model card must be filled in with intended use and known limitations, and a human with authority must sign off. Only when all three hold does the stage flip, and the whole decision becomes a permanent audit record: the score, the approver, the timestamp. That paper trail is exactly why regulated teams cannot ship without a registry. It turns "someone thought it looked good" into a record you can point a regulator at.
Here is the whole flow in code with MLflow, the most common open-source . Notice how little there is. You log the model during training, register it as a version, then transition its stage.
import mlflow
from mlflow.tracking import MlflowClient
from sklearn.ensemble import GradientBoostingClassifier
mlflow.set_tracking_uri("http://mlflow.internal:5000")
MODEL_NAME = "churn"
# 1. Train, then log the model AND its lineage in one run.
with mlflow.start_run() as run:
model = GradientBoostingClassifier(learning_rate=0.05, max_depth=8)
model.fit(X_train, y_train)
auc = roc_auc_score(y_val, model.predict_proba(X_val)[:, 1])
mlflow.log_params({"learning_rate": 0.05, "max_depth": 8})
mlflow.log_metric("val_auc", auc)
mlflow.set_tag("data_snapshot", "2026-07-01") # lineage: which data
mlflow.set_tag("git_sha", "a1b2c3d") # lineage: which code
# 2. Register the artifact as a NEW numbered version of "churn".
mlflow.sklearn.log_model(
model,
artifact_path="model",
registered_model_name=MODEL_NAME,
)
client = MlflowClient()
new_version = client.get_latest_versions(MODEL_NAME, stages=["None"])[0].version
# 3. Move it to Staging so an eval job can score it against the champion.
client.transition_model_version_stage(
name=MODEL_NAME, version=new_version, stage="Staging",
)
# 4. If it wins the eval gate, promote it. This auto-archives the old one.
client.transition_model_version_stage(
name=MODEL_NAME,
version=new_version,
stage="Production",
archive_existing_versions=True, # last Production model -> Archived
)
Serving never hardcodes a version number. It asks for the Production stage by name, so it always loads whatever is current:
# In the serving container, at load time:
model = mlflow.pyfunc.load_model(f"models:/{MODEL_NAME}/Production")
That one line is the payoff. Training decides what Production means by flipping a stage. Serving just asks for Production. The two are decoupled through the registry, which is why you can promote or roll back without touching serving code.
The mechanics are simpler than the MLflow API makes them look. A registry is really a dictionary of immutable versions plus a small set of movable stage pointers. The playground below is a tiny in-memory registry that captures exactly that core, so you can watch a promotion auto-archive the old champion and a rollback repoint the tag, with no version ever mutated. Run it and read the printed stage tags after each step.
A version number alone is not enough. Six months after you ship version 8, someone will ask why it makes a strange prediction, and you will need to know exactly what produced it. That is lineage: the links from a model back to the data, code, and parameters that created it.

Three inputs feed one training run: a pinned data snapshot (not "the users table" but a frozen version of it), an exact git commit of the training code, and the hyperparameters including the random seed. The run consumes all three and produces version 8, and the registry stamps every one of those links onto the version record. Because those edges are logged, two things become possible that were impossible with loose pickle files.
First, you can trace any live prediction back to its origins. A regulator asks what data trained the model that denied a loan, and you answer with a query instead of an investigation. The prediction was logged with its version number, the version points at its run, and the run points at the exact snapshot, commit, and seed. That is the red arrow in the diagram, and it is the entire reason lineage earns its keep. Second, you can reproduce the exact model. Same data snapshot, same commit, same params, same seed, and you rebuild version 8 byte for byte on any machine. That is what turns "it worked on my laptop" into something a whole team can rely on.
Lineage is also what makes rollback trustworthy. When you promote an archived version back to Production, the registry can tell you exactly what that older model was and how it differs from the one you are backing out of, so you are never rolling back blind.
The single clearest reason to run a registry is what happens when a deploy goes wrong. A new model can pass its offline eval and still misbehave on live traffic, because live traffic is never quite what the holdout set looked like.

Here monitoring catches that version 8 quietly collapsed fraud recall and pages the on-call engineer. Because every version is still in the catalog, the engineer does not go digging for an old file that may not even load anymore. They repoint the Production tag back to version 7, the last known-good version, which is still sitting in Archived fully intact. Serving reloads on the next cycle and recovers. Then they tag version 8 as broken so nobody promotes it again by accident.
Without a registry, this is an all-night incident: find the old pickle, pray it still deserializes with the current library versions, redeploy by hand, hope you got the right one. With a registry it is a single stage change, fully audited, done in the time it takes to type the command. The registry does not prevent bad models from ever reaching Production. It makes the cost of a bad model small.
The reason the rollback target is never a mystery is that the version history is an append-only line you can walk.

Nothing is ever overwritten, so the whole history is right there to read, ordered and scored. When v8 misbehaves you do not guess: you walk back one step to v7, the last version that carried the Production tag, and repoint. Loose files give you a folder full of hopeful names with no order and no metrics. The version line gives you an ordered, scored history where the safe move is always one step to the left.
A registry is not free. It is another service to run, another database to back up, and a discipline the whole team has to follow. It is worth being honest about when it pays off.

| Situation | Ad-hoc files | |
|---|---|---|
| One model, one person, ships once a quarter | Fine, honestly | Nice, but overkill on day one |
| Multiple models, multiple people, weekly retrains | Chaos within a month | Essential |
| Regulated domain (credit, health, insurance) | Non-starter, you cannot prove lineage | Required for audit and approval trails |
| You need to roll back fast during an incident | Hours of hunting | Seconds |
| You need to compare version 5 to version 12 |
MLflow itself was created at Databricks, which built it precisely because their own data teams kept losing track of models. The was added specifically to answer the "which version is live and how do I roll back" problem for teams managing many models at once.
Comcast is a well-known example of the registry at scale. Their data teams have publicly described using the MLflow Model Registry to manage the lifecycle of thousands of models across the company, using the Staging and Production stages plus approval steps to control what reaches customers. Before the registry, tracking that many models by hand was simply not possible, and promotion was informal. After it, every model has a numbered version, a stage, and an audit trail.
The pattern repeats everywhere serious ML runs. DoorDash and Netflix both built internal model stores that do the same core job: catalog every version, record its metrics and lineage, mark exactly one as the live one, and make rollback a stage flip rather than a redeploy. The tools differ, some homegrown, some MLflow, some the registry built into SageMaker or Vertex AI, but the shape is identical. A managed cloud registry hides the two-store split behind an API, yet underneath it is still a metadata database pointing at objects in blob storage, still one stage tag per model, still an append-only version line.
The lesson from all of them is the same one this lesson opened with. The hardest question in production machine learning is not "is the model accurate?" It is "which model is running, what made it, and how do I undo it?" A registry is the piece of the platform whose entire job is to answer that question, instantly, every time. It is the difference between a team that ships models and a team that operates them.
4 questions - Score 80% to pass
What is the core problem a model registry exists to solve?
In MLflow, what happens to the previous Production model when you promote a new version to Production with archive_existing_versions set to true?
Why are model versions immutable while stage tags are movable?
Why does model lineage (data snapshot, git commit, hyperparameters) matter in a registry?
| The hyperparameters that produced it |
| You know the exact knobs: learning rate, depth, seed |
| The lineage: data snapshot, code commit, run id | You can trace any live prediction back to what created it |
| A stage per version: Staging, Production, Archived | You always know what is live, and rollback is a stage change |
| A pointer to the artifact (the actual weights file) | The heavy bytes live in blob storage, the registry points at them |
That last row hints at a design decision worth making explicit, because it is where the split between "facts" and "bytes" lives.

The small structured records (versions, stages, metrics, who approved what) go into a real transactional database like Postgres, because they are the source of truth and you query them constantly. The heavy files (the serialized weights, the environment, the model card) go into blob storage like S3, because you do not put a 400 megabyte weights file in a database row. The registry row is the join between the two: a record for churn v8 carries the URI of its blob, so a single query resolves a stage to both the facts about the model and the location of its bytes. Serving fetches the heavy object only at the moment it actually loads the model.
| Impossible, the metrics are gone |
| One query |
The honest rule: the moment more than one person ships more than one model on any regular cadence, an ad-hoc pile of pickle files becomes a slow-motion outage. The registry does not just organize files. It replaces every "which one is live?" scavenger hunt with a query that has a definite answer. The one real cost is discipline. If people bypass the registry and copy files by hand "just this once," the guarantees evaporate. The registry is only the source of truth if everyone treats it as the only door to Production.