Core

Why ML Models Fail in Production: The Production Gap

0 of 13 complete

0%

Contents

Back|CoreWhy ML Models Fail in Production: The Production Gap
1/13
51 min left
  1. Home
  2. AI Engineering: Foundation
  3. Core Concepts
  4. Why ML Models Fail in Production: The Production Gap
Related Topics
Fine-Tuning vs RAG vs Prompting: Choosing Your ApproachLLM and GenAI OpsParameter-Efficient Fine-Tuning: LoRA and QLoRALLM and GenAI OpsEvaluating LLMs in Production: Grading Answers That Have No Right AnswerLLM and GenAI OpsPrompt Management and Versioning: Treat Prompts as Production CodeLLM and GenAI OpsVector Databases and Approximate Nearest Neighbor SearchLLM and GenAI Ops
Next lessonModel Packaging and Containerization: Killing 'Works On My Machine' for ML

System Design

1 of 13
  • Foundation
  • Intermediate
  • Advanced
  • Capstone
  • AI Engineering

    • Foundation
    • Data, RAG and Agents
    • Evaluation, LLM Ops and Security

    systemdesign.academy

    • Home
    • Glossary
    • Interview prep
    • Reviews
    • About
    • Privacy
    • Terms

    Why Do ML Models Fail in Production?

    Most ML models fail in production for a reason that has nothing to do with the model. They fail because everything around the model, the part that actually keeps it running once real users show up, was never built. That is the production gap, and closing it is what this track exists to do.

    Picture it in the concrete. A data scientist spends three weeks building a churn prediction model. It hits 94 percent accuracy on the test set. The demo goes perfectly. Everyone in the room is impressed. Six months later, that model has never made a single prediction for a real customer. It is sitting in a notebook on a laptop, exactly where it was born.

    This is the most common story in machine learning, and it is not rare. Industry surveys have repeatedly put the share of models that never reach production somewhere between half and the vast majority. Gartner has estimated that only about half of models make it from prototype to production even at mature organizations. The exact number is argued about, but the direction is not. Most models never ship.

    The natural reaction is to blame the model, or the data, or the algorithm. That reaction is wrong, and correcting it is the entire point of this track.

    The hard part of machine learning in production was never the model. It is everything around the model.

    Building the model is the part that gets taught in courses and the part that feels like the real work. But a trained model is only the engine. It cannot move anyone anywhere on its own. The discipline of building everything around that engine so it runs reliably, every day, at scale, is called ML platform engineering. That is what this track is about, and by the end of it you will be able to take a model that only exists in a notebook and turn it into a service the rest of a company can depend on.

    The Engine and the Car

    Picture a beautifully built engine sitting on a workbench in a garage. It is precision engineered. It revs perfectly. It is also completely useless for getting you to work, because an engine on a bench does not move anyone anywhere.

    To turn that engine into a car you can actually drive, you need a lot more:

    • A chassis to mount the engine on
    • A fuel system that feeds it continuously
    • A transmission that connects it to the wheels
    • A dashboard so the driver can see what is happening
    • Brakes and warning lights for when something goes wrong
    • A pit crew that services it so it keeps running

    A trained model is the engine. The car is everything else. The data scientist builds the engine and cares about how well it runs on the bench, which we measure as accuracy. The ML platform engineer builds the car and cares about whether it keeps running on a real road, next week, under load, without a mechanic riding along.

    Hold onto this picture. Every topic in this track is another part of the car. When we reach , that is the transmission connecting the engine to the wheels. When we reach monitoring, that is the dashboard and the warning lights. When we reach retraining pipelines, that is the pit crew. The analogy is not decoration, it is the map.

    Why Doesn't a Notebook Model Work in Production?

    Here is the gap made concrete. Everything that works quietly in a notebook becomes a real engineering problem the moment real users show up. The left side of this map is what a data scientist does on a laptop. The right side is the machinery each of those habits forces you to build. The empty space between them is the gap.

    The production gap drawn as a three-column map: on the left, six notebook habits such as running on your own Python and library versions, calling model dot predict in a cell, training once by hand on a CSV, computing features inline, saving one pickle file, and checking accuracy once on a test set; in the middle, an empty gap with arrows; on the right, the six production systems each habit forces you to build, namely containerization, a serving API, training pipelines, a feature store, a model registry, and monitoring with drift detection

    Read the same idea as a table, because the wording of each jump is where the engineering lives.

    In the notebookIn production you now need
    Runs on your machine, your Python, your library versionsIt must run identically on a server you will never see, which means containerization
    You call model.predict(x) in a cellTen thousand users call it per second over , which means a serving API
    You trained it once, by hand, on a CSVThe data changes every week, so it must retrain itself, which means pipelines
    Features computed inline in the notebook

    Why Research and Production Feel Like Two Different Jobs

    The deeper reason models stall is that building a model and running a model optimize for different things, and the second job barely resembles the first. A research or notebook workflow chases the highest accuracy it can get on a fixed dataset. A production workflow chases reliable answers on data that never stops moving.

    A research versus production comparison matrix with eight rows: what you optimize, the data, how often it runs, the latency budget, who operates it, what failure looks like, how versioning works, and what done means, contrasting an amber research and notebook column against a blue production and platform column, showing research optimizes peak accuracy on a static CSV run once by hand while production optimizes reliable answers on a drifting live stream run continuously by on-call engineers where failure is silent

    Walk down the two columns and the mismatch jumps out. In research, the data is a static file that never moves, the model runs once when you press a button, minutes of are fine, the author operates it on a laptop, and failure means a cell throws a visible error. In production, the data is a live stream that drifts every week, the model runs continuously and unattended, you have tens of milliseconds per request, an on-call engineer operates it at three in the morning, and failure is silent: the model keeps returning confident numbers that are quietly wrong.

    The notebook column is not wrong. Everything in it is true and necessary, and no model exists without that work. But each cell quietly assumes the world holds still. Production is the discipline of making the same model keep working after the world moves, which turns out to be a different job with different tools. That is why a brilliant model and a reliable service can be years apart, and why a strong data scientist and a strong platform engineer are complementary skills rather than the same skill.

    Why Production ML Is a Loop, Not a Line

    A normal web application is a straight line. You deploy it, and it keeps doing the same correct thing until you change the code. A machine learning model is different. It starts decaying the moment the real world drifts away from the data it learned from. Fraud patterns change. Customer taste changes. Prices change. The model does not know any of this happened, and it will keep answering with total confidence as it slowly goes wrong.

    So production machine learning is not a one-time deploy. It is a loop that has to run forever.

    The machine learning lifecycle drawn as a serpentine loop of seven stages that visibly closes back on itself: collect data, train and validate, package, deploy and serve, then turning down and back through monitor live, detect drift, and retrain on fresh data before returning to serving, with a closing note that a mature platform runs this whole loop on its own while a human steps in only for the decisions

    The platform's job is to run this loop with as little human effort as possible. Serve predictions. Watch how the model behaves on live data. Catch the moment it starts to decay. Retrain on fresh data. Redeploy the new version safely. A human steps in for decisions, not for plumbing. When people say a company has a mature ML platform, they mean this loop turns on its own.

    To see why the loop is not optional, watch what happens to a model that ships and is then left alone. The accuracy falls while nothing looks broken, which is exactly what makes it dangerous.

    A model decay timeline plotted as a line chart over six months: a red precision line starts at 94 percent when the model ships and slides steadily down to 78 percent, while a green dashed line for latency and uptime stays flat and healthy across the top the entire time, with an amber marker at month four where chargebacks finally spike and force someone to look, illustrating that the operational metrics everyone watched stayed green while the one metric that mattered rotted unnoticed

    This is the failure mode that catches teams by surprise, because every instinct from normal software says a healthy service is a correct service. A web bug crashes and pages someone within seconds. A decaying model raises no exception, keeps its flat, keeps the service green, and simply hands back worse answers on every request. In the example above, precision on real fraud fell sixteen points before anyone noticed, and the only thing that eventually raised the alarm was a business metric, chargebacks, not any engineering dashboard. The lag between the model going wrong and someone finding out is the whole danger, and closing that lag is what monitoring exists to do.

    How Does a Shipped Model Actually Break?

    When a production model goes wrong, it almost never throws an exception. It keeps returning confident numbers that are quietly less correct than they used to be. There are four families of failure, and not one of them is a bug in the model. Each is a property of the world the model now lives in.

    A taxonomy of the four ways a shipped model breaks, drawn as four cards: data and concept drift where live inputs slide away from the training distribution, train-serve skew where a feature is computed differently in training and serving, scale and latency where production load and cost overwhelm a notebook-grade setup, and silent data-quality breaks where an upstream job renames a column or changes a unit, each card noting the concrete symptom and the platform capability that catches it, with the shared theme that none of them raises an exception

    The first family is drift. The live inputs slide away from the training distribution, or the meaning of the label itself changes, and the model is left answering last year's world. The second is , and it is worth dwelling on because it is the most confusing one to debug. A feature gets computed one way when you train and a subtly different way when you serve, so the model is fed numbers at request time that it never actually saw in training.

    Train-serve skew drawn as two stacked data-flow lanes computing the same average-spend feature two different ways: the training lane runs offline in a notebook over two years of clean sorted history with a pandas rolling mean and produces 42.10, while the serving lane runs online in a separate re-implemented Java service over only the last thirty days with a different window edge and produces 37.80, so the model is trained to trust one number and served a different number on every request, a divergence that offline metrics never catch

    Look closely at what went wrong there. The offline training path and the online serving path are two different bodies of code, written by different people at different times, and they compute the same feature to two different values: 42.10 versus 37.80. The model learned on the first and is fed the second in production, and your offline evaluation will never reveal the gap, because offline you only ever run the training path. This is exactly the failure a exists to prevent, by making both paths read one shared definition instead of two copies that drift apart.

    You can watch skew happen in miniature. The tiny program below computes the same feature the training way and the serving way, shows them diverge, and then shows a single shared definition collapsing the gap to zero. No machine learning library is involved, because the bug has nothing to do with the model.

    How Is MLOps Different From DevOps?

    People often assume MLOps is DevOps with a model dropped in. The difference is exactly one thing, and that one thing is what makes it hard.

    In traditional DevOps, two things change over time: your code and your configuration. Both are things you control directly. When something breaks, it broke because a human changed code or config, and you can find that change in the history and reason about it.

    In MLOps, a third thing changes on its own: the data.

    Code you control. Data drifts whether you touch it or not. And when the data changes, the model's behavior changes without a single line of code changing.

    This is why an ML platform needs machinery that a normal backend never needs. You have to version data, not just code, because the same code trained on different data is a different model. You have to version models and track exactly which data trained each one, so that when a model misbehaves you can reconstruct precisely how it was made. You have to monitor the statistical shape of live traffic, not just error rates and , because the failure hides in the distribution rather than in the logs. And you have to be able to retrain and redeploy automatically, because manual retraining does not keep pace with data that moves every week. It is DevOps standing on a floor that is quietly moving under your feet, and every extra piece of an ML platform exists to deal with that moving floor.

    How Far Along Is Your ML Platform?

    Nobody builds the whole platform on day one, and you should not try to. There is a well-known ladder for this, described in Google's widely cited MLOps guidance, and most teams can honestly place themselves on exactly one rung. The point of the ladder is direction, not a grade. Each rung automates one more thing a human was doing by hand.

    The MLOps maturity ladder drawn as three rising stairs: Level 0 Manual where every step is script-driven and run from a notebook with rare deploys and no monitoring, Level 1 Pipeline automation where training is an automated repeatable pipeline that retrains on fresh data with a feature store keeping train and serve consistent, and Level 2 CI/CD automation where the platform builds tests and deploys pipelines automatically and monitoring triggers retraining so the full loop turns with a human only in the decisions, with each step taller than the last to show manual effort falling as you climb

    At Level 0, everything is manual. A person runs a notebook, exports a model, and hands it to someone to deploy, and this happens rarely, weeks or months apart, with no monitoring and no automated retraining. At Level 1, the training pipeline runs itself: fresh data triggers a repeatable, automated retrain, and a keeps the training and serving paths consistent. At Level 2, CI/CD builds, tests, and deploys the pipelines automatically, monitoring detects drift and triggers retraining, and new models roll out safely behind automated checks, so the full loop turns with a human only in the decisions.

    You do not have to reach Level 2 to win. The single most valuable jump for most teams is from 0 to 1, from hand-run notebooks to an automated training pipeline, because that is the move that turns a one-off model into a repeatable capability. Every lesson in this track is a rung on this ladder, and knowing which rung you are standing on tells you what to build next.

    Why Is the Model Such a Small Part of the Work?

    If the surrounding platform feels like a lot of machinery for one model, that reaction has a famous name. In 2015, a team at Google led by D. Sculley published a paper called "Hidden Technical Debt in Machine Learning Systems," and it contains the single most quoted picture in the field.

    The hidden technical debt field from Sculley and colleagues 2015, drawn as a large bordered ML system containing many surrounding concern boxes such as data collection, feature extraction, data verification, configuration, process management, serving infrastructure, monitoring, analysis tools, machine resource management, data structures, and metadata, with the actual ML code shown as one small black box in the middle, illustrating that the model is a tiny fraction of a real machine learning system while the platform around it dwarfs it in both code and effort

    The finding is blunt. In a real machine learning system, the model code is the tiny black box in the center. Everything around it, data collection, feature extraction, data verification, configuration, process management, serving infrastructure, monitoring, analysis tooling, resource management, is the actual system, and collectively it dwarfs the model in both lines of code and engineering effort. The part almost nobody plans for is the part that turns out to be the system.

    This is why "the model is done" and "the system is done" are years apart, and why it is so easy to underestimate a machine learning project. The course that taught you to train a model built the black box. This track builds everything around it, because that surrounding platform is where models actually go to fail, and it is also where the durable engineering value lives. The model is often replaced or retrained many times over; the platform outlives every model that runs on it.

    What Is Inside an ML Platform?

    Put the pieces together and you get the map for this entire track. Here is the whole platform on one canvas. Follow the numbered path to see how a model travels from a notebook experiment to a monitored production service, with data and compute as shared rails that feed every stage.

    The full ML platform anatomy drawn as an architecture map with two shared rails on the left, data and pipelines plus a feature store plus compute and GPU plus access and secrets, feeding a numbered flow that runs from experiment in a notebook, to package and register, to CI/CD checks, to serving an API under load, to live traffic from real users, to monitoring and drift detection that triggers a retrain and closes the loop back to the start

    Each component solves one specific failure of the notebook world, and each has a real industry-standard tool behind it.

    ComponentThe problem it solvesThe tool you will learn
    Packaging and containerizationThe model must run the same everywhere, BentoML
    Serving and inference APIsTurn a model file into a service under loadBentoML
    Training pipelines and orchestrationRetrain automatically instead of by handKubeflow
    Feature storesKill train and serve skew, reuse features

    Should Every Model Cross the Gap?

    Closing the production gap is expensive, so the honest first question is whether a given model should cross it at all. Not every model deserves a platform. A one-off analysis that answers a question once does not need serving, retraining, or drift monitoring, and building all of that around it would be waste dressed up as rigor.

    A decision map for what to productionize, branching from a first question of whether the model makes a repeated automated decision: on the no side, a one-off analysis that answers a question once should be left as a notebook or at most wrapped in a simple pipeline if it will be rerun periodically; on the yes side, a model that drives a live product should get serving and monitoring now, and the full loop with a feature store and retraining if the data it learns from keeps changing, with a closing rule of thumb to match the machinery to the decision rather than to the hype

    Start with one question: does this model make a repeated, automated decision, the same kind of prediction many times a day feeding a real product or process? If the answer is no, if it is a board deck or a sizing study or a one-time segmentation that you run, read, and act on once, then a notebook and a written result is the right amount of engineering. If someone will rerun it monthly on fresh data, wrap it in a simple pipeline and stop there, without serving or drift monitoring.

    If the answer is yes, if it scores fraud or ranks recommendations or sets prices thousands of times a day without a human reading each prediction, then it earns real infrastructure. Ask one more question: does the data it learns from keep changing? If it does, you need the full loop, serving and a and monitoring and retraining. If the data is stable, you can start with serving and monitoring and add retraining only when drift actually shows up.

    The rule of thumb is to match the machinery to the decision, not to the hype. Productionizing pays for itself exactly when a repeated, automated decision meets data that drifts, because that is the only regime where the platform's cost is clearly justified. Build the loop where it earns its keep, and leave a good one-off analysis as the notebook it deserves to be.

    How Did Uber, Netflix, and Meta Build Their ML Platforms?

    None of this is theoretical. Every company running machine learning at scale eventually built exactly this kind of platform, usually after feeling the pain of deploying models by hand.

    How the giants built their ML platforms, drawn as six panels: Uber built Michelangelo to standardize the whole lifecycle, Netflix built Metaflow so scientists write normal Python while the platform handles scale, Meta built FBLearner Flow to run experiments at industrial scale, Airbnb built Bighead for reproducible model development, Spotify composed its platform on top of Kubeflow from open-source parts, and a sixth panel drawing the common lesson that the model was never the bottleneck and a platform is what removes it

    Uber built Michelangelo, an internal ML platform that standardized the whole lifecycle from to training to serving. It let hundreds of teams ship models without each one reinventing the plumbing, and it powered everything from arrival-time estimates to fraud detection.

    Netflix built Metaflow, a framework for building and running data science workflows, which they later open sourced. It let scientists write normal Python while the platform handled scaling, scheduling, and moving work to the cloud.

    Meta built FBLearner Flow, which ran millions of modeling experiments and served predictions across the company. At its peak it was training and serving a huge share of Meta's production models through one platform.

    Airbnb built Bighead, and Spotify built its own ML platform on top of Kubeflow, both for the same reason. The model was never the bottleneck. The path from a good model to a reliable production service was the bottleneck, and a platform is what removes it.

    The lesson from all of them is the same. If you want machine learning to be a repeatable capability rather than a series of one-off heroics, you build the platform. That is what you are going to learn to do here, one part of the car at a time.

    Knowledge Check

    Knowledge Check

    4 questions - Score 80% to pass

    Q1

    According to this lesson, what is the main reason most ML models never reach production?

    Q2

    What single factor makes MLOps fundamentally harder than traditional DevOps?

    Q3

    A model scores well on offline evaluation but performs worse in production, and no one can reproduce the gap. What is the most likely cause?

    Q4

    Why is production machine learning described as a loop rather than a line?

    The same feature logic must run at training time and at request time, or predictions go wrong, which means a
    One file named model_final_v2_really.pklDozens of versions, and you must know which one is live and roll back in seconds, which means a
    Accuracy checked once, on the test setThe world shifts and accuracy silently rots, which means monitoring and drift detection

    Look at the right-hand column. Every row names a different piece of machinery. That is not a coincidence. An ML platform is precisely the system that closes every one of these gaps. Each becomes its own lesson in this track. The first two, getting a model to run identically everywhere and answer over HTTP, are covered in model packaging and containerization and model serving and inference.

    The reason this table matters is that none of the right-hand items are optional once the model is load-bearing. You cannot serve ten thousand requests per second from a notebook cell, you cannot debug a production incident from a filename with v2 in it, and you cannot notice a model going wrong if the only accuracy number you ever computed was on a static test set months ago. Each row is a hole that becomes a real outage if you leave it open.

    Catching that decline is only half the job. The platform also has to decide when a retrain is actually worth it, because retraining on thin or bad data can ship a worse model than the one you have. A mature loop does not retrain blindly on a timer. It retrains when the evidence says the current model has drifted far enough that a fresh one, trained on enough good new data, will genuinely do better.

    The third family is scale and . The notebook served one request in two seconds and that was fine. Production needs thousands of requests per second inside a tight latency budget, on hardware that costs real money every hour, and a model that is accurate but too slow or too expensive to serve is not a shippable model. The fourth family is silent data-quality breaks: an upstream job renames a column, starts sending nulls, or flips a unit from cents to dollars, and the model does not error. It feeds garbage in and hands confident garbage back out, on every request, until someone downstream finally notices the numbers are wrong.

    Notice the shared thread through all four: no exception is ever raised. That single fact is why monitoring is not a nice-to-have you add later. It is the only alarm you have, because the model itself will never tell you it has gone wrong.

    Feast
    and versioningKnow what is live, roll back fastMLflow
    Monitoring and drift detectionCatch silent model decayEvidently
    Scaling and GPU infrastructureServe heavy models without burning money
    Production and LLM systemsRun retrieval plus a model reliablyLLM serving stack

    Build these once as reusable infrastructure and the payoff is enormous. The next model your team ships does not rebuild any of this. It plugs into a platform that already exists, and it goes from notebook to production in days instead of quarters. That multiplier is the real reason companies invest in platforms rather than deploying each model by hand.

    The last row, production RAG and LLM systems, is where this same platform meets generative AI, and the choices there start with fine-tuning versus RAG versus prompting.

    It also helps to see the same platform as a stack rather than a flow. Users only ever touch the top layer. Every layer beneath it exists so that top layer keeps answering correctly next month.

    The same ML platform seen as a layered stack from top to bottom: product and users at the top, then serving and inference backed by BentoML, then registry and CI/CD backed by MLflow, then a feature store backed by Feast, then data and training pipelines backed by Kubeflow, and compute and GPU infrastructure backed by Kubernetes at the base, with a note that a wrong prediction the user sees might really be a feature-store mismatch or a broken pipeline several layers down, so debugging a production model means understanding the whole stack

    Reading the stack top to bottom makes the dependencies obvious. A wrong prediction that a user sees at the top might really be a feature-store mismatch two layers down, or a broken data pipeline three layers down. That is why debugging a model in production means understanding the entire stack, not just the model file, and it is why platform engineers spend far more time in the lower layers than in the model itself.