Core

The Full ML Platform: Assembling Every Piece Into One Golden Path

0 of 12 complete

0%

Contents

Back|CoreThe Full ML Platform: Assembling Every Piece Into One Golden Path
1/12
31 min left
  1. Home
  2. AI Engineering: Foundation
  3. Core Concepts
  4. The Full ML Platform: Assembling Every Piece Into One Golden Path
Prerequisites
Production RAG and LLM Systems: Running It, Not Just Building Itrequired
Related Topics
Fine-Tuning vs RAG vs Prompting: Choosing Your ApproachLLM and GenAI OpsParameter-Efficient Fine-Tuning: LoRA and QLoRALLM and GenAI OpsEvaluating LLMs in Production: Grading Answers That Have No Right AnswerLLM and GenAI OpsPrompt Management and Versioning: Treat Prompts as Production CodeLLM and GenAI OpsVector Databases and Approximate Nearest Neighbor SearchLLM and GenAI Ops
Previous lessonProduction RAG and LLM Systems: Running It, Not Just Building It

System Design

  • Foundation
1 of 12
Intermediate
  • Advanced
  • Capstone
  • AI Engineering

    • Foundation
    • Data, RAG and Agents
    • Evaluation, LLM Ops and Security

    systemdesign.academy

    • Home
    • Glossary
    • Interview prep
    • Reviews
    • About
    • Privacy
    • Terms

    Nine Tools, One Question: How Do They Fit Together?

    You have now met every part of the car. Packaging in a container so the model runs the same everywhere. A serving API that answers requests under load. A pipeline that retrains without a human. A that kills train and serve skew. A registry that knows which version is live. Monitoring that catches silent decay. Scaling and GPUs so heavy models do not bankrupt you.

    Each of those was a lesson. Each was a real tool: Docker, BentoML, Airflow, Feast, MLflow, Evidently, . But here is the thing nobody tells you when you learn them one at a time. A pile of good tools is not a platform. A drawer full of engine parts is not a car.

    The question this lesson answers is the one that actually matters at work: how do these nine pieces connect into a single system that a team can plug into and ship a model in days instead of quarters?

    A platform is not the tools. It is the wiring between them, plus the promise that a team never has to redo the wiring.

    That promise has a name in this field. It is called the golden path: the one well-worn route a model takes from a notebook to a monitored production service, paved so thoroughly that taking any other route feels like a mistake. Building that path is what this whole track has been preparing you to do.

    There is a second, quieter promise underneath the first. When the wiring is done once and owned by someone, the second model a team ships does not pay for any of it again. The first model funds the platform; every model after it rides for free. That multiplier is the entire economic case for building a platform instead of hand-deploying models forever, and it is why the companies you will read about at the end of this lesson all eventually built the same thing.

    From a Pile of Tools to a Reference Architecture

    Let me put the whole thing on one canvas. This is the reference architecture: every tool from the track, wired into the loop a model actually travels, with the real product that fills each box named underneath it.

    The complete ML platform reference architecture drawn left to right as six connected stages, data and ingest with Kafka and Snowflake, then a feature store with Feast and Redis, then a training pipeline with Airflow and Ray, then a registry and packaging stage with MLflow and Docker, then serving with BentoML and KServe, then monitoring with Evidently and Grafana, all sitting on a shared Kubernetes and GPU compute foundation, with a highlighted red retrain loop running from monitoring back to the training pipeline

    Trace the path once and every earlier lesson clicks into place. A model is authored in a notebook and committed. A pipeline trains it on features pulled from the , so the values used to train are the same values used to serve. The trained model lands in the registry, which decides what is blessed for production. The blessed version is packaged into a container, deployed to a serving API, and run on Kubernetes, where it fetches fresh features from the online store for every request. Then monitoring watches the live predictions and, the moment the world drifts, fires a retrain trigger back at the pipeline.

    That last red arrow is the whole ballgame. It is what separates an ML platform from a one-time deployment. Without it, you launched a model. With it, you built a machine that keeps a model correct. Notice also the shared compute foundation at the base: training and serving are scheduled onto one cluster, which is what lets a platform amortize expensive GPUs across every model instead of buying idle hardware per team.

    Now watch the same architecture as a choreography that unfolds over time, from the team's first commit to the retrain that closes the loop. The map shows the boxes; this shows who acts, in what order.

    An end-to-end sequence diagram with seven actor lanes, the ML engineer, Git and CI, the training pipeline, the registry, serving, monitoring, and the closing loop, showing ten numbered steps from committing model code and feature definitions, through validation, pulling point-in-time features, training and logging to the registry, eval-gated promotion to staging then production, packaging and canary deployment, per-request feature fetch and prediction, live drift scoring, and a final red step where a drift breach fires a retrain trigger back to step three

    Notice how little the ML team does on the happy path. They commit code. Everything after that, training, logging, promotion, packaging, deployment, monitoring, and the retrain trigger, runs without them touching a server. That is not an accident. That is the design goal, and the distance between step 2 and step 10 running unattended is a fair measure of how mature a platform really is.

    One Platform, Three Loops Turning at Different Speeds

    The single canvas can fool you into thinking a platform is one loop. It is three, and they run on wildly different clocks. Seeing them separately is the fastest way to reason about where a platform actually breaks.

    Three interlocking cycles shown side by side, an offline loop in violet that runs over hours to days and pulls features, trains, logs to the registry, and gates on evaluation, an online loop in blue that runs in milliseconds and receives a request, fetches online features, runs inference, and returns a prediction, and a feedback loop in green that runs continuously and collects logged predictions, tracks drift and quality, compares to a baseline, and fires a retrain, with three hinge cards below naming the feature store, the model registry, and the retrain trigger as the shared points where the loops connect

    The offline loop retrains over hours or days. Its job is to produce a better blessed model on fresh data. The online loop answers a single request in milliseconds; it fetches this entity's features and runs inference thousands of times a second. The feedback loop watches production continuously and, when quality slips past a threshold, kicks the offline loop into another lap.

    The loops interlock at exactly three hinges, and this is the part worth memorizing. The is the hinge between offline and online, which is why one shared definition kills train and serve skew. The registry is the hinge between the offline loop and the online loop, which is why the offline loop hands over a specific version and never a loose file. The retrain trigger is the hinge between the feedback loop and the offline loop, which is what lets the whole thing turn on its own.

    Here is the practical payoff. When a platform has an incident, it is almost never inside one loop. It is at a hinge: a feature computed differently between the offline and online paths, a registry pointer left on last week's model, or a retrain trigger that quietly stopped firing. Get the three hinges right and the loops mostly take care of themselves.

    The Same Platform, Seen as a Stack

    The flow diagram shows how a model moves. But when you are building the platform, it helps to see it as layers, because layers are how you assign ownership, decide what to build first, and choose the concrete tool for each job.

    A canonical open-source reference stack drawn as eight horizontal layers from top to bottom, self-service with config and templates, serving and inference with BentoML KServe and Triton, registry and CI/CD with MLflow and Docker, observability with Evidently Prometheus and Grafana, training and orchestration with Airflow Kubeflow and Ray, feature store with Feast and Redis, data and storage with Snowflake and S3, and compute and GPU with Kubernetes, each layer labelled with what it owns and the tool that owns it, and a caption noting teams only ever touch the top layer

    Read it bottom to top. Compute at the base, then data and storage, then the , then training and orchestration, then observability, then the registry, then serving, and at the very top the self-service layer that teams actually interact with. The tools named on the right are one battle-tested assembly you can copy. You do not have to pick these exact products, but you do have to fill every layer, because a gap in any layer surfaces as a wrong answer at the top.

    Two rules fall out of this picture. First, you build from the bottom up. There is no point in a slick self-service layer if there is no compute under it, so compute, data, and training come first and the polished top layer comes last. Second, and this is the one people miss, the quality of a platform is measured by how thin that top layer feels. If an ML team has to understand at the base to ship a model at the top, your platform is leaking its lower layers upward. A mature platform lets a team declare a model in config and never learn what is beneath it.

    LayerThe lesson behind itWhat it owns
    Self-serviceThis lessonThe golden path, templates, config

    Who Owns What: The Line That Makes or Breaks a Platform

    Here is where most internal platforms quietly fail, and it has nothing to do with technology. It is an org chart problem. If you do not draw a clear line between what the platform team owns and what the ML teams own, one of two disasters follows.

    An ownership map with two columns split by a vertical dashed line labelled golden path, the left column is the platform team which owns the cluster and compute, pipeline templates and orchestration, the feature store and registry, the serving framework and rollout, and the monitoring stack, with a success metric of time-to-production averaged across teams, the right column is the ML team which owns feature definitions, model code and architecture, evaluation thresholds, the business metric to move, and the config that declares the model, with a success metric of the business outcome

    The platform team owns the plumbing: the cluster, the pipeline templates, the , the registry, the serving framework, the monitoring stack. They treat all of it as a product, and their customers are the ML teams inside the company. Their success metric is not model accuracy. It is time-to-production, averaged across every team.

    The ML teams own what is specific to their model: the feature definitions, the model code, the evaluation thresholds, and the business metric the model is supposed to move. They care about fraud caught or deliveries tightened, not about how the autoscaler works. The golden path is the contract between the two, and the crispness of that line is a leading indicator of whether the platform will scale to ten teams or strangle at three.

    When the line blurs, the platform dies in one of two mirror-image ways. Both are org-shaped, not code-shaped, and both are fatal.

    A two-card anti-pattern diagram in rose, the first card is platform as a ticket queue where every deploy is fulfilled by hand so the platform team becomes a bottleneck that scales linearly with the number of models, the second card is every team rebuilds plumbing where teams stand up their own serving stacks so you end with ten half-maintained platforms and the shared investment buys nothing, and a green banner underneath naming self-service as the single cure for both

    The first failure is the platform as a ticket queue: every deploy is a request a human fulfills, so the platform team becomes a bottleneck that grows linearly with the number of models. The second is the opposite, where the path did not fit so teams rebuilt their own plumbing, and now the shared investment bought nothing. The cure for both is the same idea, and it has a name.

    Self-Service: The Team Plugs In, It Does Not Rebuild

    Self-service is the difference between a platform and a service desk. The test is simple. Can a new team ship a standard model without filing a ticket or learning your infrastructure? On the golden path the answer is yes, and the shape of that yes is a config file.

    A self-service diagram showing a dark model.yaml config on the left that a team writes, declaring a fraud-scorer model, its owner, three features, a daily training schedule with an AUC eval gate, a realtime CPU serving template with canary rollout, and monitoring with drift on and retrain-on-breach true, an arrow in the middle labelled the platform does eight jobs, and six green output cards on the right, a trained model, a registry entry, a container, a live endpoint, dashboards, and a wired retrain loop

    Follow the yes path. A team has a trained model. A serving template already fits it. The features it needs are already in the . So the team writes a config file that declares the model and its features, and the platform does the rest: it trains, logs to the registry, packages, deploys, and lands the model live, already monitored, already inside the retrain loop. The team never touched Kubernetes. That is a good day, and on a mature platform it is the common day.

    Now the no branches, because they are where the platform grows. If a serving template does not fit, say the model needs GPU inference the platform has never served, the platform team builds that capability once, as a reusable template. The next team that needs GPU inference gets it for free. Same with features: a brand-new feature is registered once and then belongs to everyone. The gap closes permanently instead of being patched per team. This is the discipline that keeps a platform from rotting into a pile of one-off exceptions.

    The maturity ladder underneath all of this is not a feeling, it is measurable. The small program below is a readiness scorer: mark each capability as missing, manual, or automated, and it reports your true rung and the one thing to build next. The rule it encodes is the one teams most often get wrong.

    Run it and change the scores to match your own team. The lesson it enforces is that your rung is your lowest honest capability, not your highest. A team with beautiful automated training but a hand-pushed retrain is not a Level 4 team; it is a Level 3 team with a great training story, and the next thing to build is the retrain trigger, not a self-service portal.

    The Readiness Scorer, Worked by Hand

    Let me walk through the scorer above by hand, with the scores it starts with, so you can see exactly why it prints what it prints.

    Part 1, the scores. The capabilities dictionary gives each capability a number. 0 means missing, 1 means done by hand, 2 means automated. The starting scores describe a team in the middle of the journey. Training, the and the registry are at 2. Deploy, canary rollback and monitoring are at 1. The closed retrain loop and self-service config are at 0.

    Part 2, the order. The order list puts the eight capabilities in the order they depend on each other. You cannot trust an automated deploy if you cannot rebuild the model it deploys, so reproducible training comes first. Self-service comes last, because it sits on top of everything else.

    Part 3, the loop. The code walks down order from the top and stops at the first capability that is not yet a 2. Position 0 is reproducible training, a 2, so it moves on. Positions 1 and 2, the feature store and the registry, are also 2. Position 3 is automated deploy, which is a 1. The loop stops there. So rung becomes 3 and next_build becomes automated deploy.

    Part 4, the output. It prints the table, then "your true rung: 3 of 8" and "build THIS next: automated_deploy". The last line names the capability right after it, canary rollback, as the one not to skip ahead to.

    Notice what the scorer ignores. Live monitoring is at 1, which may feel like progress, but it does not raise the rung. The rung is set by the first gap in the chain, not by the best score.

    Now try one change. Set automated_deploy to 2 and run it again. The loop now stops at canary rollback, position 4, so the rung becomes 4. One capability moved from manual to automated, and the team moved up exactly one rung. That is the ladder from the maturity slide, in a few lines of code.

    Build vs Buy: Do You Even Assemble This Yourself?

    Everything so far assumed you are assembling the platform out of open-source parts. You do not have to. A cloud vendor will sell you most of these boxes behind one API. So before you build, decide honestly whether you should, and the honest answer is decided per layer, not for the platform as a whole.

    A build-versus-buy decision matrix with six component rows, compute and GPU, feature store, orchestration, registry and CI/CD, serving, and monitoring, and three columns showing the open-source build option, the managed buy option with real products like EKS, GKE, Tecton, SageMaker Feature Store, Vertex Pipelines, Databricks, SageMaker, and Arize or Fiddler, and a when-to-pick-which column that recommends hybrid for compute and serving, buy-first for feature store and monitoring, and build for orchestration and registry

    There is no universal answer, only a right answer for your scale, and it usually differs by layer. Buy the layers where a managed product is good enough and cheap enough: orchestration and monitoring rarely justify a from-scratch build early on. Self-host the one or two layers where cost or control actually bites, which in practice is compute and the serving hot path. Most real platforms end up hybrid, and that is not a compromise, it is the correct design.

    The deeper question is when the crossover from buy to build happens at all, and to answer that you have to look at where the money goes.

    A cost breakdown for a representative mid-size platform running dozens of models, a horizontal one-hundred-percent stacked bar split into training GPU compute at forty-six percent, inference serving at eighteen percent, storage at twelve percent, the online feature store at eight percent, monitoring at six percent, and data pipelines plus platform team tooling at ten percent, with a labelled legend and a side note explaining that because compute is roughly two thirds of the bill, the buy-to-build crossover is driven by GPU cost, which is why Uber and Netflix build

    Treat those numbers as illustrative shape, not a benchmark, because the exact split swings with your workload. But the shape is stable and it is the whole point: compute, and GPU compute in particular, is roughly two thirds of the bill. That single fact decides build versus buy at scale. A managed platform's per-training-hour and per-endpoint fees are a fine trade when you have three models. At three hundred, those same fees dwarf a platform team's salaries, and owning GPU scheduling and latency becomes worth the headcount. That is why the giants build, and it is almost never the or the dashboards that pushed them over the line.

    A Maturity Model: You Do Not Start at the Top

    If the reference architecture looks like a lot, that is because it is. And the single most useful thing to understand about it is that nobody builds it all at once. Platforms grow up a ladder, one rung at a time, and each rung earns its place by removing the specific pain of the rung below.

    A phased rollout roadmap drawn as five timeline segments, phase zero in weeks zero to two adds reproducible training in version control with MLflow to kill the which-pickle-is-live problem, phase one in weeks two to six adds a scheduled pipeline and feature store with Airflow and Feast to kill manual retrains and skew, phase two in weeks six to ten adds a registry, packaging, and canary deploy with Docker and KServe to kill the deploy ticket, phase three in weeks ten to sixteen adds monitoring and drift with Evidently and Grafana to close the loop, and an ongoing phase four adds a self-service layer and reusable templates

    Find your rung honestly. Phase 0 gets training into version control so any model is reproducible. Phase 1 moves training onto a scheduled pipeline with a , so a machine retrains instead of a person remembering to. Phase 2 automates safe deployment, so shipping a version is a promotion in the registry rather than a ticket. Phase 3 wires monitoring's drift signal into the retrain trigger, and the loop finally turns on its own. Phase 4 puts the thin self-service layer on top. The weeks are a realistic order of magnitude, not a promise, and the crucial property is that each phase depends on the one before it.

    The most common self-inflicted wound in this field is a team that believes it is at the top rung while it is actually deploying pickle files by hand. Do not build the automated retrain loop before you have reproducible training. Build the one machine that moves you up exactly one rung, ship it, feel the pain lift, then build the next.

    To make that self-assessment concrete, here is the scorecard the readiness scorer was based on. Print it, mark where you actually are, and let your lowest honest row set your agenda.

    A capstone maturity scorecard as a table with ten ability rows, reproducible training, feature store live, model registry, automated deploy, canary and rollback, live monitoring, retrain loop closed, self-service config, ownership drawn, and cost visibility, and three status columns, missing, manual, and automated, with each row marked to show a realistic mid-journey team that has automated its training, registry, and feature store but is still manual on deploy, rollback, monitoring, and ownership, and missing the closed retrain loop, self-service, and cost visibility

    Every Way a Model Fails, and the Box That Stops It

    There is one more way to hold the whole platform in your head, and it is the one that makes every component feel necessary rather than optional. Read the architecture backwards, as a defense map. Each box you built exists to stop one specific way a shipped model quietly breaks.

    A failure-to-defense map with six rows, each pairing a silent production failure on the left with the platform component that defends it on the right via an arrow, offline metrics looking great while live predictions are worse maps to the feature store, a model rotting from ninety-four to seventy-eight percent while dashboards stay green maps to monitoring and the retrain trigger, an upstream column rename or unit flip maps to data validation, p99 latency blowing the SLA or the GPU bill tripling maps to serving and autoscaling, an unreproducible past result maps to experiment tracking and lineage, and a worse model reaching production with no way back maps to the registry with canary rollback

    Read it as a promise. For every way a shipped model quietly breaks, some box on the golden path is the thing standing in the way. The defends against skew. Monitoring and the retrain trigger defend against silent decay. Data validation defends against a broken upstream job. Serving and autoscaling defend against latency and cost blowouts. Experiment tracking and lineage defend against an unreproducible result. The registry with canary rollback defends against a bad version reaching users with no way home.

    Notice the shared thread: none of these is a bug in the model, and none of them is fixed by a smarter algorithm. Every one is a property of the world the model now lives in, and every one is caught by a box you built. That is the shift from doing machine learning to running it, and it is where the durable engineering value lives.

    How the Giants Assembled It

    None of this is theoretical, and the famous platforms map almost box for box onto the architecture you just traced.

    Uber built Michelangelo, and it is the canonical example of this whole picture assembled into one product. It standardized the entire lifecycle: a (Uber's team helped popularize the very idea), managed training, a model registry, one-click deployment to serving, and monitoring, all behind a self-service interface. Hundreds of teams shipped models on it without rebuilding any plumbing, powering arrival-time estimates, fraud detection, and pricing. Michelangelo is the golden path made real, and Uber later open sourced the feature store as Michelangelo Palette, which fed the ideas behind Feast.

    Netflix built Metaflow, which they open sourced. Its bet is on the self-service layer above all else: a scientist writes ordinary Python, and Metaflow handles versioning, scheduling, moving work to the cloud, and scaling to GPUs underneath. It is the top layer of our stack taken very seriously, sitting on top of Netflix's own compute and data layers.

    Meta built FBLearner Flow, which at its peak trained and served a large share of the company's production models through one pipeline-and-serving platform. Airbnb built Bighead, and Spotify built its platform on top of Kubeflow. Different tools, same architecture, same reason. The model was never the bottleneck. The path from a good model to a reliable service was, and a platform is what paves that path.

    You now know how to build that path. You know every box, what it solves, how they wire together into three interlocking loops, who owns what, when to build versus buy, which failure each component defends against, and which rung to stand on. That is ML platform engineering, and that is the whole job.

    Knowledge Check

    Knowledge Check

    3 questions - Score 80% to pass

    Q1

    In the reference architecture, what single connection turns a one-time model deployment into an ML platform?

    Q2

    What does 'self-service' mean on a mature ML platform?

    Q3

    A five-person startup needs one fraud model in production this quarter and has no platform team. What does the build-vs-buy guidance suggest?

    Serving + packagingLessons 2 and 3, rollout, rollback
    ObservabilityLesson 7Drift, quality, the retrain trigger
    Training + registryLessons 4 and 6Reproducible runs, versioning
    Data + featuresLesson 5Feature consistency, no skew
    ComputeLesson 8Scheduling, GPU sharing, autoscaling

    The two mistakes are symmetric. Do not build all of it on day one for three models. And do not stay fully managed when you are running three hundred and the bill has quietly become someone's whole budget. Start by buying, then self-host the specific layer whose cost or control wall you actually hit.

    That is how every mature platform was actually built, one rung at a time, no matter how finished it looks in the architecture diagram today.