This idea carries a full system design question on its own. Each walks through the full answer.
A team builds a support agent. You type "the customer says their order never arrived, sort it out," and the agent looks up the order, checks the shipping status, and issues a refund. In the demo it works perfectly. Everyone claps.
Two weeks after launch, an on-call engineer gets paged. One customer has been refunded four times for the same 42 dollar order. The logs tell the story. The agent called issue_refund, the payment API timed out, the agent read the timeout as "that did not work," and it tried again. And again. And again. The refund actually went through every single time. The API was slow, not broken.
Nobody wrote a bug. The code was fine. The model made a reasonable-looking decision at every step, and the sum of those reasonable decisions was four refunds and an angry finance team. This is the defining property of an agent in production: it does not fail loudly with a stack trace, it fails quietly by making a chain of locally sensible choices that add up to a bad outcome.
A chatbot that gives a wrong answer wastes a few seconds. An agent that takes a wrong action moves money, deletes data, or emails the wrong person. The stakes are completely different.
The design of the agent loop is covered elsewhere in this track. This lesson is the operator's counterpart: not how an agent is built, but how you run one without fear. Deployment shape, the budgets that bound a run, the tracing that makes a failure legible, the rollout process for a change you cannot prove safe in staging, and the handful of numbers that tell you a fleet of agents is healthy. This is the last lesson in the LLM and GenAI Ops chapter, and it is the one that separates a slick demo from something you can run on a Friday night.
The first thing to unlearn is that an agent is a clever prompt. In production it is a service, and it has the anatomy of one. A user message arrives at a gateway that authenticates it, rate limits it, and stamps it with a request id. The gateway turns that message into a run, and it attaches a budget to the run before a single token is generated. From there the request moves through an orchestrator that executes the loop, a governor that can stop it, a sandbox where tools actually run, and a model plane that routes and caches the model calls. Underneath all of it sits an observability plane that taps every step.

Read the planes as a division of responsibility. The control plane is code you own end to end: the gateway, the orchestrator that builds each prompt and reads each tool call, and the governor that is the only component allowed to stop a run. The execution plane is where side effects happen, and every one of them is contained: tools run in a sandbox with a timeout and least-privilege credentials, and the tools that change state or spend money sit behind a tighter scope than the tools that only read. The model plane treats the model as an interchangeable dependency, routing cheap steps to a small model and hard steps to a frontier one, failing over to a second provider so a single vendor incident does not take the agent down, and the stable system and tool block so the prompt you rebuild every turn is billed at a fraction.
The plane that people skip, and then regret skipping, is the observability plane. Because an agent fails quietly, "the model got it wrong" is a useless bug report. You need to see which of a dozen steps went sideways, what the model was thinking, which tool it chose, what came back. The trace is what turns that four-refund mystery into a one-line answer. Everything after this slide is one of these boxes, examined up close.
A bare agent loop has no natural end. The model keeps proposing actions until it decides the goal is met, and a confused model may never decide that. A normal endpoint returns exactly once; an agent decides its own length, which means without a boundary it can run forever. So in production you wrap the loop in a governor, and you make the governor the single component with the authority to halt a run.

The governor checks three budgets before every model call: how many steps the run has taken, how much money it has spent, and how long it has been running on the wall clock. The loop continues only if all three still have room. If any one is exhausted, the run does not crash and it does not spin, it halts, checkpoints its state, returns an honest "I could not finish this," and escalates to a human. A safe abort is far better than a silent loop. An agent that stops at 15 steps and admits defeat is worth ten of an agent that spins 400 times and bills you for every spin.
Here is the loop with the governor made explicit. This is the skeleton every agent framework runs, stripped to the part that matters for operating it.
Notice that the budgets are not constants you copy from a blog post. You set them from your own traces. Look at the distribution of steps and spend across real runs, then put each cap a little above p95. Set them too tight and you kill honest long tasks; set them too loose and a runaway bills you before it trips. The cap is a product decision informed by data, and it is the first line of the whole platform.
A normal endpoint has a roughly fixed cost per request. An agent does not, and this breaks the intuition of anyone pricing one for the first time. You cannot predict the cost of a run because you do not know how many steps it will take, and worse, each step costs more than the last. The model is stateless, so on every turn the orchestrator rebuilds the whole prompt, and that prompt includes the entire transcript so far. A run that takes ten steps rereads a growing history ten times, and the later reads are the most expensive.

Read the shape. The curve is not a straight line, it bends upward, because the cost of step twenty carries the weight of the whole nineteen-step transcript before it. A typical healthy run finishes in four to seven steps for a few cents, nowhere near either cap. But the two backstops from the previous slide are not redundant. The step cap catches most bad runs early, at a fixed number of iterations. A pathological run, though, can stay under the step cap and still get expensive, because the late steps are so heavy. The spend cap is the second net: it trips that run around step twenty-one, long after the honest ones have gone home. One cap is not enough; you want both.
The lever that flattens the whole curve before either cap has to fire is prompt . The system prompt and the tool descriptions are stable across turns, so a provider can cache them and bill the repeated read at a fraction of the input price. On a long agentic run, where the same tool block is re-sent on every one of a dozen turns, caching is not a micro-optimization, it is the difference between a sustainable cost per run and a bill that grows with the square of the task length. Measure cost per run, not just total spend, and alert on the p99, because a cost blowup shows up as a fat tail on a slice of traffic long before it moves the average.
A request-response service fails in ways you already know how to operate: a 500, a stack trace, a saturated queue, a page. An agent adds a set of failure modes that a normal service simply does not have, and most of them are quiet. The skill of operating agents is knowing what each one looks like in your telemetry and which control contains it, so you can sort a symptom to a fix instead of reaching for "use a bigger model" every time.

Walk the high-severity ones. A runaway loop shows up as a step count climbing past p95 while the same tool is called over and over; it burns tokens and rate limit and blocks a worker, and the fix is a step cap plus a no-progress . A cost blowup shows up as spend per run spiking on a slice of traffic while the step count looks ordinary; a handful of pathological runs can outspend all the healthy ones combined, and the fix is a per-run spend cap plus . A misread tool error is the four-refund bug: a tool returns a 504 or a timeout, the agent reads it as failure, and it retries a write that already succeeded. The blast radius is duplicate side effects, and the fix is keys on every write so that a slow success is safe to retry.
The quieter modes are the dangerous ones because nothing throws. SLO breach is a p95 that drifts up as tasks need more steps, contained by a wall-clock budget and by streaming partial progress to the user. Silent quality drift is the worst of all: nothing errors, but escalations and thumbs-down creep up after a model or prompt change, and you find out late because no exception is ever raised. The only defense is to treat escalation rate as a first-class metric and run online evals on sampled runs. And an orphaned side effect is a run that aborts mid-task after a write, leaving the refund posted but the ticket never closed, contained by durable execution that logs each step and can resume or compensate on restart. Read the last column and notice what is absent: not one entry is "use a smarter model." Every control is an engineering wrapper around a model that will keep making the occasional wrong call.
Users feel , and an agent feels slow for a specific structural reason. A run is not one model call, it is a chain of them, and they run in sequence because each step depends on the result of the last. You cannot parallelize a chain of decisions where step three is chosen based on what step two returned. So the seconds stack.

Look at where the time goes. The tools are fast, tens to a few hundred milliseconds each. The seconds are the model calls, and roughly ninety percent of a run's wall-clock time is the model thinking across four sequential calls. This is why the highest-leverage latency optimization is not a faster model or a faster tool, it is fewer steps. Cutting a task from six steps to four removes a whole model call from the critical path, which dwarfs any speedup you could win on an individual call.
The number you operate against is the tail, not the mean. A p50 of about seven seconds feels fine to a user. But the p95 is what times out an upstream client, abandons a session, and triggers a cascade of client retries that make everything worse. Here the p95 sits at 10.8 seconds against a 12-second SLO, close enough that a longer class of tasks would breach it. Budget for the tail: cap the wall clock in the governor, cut steps wherever the task allows, run genuinely independent tool calls in parallel, and above all stream tokens to the user so they see motion long before the run is actually done. A user watching output appear tolerates far more latency than a user watching a spinner.
When an agent misbehaves, the final answer contains almost no information about which step failed. So you record the run the way you would record a distributed transaction: as a tree of spans, one span per model call, per tool call, per retry, each carrying its inputs, outputs, token count, , and cost. Modern agent observability tools like LangSmith and Langfuse, and the OpenTelemetry GenAI semantic conventions, exist to capture exactly this tree.

The tree on the left is the four-refund incident, finally legible. Open the run and the story is right there: step two calls issue_refund, the tool returns a 504 after two seconds, and a retry fires under the same key and comes back deduped, confirming the customer was charged once, not twice. Without that trace you are grepping raw logs and guessing which of nine steps moved the money. With it, the diagnosis takes ten seconds instead of two hours. The trace is not a nice-to-have; it is the primary operational surface for an agent.
The right side is the pipeline that turns spans into something you can page on. The instrumented run emits spans, an OpenTelemetry exporter ships them in a vendor-neutral format, a collector samples and batches and fans them out to a trace store and a metrics store, and dashboards render both. The design principle to hold onto is that a metric and a trace are two views of the same event: when a golden signal leaves its band, you should be one click away from the exact run behind the number. A dashboard that shows you a bad number but cannot show you the run that caused it is only half an observability system.
The single highest-leverage control on a risky agent is a human gate, but you cannot pause on everything or you have destroyed the automation you built. The resolution is to classify every proposed tool call by risk, and let the class decide whether a human is in the loop. Read-only lookups run freely because they have no side effect. Anything that moves money, deletes data, or emails a customer pauses and surfaces the exact call, with its exact arguments, to a person.

This is the pattern that lets a support agent handle thousands of tickets while a human still owns every payment. The read-only calls, which are the large majority, never wait. The side-effecting calls go to an approval queue where a person can approve the call so it runs exactly as drafted, reject it, or let it time out. The important design detail is what happens on a reject: it is not an error and it does not crash the run. The rejection is fed back to the model as an observation, "that action needs approval," so the agent can reason its way to a safer alternative instead of dying. should steer the agent, not just kill it. And a time-out must not block a worker forever, so an unattended paused run parks in a durable state and pages rather than holding a thread hostage.
Two properties make this safe rather than theatrical. First, the risk class of a tool is configuration your code owns, never something the model decides for itself; you do not ask the agent whether its own action is dangerous. Second, the human sees the concrete call and arguments, not a vague summary, so the approval is a real review and not a rubber stamp. This is how you let an agent draft a five thousand dollar refund without ever letting it issue one unsupervised.
The step cap guarantees a run ends, but at 15 wasted steps it is a blunt instrument. A confused agent almost always shows a pattern long before it hits the cap: it calls the same tool with the same arguments, gets the same result, and makes no progress toward the goal. A watches for that stall and trips early, so you pay for four spun steps instead of fifteen.

The strip on the left is the four-refund pattern seen as a stall: the model misread a slow refund as a failure and looped, calling the same two tools with the same arguments and learning nothing new. The breaker fires at step five, when the no-progress window fills, rather than letting the run grind to the hard cap. The trip conditions are heuristics, and any one is enough: the same tool and arguments repeated three times, tool results that stop changing the state the model can see, spend accelerating faster than progress, or a wall clock past its budget. On a trip, the breaker does the same graceful thing the governor does: halt, checkpoint, escalate, and tell the user honestly that the task could not be completed.
The detector is a small amount of code that lives in the orchestrator, tracking a short window of recent tool signatures.
Ship both limits, because they do different jobs. The step and spend caps are hard guarantees that a run terminates no matter what. The circuit breaker is a cheaper, earlier heuristic that catches the common stall before the hard cap has to. The breaker saves money and on the frequent case; the caps guarantee safety on the case the breaker misses.
Reliability in an agent compounds instead of averaging, and this is the fact that surprises every team. A run succeeds only if every step succeeds, and for roughly independent steps that means multiplying the per-step rates. A 97 percent step sounds excellent, but over ten steps the whole-task success is 0.97 to the tenth, about 74 percent. Roughly a quarter of runs have at least one bad step. Left alone, that missing quarter is invisible: wrong actions, half-finished tasks, quiet retries you never see.

The job of a production platform is not to make that quarter vanish, it is to make it legible. Wrap the same traffic in the controls from this lesson and every run lands in a bucket you can count. Retries absorb the transient errors and move those runs into "recovered." The human gate escalates the calls the agent should not take alone. The budgets abort the pathological runs safely. Nothing disappears into silence; the 2,000 runs in a day split into five named outcomes, each with an owner and an alert. "74 percent, and I cannot see the rest" becomes five numbers on a dashboard.
Once failures are legible, you can attack them. The two ways to move the buckets left are fewer steps and a higher per-step rate, and shortening the task is almost always the cheaper lever, because you are fighting an exponent. You cannot prompt your way out of 0.97 to the tenth; you can decompose a twenty-step task into two ten-step subtasks with a checkpoint between them, and turn one long fragile chain into two shorter robust ones. But even before you shorten anything, converting silent failure into visible outcomes is the win that makes the rest possible.
A normal service is watched through the four golden signals: , traffic, errors, and saturation. An agent needs a few of its own, because unlike a normal service, its cost and its correctness are not fixed per request. These are the signals to put on the wall and page on. The discipline that matters is to watch the band, not the single run: one bad run is noise, a drifting band is an incident.

The first three are cheap to measure and catch the loud failures. Steps per run flags a stalling change before anything else does; a rising p95 here is your earliest warning. Spend per run catches the cost blowup, and you watch its p99 rather than its average because the blowup lives in the tail. p95 run latency is your SLO signal, watched against the tail and not the mean for the reasons the latency slide laid out.
The other three are proxies for the failures that do not throw. Escalation rate, the fraction of runs that hit the human queue, moves when the agent is out of its depth, and a jump is a signal that quality has drifted even though nothing errored. Abort rate, the fraction a budget stopped, spikes when a runaway pattern appears. Tool error rate catches a flaky downstream before it drives a wave of retries; on this dashboard it is the one signal drifting up, flagged as a watch, which is exactly the kind of early creep you want to catch before it becomes an incident. Alert on the band each signal should sit in, sampled over a window, so a single unlucky run never wakes anyone at two in the morning.
A one-word change to a prompt, a new model version, or a tweaked tool description can swing behavior across every run. And because the effect is non-deterministic, you cannot prove the change safe in staging the way you would prove a pure function correct with a unit test. The same input can produce a different chain of steps today than it did yesterday. So you roll a change out the way you roll out code: gradually, on real traffic, watching the signals decide.

Start in shadow, where the new version runs beside the live one on real requests but its output is scored, never served to a user. If its outputs agree with or beat the incumbent, promote to a small live canary slice, then a wider ramp, then full, and at each stage watch escalation, abort, cost, and p95 against the version that is currently live. The gate that promotes is the same gate that reverts: at any stage, if a signal breaches its band against the baseline, traffic snaps back to the last known good version automatically, before a human is even paged.
The subtle point is what the baseline is. You are not asking "is the new prompt good," which is unanswerable in the abstract, you are asking "is it at least as good as what is already live," which your signals can answer on real traffic. That comparison is the entire reason shadow mode exists: it scores the candidate on the same requests the incumbent just handled, so the comparison is apples to apples. A prompt change without a rollout process is a bet you cannot see the odds of; a canary turns it into a measured, reversible deploy.
None of this is hypothetical. The teams shipping agents in 2024 and 2025 converged on the same playbook, usually after getting burned once. Klarna put an assistant in front of customer support and reported it doing the equivalent work of hundreds of full-time agents in its first month. The headline was the volume; the engineering story was the containment, tight tool scopes and human escalation on anything touching a payment. Klarna later pulled some of that work back to human staff, which is its own lesson: the automation is only ever as good as the and the quality bar behind it.
GitHub Copilot grew from autocomplete into agentic workflows that can take an issue, plan a fix, edit files, and open a pull request. Notice the guardrail baked into the design: the agent produces a pull request, and a human reviews and merges. The risky action, changing the main branch, stays behind human approval. And when Anthropic and OpenAI both ship agent frameworks whose core features are step limits, tracing, and human-in-the-loop approval hooks, the model makers are telling you where the hard part is. It is not the cleverness of the loop; it is the control around it.

Here is that convergent playbook as a go-live gate, the checklist to satisfy before an agent touches a real credential. Step and spend budgets set from your own traces, so one confused run cannot loop forever and bill you for every spin. Human approval on the risky tools, so the agent does the volume while a person owns the irreversible actions. Sandboxed tools with least-privilege credentials, so a wrong or prompt-injected call cannot reach secrets or production. Full-run tracing, so a failure is a ten-second diagnosis. Canary rollout with auto-rollback, so a non-deterministic change ships like code. And alerting on the golden signals, so the quiet failures surface in escalation and abort rate before a customer finds them.
The impressive part of an agent is the loop. The part that lets you run it on real credentials on a Friday night is everything around the loop. Build these six and an agent becomes a tool you can trust with money and data. Skip them and you get four refunds and a page at two in the morning. Nothing on the list is about a smarter model, and that is the whole point of operating agents in production.
3 questions - Score 80% to pass
Why is a governor that caps steps, spend, and wall-clock time considered essential rather than a nice-to-have for a production agent?
A single agent run to 'refund an order' costs far more than a simple chatbot reply, and later steps cost more than earlier ones. Why?
After a model change, nothing in the agent throws an error, but over a week escalations and thumbs-down creep up. Which control is designed to catch this?