A cascade is only as good as the check that decides accept or escalate. The lesson is direct about which check to trust.
Do not escalate on the model's own confidence. If you ask the cheap model how sure it is, the number often does not match whether the answer is right. Models tend to be overconfident.
Escalate on a validator. A validator is a check in code on the cheap answer. Does the JSON parse? Does the code compile and pass its tests? Does the field match the schema? Is the claim supported by the source document? In the lesson's illustrative example of 1,000 requests, both signals escalate the same 220. The validator catches 180 of the 200 real failures. Self-reported confidence catches 60. The cost is the same, and the validator catches three times as many.
Cost is easy to write down. Count one cheap call as 1 and one strong call as R, the price ratio. Every request pays 1. A fraction p escalates and also pays R. So a cascade costs 1 + p × R per request, against R for sending everything to the strong model.
In the lesson's illustrative model, with R at 16 and 90 percent easy traffic, the cascade cost about 1.35 dollars per thousand requests against 7.20 for all-strong, with accuracy of 0.931 against 0.963. As traffic gets harder, p rises. Near 48 percent easy traffic, the cascade escalates about 44 percent of requests and costs more than half the all-strong bill. At that point the cheap first pass is mostly wasted work. A cascade is a bet that most requests are easy.