A vendor's needle grid is not your number. You need a recall-versus-fill curve for your task, on your data, with your distractors. Your production traces already hold the raw material.
Take real runs where you know the correct answer. Group them by how full the context was, for example 10, 30, 50, 70 and 90 percent. In each group, ask the known question again against the real context and score whether the model recovers the fact. Plot recall against fill.
The point where recall starts to slide is the knee, and the knee is your working budget. In the lesson's illustrative model, recall stays within 85 percent of its near-empty level only up to about 30 percent fill. On a 200k window that is roughly 60k usable tokens, not 200k.
The knee is a property of the model, so it moves when the model moves. A provider can ship a new version under the same name and the same window size, and your knee can shift up or down. Re-run the measurement on every model or version change.
To run the degradation model and the budget arithmetic yourself, the Context Rot lesson is the next step.