Decomposition changes how many steps must be right in a row. Take the twenty step task at 0.95 per step, about 36 percent end to end. Split it into four segments of five steps, with a check at each boundary. When a segment fails and the check catches it, only that segment is retried.
With a strong verifier that catches 90 percent of bad segments, the lesson's model lifts the task from about 36 percent to roughly 75 percent. It costs about 1.4 times the tokens. A failure caught at a boundary also retries five steps instead of twenty.
The lift lives in the catch rate, not in the act of adding a check. With a catch rate of zero, which is roughly what a self-opinion gives on correctness, task success stays at the baseline. You still pay for every gate. The lesson calls this reflection theater.
More splitting is not always better. Every boundary adds a verifier call, extra context, and one more place where the split itself can be wrong. In the lesson's second model, success peaks at around four or five steps per segment and falls after that. Token cost is lowest at that same point. Split at the natural checkable seams, and stop.