Run the corpus and you get a pass or fail for each attack. Attack success rate, or ASR, is the fraction of attacks that got through. It is the one number a security test tracks over time.
You split ASR by family, because the families do not carry the same harm. In the lesson, a merge check treats them differently. A single leaked canary or a single forbidden tool call blocks the merge, because the damage is real. A jailbreak that produces bad text but cannot act gets a small budget instead, which you push down over time. Borderline results go to a human.
One overall number can also hide problems. Suppose one threat has zero attacks in the corpus. Its success rate shows as zero, which looks perfect, but it means untested, not safe. That is why the lesson lays ASR against the threat catalog, one row per threat, so empty rows stand out.
If this sounds like model evaluation, it is. The corpus is a golden dataset of adversarial cases. The success check is the scorer. The only new step is running it as a required check in CI/CD, so a change to the prompt, the tools, the model version or the guardrails reruns every attack.