Ice-rink scoring still returns 200s, but p99 model latency crosses the dispatch budget before CPU looks high. Ops wants capacity to follow latency, not a vanity CPU chart. Which auto scaling metric fits?
Select an answer to reveal the explanation.
Short Explanation
Scoring still returns 200s, but p99 latency crosses the dispatch budget before CPU looks high. Scale on model latency. CPU-only misses the SLA, and this is Domain 3 metric selection, not a Domain 4 troubleshooting essay.
Full Explanation
Choose model latency as the auto scaling metric when the SLA is latency and CPU or invoke charts are not the first signal. That is metric selection for scaling, not Domain 4 troubleshooting. Rekognition is not that metric.