Quiz 27 Question 10 of 20

An AI operations team notices that during peak hours, the latency of their LLM inference service spikes due to a sudden surge in user requests. Which Kubernetes component should they configure to automatically spin up additional replicas of their inference pods to handle the increased load based on CPU, memory, or custom metrics?

Select an answer to reveal the explanation.

Motivation