Quiz 36 Question 5 of 20

An AI operations team is deploying a large language model (LLM) for real-time customer service chatbot inference on an NVIDIA GPU. The team needs to maximize the system's inference throughput (tokens processed per second) while maintaining acceptable latency. Which of the following adjustments will most directly optimize GPU execution efficiency and throughput?

Select an answer to reveal the explanation.

Motivation