Quiz 35 Question 20 of 20

You are deploying a large language model (LLM) to handle real-time user queries, and the current latency is too high for a good user experience. Which of the following optimization techniques will most effectively reduce the time-to-first-token and overall inference latency on your NVIDIA GPU?

Select an answer to reveal the explanation.

Motivation