Quiz 12 Question 10 of 20

You're tuning inference performance for large language models running on NVIDIA GPUs in production, and you need to reduce latency and increase throughput. Which component is most critical for this optimization work?

Select an answer to reveal the explanation.

Motivation