Quiz 23 Question 6 of 20

You're tuning inference performance for large language models running on NVIDIA GPUs in production, and you need to reduce latency and increase throughput. Which component is most critical for this optimization work?

Select an answer to reveal the explanation.

Motivation