Quiz 37 Question 10 of 20

Your team needs to run a large language model (LLM) for real-time customer support, but you are limited to a single, lower-spec NVIDIA GPU with limited VRAM. Which of the following approaches is most effective for shrinking the model's memory footprint and speeding up its response times (inference latency)?

Select an answer to reveal the explanation.

Motivation