Quiz 28 Question 8 of 20

When optimizing a Large Language Model (LLM) for low-latency inference, which compression technique offers the most effective balance between shrinking model footprint and maintaining high accuracy?

Select an answer to reveal the explanation.

Motivation