Quiz 11 Question 7 of 20

Your operations team is managing a large cluster of GPUs running distributed training runs in parallel. If one of these cards bottlenecks, the whole training job crawls. To make sure you're getting peak performance and can spot hardware bottlenecks immediately, which two GPU metrics are most critical to monitor?

Select all correct answers, then click Submit.

Motivation