Quiz 11 Question 8 of 20

Your operations team is tasked with monitoring a large-scale AI infrastructure where multiple GPUs are running heavy training workloads in parallel. Since these GPUs are highly interdependent, a slowdown on one card can cause the entire training run to stall. Which two metrics are most essential to monitor on the GPUs to ensure optimal performance and catch bottlenecks early? (Select two)

Select all correct answers, then click Submit.

Motivation