Quiz 14 Question 20 of 20

An operations team notices that several nodes in a GPU-accelerated Kubernetes cluster are experiencing performance bottlenecks, while other nodes sit completely idle. To properly diagnose the load imbalance and gather granular GPU metrics like Tensor Core activity, memory usage, and SM utilization, which tool should they implement?

Select an answer to reveal the explanation.

Motivation