Quiz 13 Question 12 of 20

You are managing an enterprise AI data center where several GPU nodes have experienced unexpected shutdowns and hardware degradation during prolonged 72-hour deep learning training runs. To implement a proactive alerting system and prevent hardware failure from thermal stress, which telemetry metric must you monitor and set critical thresholds for?

Select an answer to reveal the explanation.

Motivation