Quiz 19 Question 13 of 20

An infrastructure manager at a hyperscale AI data center running thousands of NVIDIA Tensor Core GPUs wants to transition from reactive troubleshooting to a proactive operations model. The goal is to detect early-stage hardware anomalies and predict GPU failures before they cause training jobs to crash. Which strategy should they implement to achieve this?

Select an answer to reveal the explanation.

Motivation