Quiz 37 Question 12 of 20

You are building an on-premise GPU cluster to support multi-node deep learning training. You need a platform that can manage the complete machine learning lifecycle, schedule distributed training jobs across multiple servers, maximize GPU utilization, and handle node failures automatically. Which tool is designed specifically for this orchestrating role?

Select an answer to reveal the explanation.

Motivation