Quiz 19 Question 6 of 20

An operations engineer manages a shared GPU cluster running a mix of training and inference workloads. The job queue contains critical, real-time jobs alongside flexible, multi-node training tasks that can utilize parallel GPU execution, and single-GPU batch jobs. Which queue scheduling policy will minimize latency for critical workloads while maintaining high overall cluster utilization?

Select an answer to reveal the explanation.

Motivation