Quiz 15 Question 19 of 20

A engineering team is training a large language model across an NVIDIA DGX cluster using distributed data parallel training. During monitoring, they notice poor scaling efficiency: as they add more nodes, the GPUs remain underutilized, and training throughput does not increase linearly. Which of the following issues are the most likely causes of this distributed training bottleneck? (Select two)

Select all correct answers, then click Submit.

Motivation