Quiz 15 Question 14 of 20

A data center engineer deploys NVIDIA BlueField DPUs (Data Processing Units) in a GPU cluster to optimize infrastructure operations. However, the development team notices that when they offload neural network inference tasks directly to the DPU's general-purpose ARM cores, the inference latency spikes dramatically compared to running them on the host CPUs or GPUs. What is the underlying cause of this performance bottleneck?

Select an answer to reveal the explanation.

Motivation