Quiz 10 Question 11 of 20

A PySpark job joins a very large fact table of digitised-object records against a small lookup table of the network's dozen branch codes and names. The job spends most of its time shuffling the huge fact table across the cluster just to match it with a handful of lookup rows. Which technique is most directly aimed at eliminating that unnecessary shuffle?

Select an answer to reveal the explanation.

Motivation