Quiz 3 Question 3 of 20

During pretraining a transformer can score many token positions of a packed sequence in one pass, but at serving time a decoder still emits tokens one after another. How do those time-shapes differ?

Select an answer to reveal the explanation.

Motivation