While writing the next word of a lock announcement, a decoder must not already see the words it has not written yet. Which rule enforces that?
Select an answer to reveal the explanation.
Short Explanation
A decoder writing a lock announcement cannot peek at words it has not written yet. Causal masking is the rule: look only at yourself and what already came out. Bidirectional attention is the encoder’s trick, and batch size or a CUDA barrier will not hide the future.
Full Explanation
Causal masking stops a decoder from attending to tokens it has not generated yet. Bidirectional unmasked attention is encoder behavior. Batch size and CUDA barriers are not that rule.