A community observatory wants educators to pick the better of two star-chart explanations, then those preferences shape later behavior. Which loop is that?
Select an answer to reveal the explanation.
Short Explanation
Think of two star-chart writeups and a docent circling the better one. Those preference picks steer later behavior. That is RLHF, not unlabeled next-token training, a cache, or a dashboard vote.
Full Explanation
RLHF uses human preference comparisons to steer a model. It is distinct from next-token pre-training. Caching and a dashboard vote are not that preference loop, and the item does not require algorithm coding.