A parks-permit desk declares model B the winner, but A was scored on last year’s questions with greedy decoding and B on this year’s questions with a looser sampler. What is wrong with that comparison?
Select an answer to reveal the explanation.
Short Explanation
A on last year’s questions with greedy decoding, B on this year’s with a looser sampler. That is not a model comparison. Hold the split, the metric script, and the generation settings constant. Fairness is an eval protocol, not an NCCL setting.
Full Explanation
A fair comparison holds the split, the scorer, and the decoding settings constant so only the model — or the intended treatment — changes. Scoring A on last year’s questions with greedy decoding and B on a new set with a looser sampler mixes three factors. The winner is then a protocol artifact, not a model result. Sampler-always-wins, leaked forever-tests, and NCCL “fairness” do not repair the protocol.