A public-works office wants to know whether this week's checkpoint beats last month's on 'which pump tripped?' What makes that head-to-head valid?
Select an answer to reveal the explanation.
Short Explanation
This week’s checkpoint versus last month’s on “which pump tripped?” Lock the same gold questions, the same context windows, and the same scoring script. Favorite question lists, a longer window for the new run, and mixed hand-versus-script scoring manufacture a fake win.
Full Explanation
A valid checkpoint comparison holds the evaluation constant. The gold questions, the context windows, and the scoring script must be locked so the only changing piece is the model. Different question lists, a longer window for the new run, or mixed hand-versus-script scoring can manufacture a fake win. The office needs one gold set and one script.