A transit delay model looks strong on last April's depot notes and fails on this January's ice-storm notes from another yard. What eval belongs in the reported experiment?
Select an answer to reveal the explanation.
Short Explanation
Strong on last April’s depot notes, weak on this January’s ice-storm notes from another yard. Report a domain-shifted holdout, not only a random April slice. Peeking at January and then training on it is not a clean test.
Full Explanation
When deployment will not match the training week, a random in-week slice can flatter the model. A domain-shifted holdout from the new season and yard is the experiment that matters. Peeking at those January answers and then training on them is not a clean test. The report should show the shifted score beside the in-week score.