A loom-thread tension model looks strong on the validation slice. Testers run an independent test pile that includes uncommon weave patterns the fit never saw. The test score is much worse than validation. What should that gap signal?
Select an answer to reveal the explanation.
Short Explanation
Strong on the validation slice, then much worse on an independent test pile that includes uncommon weave patterns the fit never saw. That gap signals overfitting. Underfitting would have been poor on training and validation already, static drift testing is a distribution-shape watch, and a strong validation slice is not the official score.
Full Explanation
Detect overfitting when independent test performance, including less common examples, is significantly worse than validation. Reading the gap as underfitting, as static drift, or as a pass misses that signal.