Exploratory work shows that duplicate and missing source records explain most of an apparent performance spike, yet modelers want to proceed with a complex predictive model anyway. What sequencing decision is correct?
Select an answer to reveal the explanation.
Short Explanation
If dirty data is doing most of the talking, building a fancy model is like polishing a cracked windshield and then racing. Fix the duplicates and missingness at the source, then model what's left. Otherwise you're predicting the defects, not the service.
Full Explanation
When quality defects dominate observed variation, models primarily fit noise and process errors rather than the intended phenomenon. Correct sequencing remediates source data fitness before investing in complex modeling. This preserves analytic resources and avoids operationalizing flawed signals.