A consultant is helping a sales director evaluate whether an AI-driven opportunity scoring model deployed six months ago is still performing well. The director asks how to determine, on an ongoing basis, whether the model's predicted win probabilities remain trustworthy as market conditions shift. What should the consultant recommend as the primary practice?
Select an answer to reveal the explanation.
Short Explanation
Trusting a model just because it passed its test on day one is a bit like trusting a weather forecaster's track record from last year to predict tomorrow, conditions change, and a model that was sharp six months ago can quietly drift out of step with reality. The way you actually find out if it is still doing its job is by checking its predictions against what really happened: did the deals it called likely to win actually win, and did the ones it flagged as risky actually fall through. That comparison, done regularly, tells you the truth in a way that a single early validation or a few hallway comments from sellers never can, because opinions are inconsistent and a one-time check goes stale. Cranking up how often the model spits out a score does not touch accuracy at all, it just makes more predictions, right or wrong. Ongoing comparison against real outcomes is what keeps you honest about whether the model still deserves the team's trust.
Full Explanation
The correct answer is D. Ongoing trustworthiness of an AI scoring model is established by monitoring performance over time, specifically comparing predicted win probabilities against actual won and lost outcomes, which reveals whether the model's accuracy is holding up or drifting as market conditions and buying behavior change. Option A is incorrect because a one-time validation at deployment reflects conditions at that moment only; markets, product lines, and buyer behavior evolve, and a model that was accurate six months ago can degrade without anyone knowing unless it is monitored continuously. Option B is incorrect because informal, unstructured seller opinions are subjective and inconsistent, and they do not provide the systematic, outcome-based evidence needed to detect real accuracy drift. Option C is incorrect because generating scores more frequently changes how often output is produced, not whether that output is accurate, so it does nothing to validate the model's underlying performance. Comparing predictions to actual outcomes on a recurring basis is the concrete, evidence-based practice that lets the director know whether the model still deserves the sales team's trust.