An organization's AI model monitoring system shows stable AUC and F1 scores for a consumer lending model over 12 months. However, the risk team receives complaints from relationship managers that the model's recommendations are increasingly inconsistent with their expert judgment on similar cases. What does this discrepancy MOST suggest?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Here's the deal — b is correct because aggregate performance metrics like AUC and F1 can remain stable at the population level while the model underperforms for specific subpopulations or exhibits qualitatively different behavior than it did historically. Relationship managers working with individual cases daily may detect shifts in model reasoning that aggregate metrics do not reveal.
Full explanation below image
Full Explanation
B is correct because aggregate performance metrics like AUC and F1 can remain stable at the population level while the model underperforms for specific subpopulations or exhibits qualitatively different behavior than it did historically. Relationship managers working with individual cases daily may detect shifts in model reasoning that aggregate metrics do not reveal. This is a monitoring gap where human expert feedback is providing risk intelligence that automated metrics miss. The relationship managers' concern deserves investigation, not dismissal (A, C). Switching metrics (D) addresses measurement selection, not the discrepancy between metrics and expert observation.