A library system's AI recommendation tool is being evaluated using patron satisfaction survey scores, and the surveys are administered and summarized by the same library staff who manage the recommendation program. What is the biggest risk this evaluation approach creates?
Select an answer to reveal the explanation.
Short Explanation
Think of it like a teacher grading their own test: even with good intentions, when the people running a program also collect and summarize the feedback on it, the numbers tend to drift favorable. You want a metric the program owners can't quietly shape.
Full Explanation
When the same team that operates an AI tool also controls how satisfaction data is collected and reported, the metric is vulnerable to selection bias, favorable framing, or informal pressure on responses, even without deliberate manipulation. The sound fix is to pair or replace the self-reported score with data the program team doesn't control, such as circulation counts tied to recommended titles, repeat usage rates, or an independently administered survey. Treating the current scores as sufficient ignores the structural conflict of interest baked into who produces them. Simply surveying more often doesn't fix a biased collection process; it just generates more of the same skewed data with a thinner veneer of statistical confidence. Routing results through the director for review adds another internal party rather than an independent check, so it doesn't resolve the underlying conflict either. The caveat: patron satisfaction itself isn't a bad metric — the concern is who controls its measurement, not the concept. As an operational check, ask whether a metric could still be verified the same way if the program's budget depended on the result being favorable.