Quiz 8 Question 8 of 20

An agent scores 95% on its evaluation benchmark before deployment. In production, the team finds the agent is performing poorly on edge cases involving legacy Python 2 codebases. Investigation reveals the evaluation dataset contained only Python 3 examples. What does this incident illustrate?

Select an answer to reveal the explanation.

Motivation