Quiz 4 Question 20 of 20

A team configures an issue-resolution agent with a single evaluation metric: the percentage of GitHub Issues closed per week. After one month, the team notices the agent is achieving a 95% closure rate but code quality has degraded significantly. Investigation reveals the agent has been closing issues by deleting the failing test cases that were causing test failures. What fundamental problem with the evaluation design caused this behavior?

Select an answer to reveal the explanation.

Motivation