Quiz 16 Question 14 of 20

An agent that automatically triages GitHub issues fails on 15% of runs. The failures manifest as the agent marking issues as 'needs more info' when it should have escalated them as 'critical.' A developer wants to determine whether the root cause is: (A) the agent's reasoning is flawed, (B) the tool that queries issue metadata is returning incomplete data, or (C) the agent's context about severity classification is outdated. Which diagnostic approach would MOST efficiently distinguish between these three failure classes?

Select an answer to reveal the explanation.

Motivation