Quiz 14 Question 6 of 20

An AI agent built for an HR department is flagged because it frequently answers questions about company benefits by providing accurate historical data from its knowledge base, but users report the answers do not address what they actually asked. For example, a user asking about dental coverage receives a detailed response about vision benefits instead. Which Azure AI Foundry evaluation metric is most appropriate for diagnosing this agent quality problem?

Select an answer to reveal the explanation.

Motivation