Apex Capital Management has deployed a Copilot Studio agent to assist wealth advisors with client portfolio questions. During evaluation, the team notices that the agent sometimes responds with technically accurate information about investment concepts — but the information does not actually address what the advisor asked. The agent answers a question about bond duration by explaining equity volatility. The evaluation team needs to select a metric that will flag this specific type of failure. Which evaluation metric directly measures whether the agent's response addresses the user's actual question?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Relevance is like a teaching assistant scoring whether a student answered the actual exam question — not whether the answer was accurate in isolation, but whether it directly addressed what was asked. If a wealth advisor asks about bond duration and gets a lesson on equity volatility, relevance will score that low even if the equity content is factually perfect. Correct answer: A.
Full explanation below image
Full Explanation
Relevance as an evaluation metric specifically quantifies the degree to which an agent's response addresses and is pertinent to the user's original query. It captures the 'did the agent answer what was actually asked?' dimension of quality. The scenario at Apex Capital Management — where technically accurate content fails to address the specific question — is a textbook relevance failure.
Option A is correct because relevance is explicitly designed to surface this type of response failure. Even if the agent's response about equity volatility is perfectly grounded in the knowledge source, well-written, and coherent, it scores low on relevance because it misses the point of the question. AI evaluation frameworks score relevance by comparing the semantic relationship between the user query and the agent response.
Option B is incorrect because groundedness measures source fidelity, not query alignment. A response can be perfectly grounded (every sentence traceable to a document) while being completely irrelevant to what the user asked. In this scenario, the equity volatility answer may be grounded in a real knowledge base article — it just doesn't address the bond duration question.
Option C is incorrect because fluency measures linguistic quality: grammar, sentence clarity, and readability. A response can be fluent and irrelevant simultaneously — smooth prose does not guarantee topical match. Fluency says nothing about whether the right topic was addressed.
Option D is incorrect because coherence evaluates whether the response's internal logic flows consistently. A coherent response is one where the ideas connect logically — but coherence is indifferent to whether those ideas match the user's query. The equity volatility response could be internally coherent and still score zero on relevance.
Exam tip: Remember the practical distinction — relevance answers 'Did the agent answer the right question?', groundedness answers 'Did the agent make things up?'. Scenarios describing off-topic responses, query mismatch, or 'accurate but irrelevant' content always point to relevance as the key metric.