Stellartech Retail has deployed a Copilot Studio agent to handle return and exchange policy questions for store associates. During evaluation review, the team notices that some agent responses are on-topic and drawn from the correct policy documents — but the explanations jump between unrelated points, contradict themselves mid-paragraph, and leave associates more confused than before. An evaluator named Delia wants to flag this specific quality failure using the correct evaluation metric. Which evaluation metric is Delia looking for?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Coherence is like rating a professor's lecture: even if they're teaching the right chapter and the facts are accurate, a lecture that jumps randomly between topics and contradicts itself earlier statements earns a low coherence score. The response at Stellartech is on-topic (not a relevance issue) and source-accurate (not a groundedness issue) — it's just poorly organized. Correct answer: D.
Full explanation below image
Full Explanation
Coherence as an evaluation metric measures whether a response is logically structured, internally consistent, and presented in a way that flows naturally from one point to the next. It captures the quality dimension of 'does this response make sense as a piece of reasoning?' distinct from whether it's accurate or on-topic.
In Stellartech's scenario, the responses are on-topic (the agent is addressing return and exchange policy — relevance is passing) and drawn from correct documents (groundedness is likely passing). The failure is specifically that the response is disorganized and self-contradictory — a coherence failure.
Option D is correct because coherence specifically targets logical flow, internal consistency, and structured reasoning. A response that jumps between unrelated points or contradicts itself within the same answer will score low on coherence regardless of its accuracy or topical relevance. This is exactly the failure pattern Delia is observing.
Option A is incorrect because relevance measures whether the response addresses the user's question — not whether the answer is internally organized. In this case, the agent is answering about return policies (the right topic), so relevance scores would be acceptable. The problem is not topic mismatch; it's structural chaos within a topically correct response.
Option B is incorrect because groundedness measures source fidelity — whether claims are supported by knowledge base content. The scenario specifies that the information is drawn from the correct policy documents, so groundedness is not the primary issue. An agent can produce a grounded but incoherent response.
Option C is incorrect because fluency measures linguistic quality at the sentence level — grammar, word choice, and readability. The issue described is not poor grammar but logical disorganization and self-contradiction, which is a higher-order structural problem. A response can be grammatically fluent while being completely incoherent in its reasoning.
Exam tip: To distinguish the four metrics on the exam: Groundedness = did it come from the source? Relevance = did it answer the question? Coherence = does it make logical sense? Fluency = is the language clean? Scenario keywords like 'contradicts itself', 'jumps between topics', 'confusing organization' all point to coherence.