Pinnacle Manufacturing has deployed a Copilot Studio agent to assist maintenance engineers with equipment troubleshooting. The agent uses a curated SharePoint library of OEM service manuals as its knowledge source. During evaluation, the QA engineer Lena observes two different failure modes: in some cases, the agent responds with accurate answers that don't actually address the engineer's question; in other cases, the agent's response addresses the question but includes claims that cannot be traced to any manual in the library. Lena needs to choose separate evaluation metrics to detect each failure mode independently. Which statement correctly identifies the key distinction between groundedness and relevance?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Groundedness and relevance are two different lenses on quality. Think of groundedness as the librarian's citation check — 'Does every claim trace back to a real document?' — and relevance as the reader's satisfaction check — 'Did this actually answer my question?' Lena needs both to catch the two separate failure modes she observed. Correct answer: B.
Full explanation below image
Full Explanation
This question tests a critical distinction that frequently appears on the AB-620 exam: understanding that groundedness and relevance are orthogonal quality dimensions, each measuring a different axis of agent response quality.
Option B is correct. Groundedness answers the question: 'Is the content of this response supported by the knowledge source?' It specifically targets the hallucination failure mode — where an agent generates claims not traceable to any source document. In Lena's scenario, the failure where the agent includes claims not found in any OEM manual is a groundedness failure.
Relevance answers a different question: 'Does this response actually address what the user asked?' It targets the mismatch failure mode — where an agent produces accurate, source-supported content that doesn't fit the user's intent. In Lena's scenario, the failure where responses are accurate but don't address the engineer's question is a relevance failure.
Option A is incorrect on both counts. Groundedness does not measure logical organization — that is coherence. Relevance does not measure tone matching — that is outside the scope of standard AI evaluation metrics in Copilot Studio.
Option C is incorrect because it reverses the definitions. Groundedness does not measure whether the response answers the question (that is relevance), and relevance does not measure source citation accuracy (that is groundedness).
Option D is incorrect because groundedness and relevance are not the same metric and do not measure the same thing. A response can be highly grounded (every sentence traceable to a source) but have low relevance (doesn't answer the question). Conversely, a response can have high relevance (perfectly addresses the question) but low groundedness (the answer was hallucinated). Both failures are distinct and detectable only with their respective metrics.
Exam tip: The most common trap on AB-620 is conflating groundedness and relevance because both relate to 'accuracy.' Remember the mnemonic: Groundedness = G for Ground (source/document), Relevance = R for Response-to-Query. If the scenario mentions 'made-up content' or 'hallucination', think groundedness. If it mentions 'off-topic' or 'didn't answer my question', think relevance.