Quiz 6 Question 18 of 20

After running a batch evaluation of a Copilot Studio agent's generative AI responses against a golden dataset of 100 expected Q&A pairs, the evaluation report shows: Precision = 0.91, Recall = 0.74, F1 = 0.82. The agent frequently fails to retrieve relevant answers even though when it does answer, it is usually correct. Which metric most directly reflects this 'misses answers' problem?

Select an answer to reveal the explanation.

Motivation