Quiz 2 Question 11 of 20

A product team runs a multi-agent claims workflow in Microsoft Foundry. Human spot-checks cannot keep up with nightly regression volume. They need an automated evaluation loop that scores agent final answers for faithfulness to retrieved knowledge and for policy compliance, then flags regressions before promotion. Which approach best meets this continuous-improvement need?

Select an answer to reveal the explanation.

Motivation