Your team must evaluate a multi-agent claims solution in Microsoft Foundry before wider release. Stakeholders want structured human review of agent outcomes, not only automated scores. Which approach best designs the human review process?
Select an answer to reveal the explanation.
Short Explanation
Human review in Foundry-style evaluation is structured: representative cases, rubrics for correctness/groundedness/policy, multiple reviewers, and defects that become backlog work. Option D. App-store-after-ship (A) is not a release gate. Informal skims (B) miss systematic failure modes. CPU charts (C) do not measure answer quality.
Full Explanation
Evaluate-domain skills include designing and implementing human review processes to evaluate solutions in Foundry. Option D provides rigor, coverage, and a feedback loop into engineering.
Option A postpones quality control until customer harm is possible.
Option B lacks inter-rater consistency and auditability.
Option C measures infrastructure, not multi-agent decision quality.
Combine human review with automated evaluations for memory, tools, and prompts where possible.
Exam tip: Human review needs rubrics and sampling strategy, not vibes.