Before promoting a new multi-agent configuration in Microsoft Foundry, the responsible AI lead requires structured human review of sampled traces: correctness, tool appropriateness, and policy adherence—not only automated scores. Which evaluation practice should the team implement?
Select an answer to reveal the explanation.
Short Explanation
Foundry-oriented evaluation expects a real human review process—sampled runs, rubrics, recorded decisions—especially for policy-sensitive multi-agent systems. Automated scores alone miss nuanced tool misuse. Unstructured intern skims and model self-praise are not review programs. B is the practice to implement.
Full Explanation
Correct answer: B. Evaluation strategies for multi-agent solutions include designing and implementing a human review process to evaluate solutions in Foundry. Structured rubrics and recorded outcomes complement automated judges and catch issues scores miss.
A is incorrect because automated n-gram or similar scores are insufficient alone for tool-using multi-agent policy compliance.
C is incorrect because informal, unrecorded review lacks rigor and auditability.
D is incorrect because the model’s own confidence language is not an independent human evaluation.
Exam focus: human review in Foundry is an explicit evaluation skill, not optional theater.