Before promoting a multi-agent workflow in Microsoft Foundry, product owners must score groundedness and task success on a labeled set. Automated metrics alone are not trusted yet. What should you design?
Select an answer to reveal the explanation.
Short Explanation
B is correct. Evaluation skills start with human review processes in Foundry to judge multi-agent solutions on real cases—groundedness, task success, safety—before promotion. A compile-only gates miss agent quality. C UI screenshots ignore agent behavior. D self-confidence is not a substitute for labeled human evaluation. Pair human review with later automated judges, but do not skip the human process when the stem says owners must score quality.
Full Explanation
Correct Answer — B
Designing and implementing a human review process to evaluate solutions in Foundry is a core evaluation skill. Curated cases, human scores, and promotion gates establish quality baselines for multi-agent systems.
Why A is wrong: Compilation does not measure agent task success or groundedness.
Why C is wrong: UI cosmetics are not multi-agent quality evaluation.
Why D is wrong: Model self-confidence is unreliable as a sole gate.
Exam tip: Foundry human review is an explicit AI-500 evaluation lever.