Quiz 12 Question 19 of 20

A multi-agent release checklist currently only unit-tests Python helpers. Failures in production are often caused by bad tool schemas, weak retrieval, prompt regressions, or memory mis-scoping—not helper functions. What evaluation design should the team add before promotion?

Select an answer to reveal the explanation.

Motivation