A regulated insurer must approve any change to a multi-agent underwriting assistant before production. Automated Foundry evaluations score the suite green, but risk stakeholders require evidence that humans reviewed edge cases the metrics might miss. Which process design best satisfies both speed and governance?
Select an answer to reveal the explanation.
Short Explanation
The correct answer is A. Regulated underwriting needs both machine scale and human judgment on the scary edge cases. Sample fails and high-risk items, capture reviewer decisions with the evaluation records, and require automated thresholds plus human sign-off before promote. A 0.99 auto score alone will miss rare harm modes. Laptop cowboy deploys skip governance. Waiting for production incidents is reverse quality control.
Full Explanation
Option A is correct. AI-500 skills call for designing human review processes to evaluate solutions in Foundry alongside automated evaluations. Sampling high-risk and failing cases, recording human judgments, and dual-gating release balances velocity with governance for regulated domains.
Option B is incorrect because high composite scores do not guarantee coverage of rare, high-impact edge cases that regulators and risk teams care about.
Option C is incorrect because uncontrolled local promotion bypasses evaluation evidence, audit trails, and separation of duties expected in regulated underwriting systems.
Option D is incorrect because post-incident customer feedback is lagging and externalizes risk; governance requires pre-production review for underwriting assistants.