An administrator is troubleshooting a pilot deployment and wants the most appropriate next step. The team is focused on generative AI evaluation. Which recommendation is most appropriate?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Generative AI evaluation is like quality assurance for a restaurant—you check taste, cleanliness, and consistency, not just portions! Evaluate for groundedness (is it factual?), relevance, harmful content, and policy compliance. Comprehensive checks catch problems before customers do. Evaluate thoroughly, deploy confidently!
Full explanation below image
Full Explanation
The correct answer is C. Comprehensive evaluation of generative AI outputs for groundedness (factual accuracy), relevance to the query, absence of harmful content, and compliance with organizational policies ensures safe and trustworthy deployment. Option A is naive; evaluating only length says nothing about quality or safety. Option B is dangerous; source grounding is critical for enterprise use because ungrounded hallucinations can harm customers and expose the company to liability. Option D is incomplete; while user feedback is valuable, formal governance evaluation by qualified reviewers is essential to catch risks before broad deployment. IBM watsonx.governance requires structured evaluation criteria for generative AI because the risks—hallucination, bias, misuse—are distinct from traditional ML risks.