A team is preparing a production rollout and needs to avoid a design mistake. Which approach best demonstrates generative AI evaluation in a IBM Certified watsonx Governance Lifecycle Advisor environment?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Generative AI is powerful but tricky—a long answer sounds confident but might be total fiction! Option C gets it right: evaluate for groundedness (is it based on real sources?), relevance (does it actually answer the question?), harmful content (is it safe?), and policy compliance (does it follow your rules?). These checks keep your LLM honest and your users safe!
Full explanation below image
Full Explanation
The correct answer is C. Generative AI evaluation requires multiple dimensions: groundedness (answers backed by trusted sources, not hallucinated), relevance (outputs match user intent and context), harmful content assessment (toxicity, bias, inappropriate material), and policy compliance (aligns with organizational rules like data classification). Option A (length only) is a poor proxy for quality. Option B (ignoring grounding) invites hallucinations and misinformation in enterprise settings. Option D (user-only evaluation) is reactive and doesn't catch problems at deployment time.