A team is preparing a production rollout and needs to avoid a design mistake. The team is focused on generative AI evaluation. Which recommendation is most appropriate?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Evaluating a generative AI answer only by length is like grading an essay by word count — you reward verbosity while missing whether anything written is actually true. Option C is the complete rubric: evaluate for groundedness, relevance, harmful content, and policy compliance where applicable. Length is a proxy, not a quality signal. Ignoring source grounding in enterprise answers invites confident hallucination. Delegating all safety evaluation to end users outsources quality control to people who lack the context. Structured, multi-dimension GenAI evaluation — option C!
Full explanation below image
Full Explanation
The correct answer is C. Evaluate generative AI for groundedness, relevance, harmful content, and policy compliance where applicable. aligns directly with the 4.0 Configure Evaluation and Monitoring exam objective and reflects how IBM watsonx.governance operationalizes this practice in enterprise environments. IBM watsonx.governance is designed to provide the tooling, workflows, and evidence collection mechanisms that make AI trustworthy and auditable from intake through retirement. Choosing this option ensures the implementation is both defensible to regulators and maintainable by governance teams across the full AI lifecycle. Option A is incorrect because it narrows the scope to a single artifact or metric and misses the full governance requirement. Option B is incorrect because it eliminates a required control and leaves the governance program without a critical safeguard. Option D is incorrect because it offloads or obscures a governance responsibility in a way that breaks accountability and auditability. On scenario-based exam questions, the correct answer is consistently the one that operationalizes the IBM governance practice with structured controls rather than approximating it with shortcuts, manual workarounds, or oversimplified rules.