4.0 Configure Evaluation and Monitoring
IBM Certified Specialist - watsonx.governance v2 · 59 questions
- Which task type in watsonx.governance Evaluation Studio is specifically designed to assess retrieval-augmented generation pipelines?
- Which three metric dimensions does watsonx.governance Evaluation Studio provide for LLM evaluation?
- In watsonx.governance LLM-as-a-Judge evaluation, what is the valid score range for metrics such as Faithfulness and Answer Relevance?
- In watsonx.governance LLM-as-a-Judge evaluation, which metric tracks judge LLM calls that fail to return a valid numeric score?
- Which feature in watsonx.governance Evaluation Studio allows a user to designate one prompt or LLM configuration as a fixed baseline during multi-asset comparisons?
- When using multi-asset testing in watsonx.governance Evaluation Studio, what is the primary purpose of assigning weighted factors to evaluation metrics during custom ranking?
- In watsonx.governance, which text quality metric measures n-gram overlap between generated text and reference texts with an emphasis on recall?
- Which text quality metric specifically evaluates text simplification by comparing model output against both a reference text and the original source input?
- In watsonx.governance text evaluation, which metric is defined as the size of the intersection of two token sets divided by the size of their union?
- A data scientist needs a metric that scores 1 when model output is character-for-character identical to the reference and 0 otherwise. Which metric should they configure in watsonx.governance?
- In watsonx.governance, the HAP (Hate, Abuse, and Profanity) index belongs to which evaluation category and monitors which data streams?
- Which statement about PII detection in watsonx.governance is correct?
- Which quality metric is most appropriate when evaluating a binary classification model where false negatives and false positives carry equal business cost?
- In watsonx.governance fairness monitoring, what does the disparate impact ratio measure?
- When configuring a fairness monitor in watsonx.governance, what is the correct definition of the reference group?
- A deployed fraud detection model's input feature distributions have shifted significantly relative to training data, but labeled feedback data shows quality metrics remain within acceptable thresholds. Which type of drift does this scenario describe?
- Which statement best distinguishes LIME from SHAP in the context of watsonx.governance explainability?
- In watsonx.governance explainability, what is a contrastive explanation?
- In the recommended watsonx.governance monitoring configuration sequence, which step immediately precedes setting alert thresholds?
- In the watsonx.governance monitoring configuration workflow, which action must be completed before configuring drift detection?
- A model risk manager wants watsonx.governance to automatically initiate a reassessment workflow whenever a fairness metric remains below 0.80 for three consecutive evaluation windows. Which capability supports this requirement?
- How does watsonx.governance 2.0 differentiate its monitoring approach from earlier-generation model risk management platforms?
- Which watsonx.governance artifact is automatically generated to provide a standardized summary of a model's intended use, training data, performance metrics, and known limitations for stakeholder and regulatory review?
- A monitoring dashboard shows that AUC-ROC for a binary classifier has remained above 0.85 but statistical parity difference has deteriorated to -0.22 over the past 30 days. What does this combination of signals indicate?
- A team is preparing a production rollout and needs to avoid a design mistake. For IBM Certified watsonx Governance Lifecycle Advisor, the topic is fairness evaluation. What should the team do?
- A senior engineer reviews the current plan and notices that one key control is missing. Which approach best demonstrates drift monitoring in a IBM Certified watsonx Governance Lifecycle Advisor environment?
- A consultant is advising a customer that wants a scalable implementation. The team is focused on generative AI evaluation. Which recommendation is most appropriate?
- A project lead is turning a proof of concept into a governed production workflow. What is the best way to handle threshold and alert design while staying aligned with the certification objectives?
- An administrator is troubleshooting a pilot deployment and wants the most appropriate next step. Which approach best demonstrates fairness evaluation in a IBM Certified watsonx Governance Lifecycle Advisor environment?
- During a design review, an architect asks which choice best matches the IBM guidance. The team is focused on drift monitoring. Which recommendation is most appropriate?
- A team is preparing a production rollout and needs to avoid a design mistake. What is the best way to handle generative AI evaluation while staying aligned with the certification objectives?
- A senior engineer reviews the current plan and notices that one key control is missing. For IBM Certified watsonx Governance Lifecycle Advisor, the topic is threshold and alert design. What should the team do?
- A consultant is advising a customer that wants a scalable implementation. The team is focused on fairness evaluation. Which recommendation is most appropriate?
- A project lead is turning a proof of concept into a governed production workflow. What is the best way to handle drift monitoring while staying aligned with the certification objectives?
- An administrator is troubleshooting a pilot deployment and wants the most appropriate next step. For IBM Certified watsonx Governance Lifecycle Advisor, the topic is generative AI evaluation. What should the team do?
- During a design review, an architect asks which choice best matches the IBM guidance. Which approach best demonstrates threshold and alert design in a IBM Certified watsonx Governance Lifecycle Advisor environment?
- A team is preparing a production rollout and needs to avoid a design mistake. What is the best way to handle fairness evaluation while staying aligned with the certification objectives?
- A senior engineer reviews the current plan and notices that one key control is missing. For IBM Certified watsonx Governance Lifecycle Advisor, the topic is drift monitoring. What should the team do?
- A consultant is advising a customer that wants a scalable implementation. Which approach best demonstrates generative AI evaluation in a IBM Certified watsonx Governance Lifecycle Advisor environment?
- A project lead is turning a proof of concept into a governed production workflow. The team is focused on threshold and alert design. Which recommendation is most appropriate?
- An administrator is troubleshooting a pilot deployment and wants the most appropriate next step. For IBM Certified watsonx Governance Lifecycle Advisor, the topic is fairness evaluation. What should the team do?
- During a design review, an architect asks which choice best matches the IBM guidance. Which approach best demonstrates drift monitoring in a IBM Certified watsonx Governance Lifecycle Advisor environment?
- A team is preparing a production rollout and needs to avoid a design mistake. The team is focused on generative AI evaluation. Which recommendation is most appropriate?
- A senior engineer reviews the current plan and notices that one key control is missing. What is the best way to handle threshold and alert design while staying aligned with the certification objectives?
- A consultant is advising a customer that wants a scalable implementation. Which approach best demonstrates fairness evaluation in a IBM Certified watsonx Governance Lifecycle Advisor environment?
- A project lead is turning a proof of concept into a governed production workflow. The team is focused on drift monitoring. Which recommendation is most appropriate?
- An administrator is troubleshooting a pilot deployment and wants the most appropriate next step. What is the best way to handle generative AI evaluation while staying aligned with the certification objectives?
- During a design review, an architect asks which choice best matches the IBM guidance. For IBM Certified watsonx Governance Lifecycle Advisor, the topic is threshold and alert design. What should the team do?
- A team is preparing a production rollout and needs to avoid a design mistake. The team is focused on fairness evaluation. Which recommendation is most appropriate?
- A senior engineer reviews the current plan and notices that one key control is missing. What is the best way to handle drift monitoring while staying aligned with the certification objectives?
- A consultant is advising a customer that wants a scalable implementation. For IBM Certified watsonx Governance Lifecycle Advisor, the topic is generative AI evaluation. What should the team do?
- A project lead is turning a proof of concept into a governed production workflow. Which approach best demonstrates threshold and alert design in a IBM Certified watsonx Governance Lifecycle Advisor environment?
- An administrator is troubleshooting a pilot deployment and wants the most appropriate next step. What is the best way to handle fairness evaluation while staying aligned with the certification objectives?
- During a design review, an architect asks which choice best matches the IBM guidance. For IBM Certified watsonx Governance Lifecycle Advisor, the topic is drift monitoring. What should the team do?
- A team is preparing a production rollout and needs to avoid a design mistake. Which approach best demonstrates generative AI evaluation in a IBM Certified watsonx Governance Lifecycle Advisor environment?
- A senior engineer reviews the current plan and notices that one key control is missing. The team is focused on threshold and alert design. Which recommendation is most appropriate?
- A consultant is advising a customer that wants a scalable implementation. For IBM Certified watsonx Governance Lifecycle Advisor, the topic is fairness evaluation. What should the team do?
- A project lead is turning a proof of concept into a governed production workflow. Which approach best demonstrates drift monitoring in a IBM Certified watsonx Governance Lifecycle Advisor environment?
- An administrator is troubleshooting a pilot deployment and wants the most appropriate next step. The team is focused on generative AI evaluation. Which recommendation is most appropriate?