In watsonx.governance LLM-as-a-Judge evaluation, what is the valid score range for metrics such as Faithfulness and Answer Relevance?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Think of it this way: in real-world AI governance, 0 to 1 is exactly what teams reach for when they need to handle this scenario. LLM-as-a-Judge metrics in watsonx. On the exam, remember that this falls squarely under the 4.0 Configure Evaluation and Monitoring domain.
Full explanation below image
Full Explanation
LLM-as-a-Judge metrics in watsonx.governance use a 0-to-1 scale. A score of 1 represents perfect performance and 0 indicates complete failure for that metric. The correct answer, "0 to 1", directly addresses the scenario described because it aligns with the specific governance requirement in question. The incorrect options ("0 to 100", "-1 to 1", "1 to 5") may seem plausible but do not satisfy the core requirement. Understanding the distinction between these concepts is critical for IBM watsonx.governance implementations and is frequently tested in the 4.0 Configure Evaluation and Monitoring section of the certification exam.