A tool-library FAQ helper is scored on the share of answers a reviewer marks factually correct. Leadership treats that share as the success target. They are not asking for a homework confusion matrix. Which metric is this?
Select an answer to reveal the explanation.
Short Explanation
Think of a reviewer putting a check mark on answers that are actually true. That share is business-facing accuracy. A homework confusion matrix, BLEU on folder colors, and GPU heat do not replace the reviewer's mark.
Full Explanation
Human-judged factual correctness is a business-facing accuracy metric for a GenAI helper. A confusion-matrix exercise without a reviewer is not what leadership asked for. BLEU and GPU temperature do not measure whether a reviewer marked the answer correct.