Quiz 7 Question 19 of 20

A 311 chatbot team measures response quality by comparing generated answers against reference answers using BLEU and ROUGE scores. A reviewer notices these scores penalize a response that used different wording but conveyed the exact same information as the reference answer. What does this reveal about BLEU and ROUGE as evaluation metrics?

Select an answer to reveal the explanation.

Motivation