A water desk scores extractive 'where is valve 9?' answers with a translation-style n-gram overlap number and crowns a chatty restatement. What is wrong with that primary metric?
Select an answer to reveal the explanation.
Short Explanation
A translation-style n-gram can crown a chatty restatement of the wrong valve. For extractive span QA, prefer span match or token F1, plus a human check. Perplexity and the longest aisle description are not location scores.
Full Explanation
Extractive span QA is scored by whether the predicted span matches the gold location, not by how much the prose overlaps a reference. Translation-style n-gram scores such as BLEU or ROUGE can reward a chatty restatement and hide a wrong valve. Span exact match or token F1, with a human check on a sample, is the associate primary. Perplexity and length are not location scores.