A city clerk must recast each parking-appeal ruling into a statute-mandated one-sentence formula that already has a gold restatement. Which automatic metric is a reasonable primary score for that constrained rewrite, yet a weak primary for open-ended chat?
Select an answer to reveal the explanation.
Short Explanation
The parking-appeal ruling already has a gold one-sentence restatement. A BLEU-style n-gram overlap is a reasonable primary for that constrained rewrite. It is a weak primary for open-ended chat, where many replies are fine and few references stay stable. Perplexity, a mystery leaderboard, and NCCL bandwidth are not that scoring choice.
Full Explanation
BLEU measures n-gram overlap with one or more gold references, so it fits a constrained restatement that already has a target phrasing. Open-ended chat has many acceptable replies and few stable references, so BLEU is a weak primary there. Associate evaluation matches the metric family to the task instead of reusing one number for every generation job. Raw perplexity, a mystery leaderboard, or a Professional collective measurement does not answer this scoring choice.