A county elections office pulls one pretrained LLM and wants four jobs: token tags on polling-place names, a label for each precinct note, a short recap of a canvass memo, and a which-site question. How should they score the run?
Select an answer to reveal the explanation.
Short Explanation
Four jobs, four scoreboards: token tags, a precinct-note label, a canvass recap, and a which-site question. Those families do not share one mystery number. BLEU on the tags and an NCA-GENM image board hide task-specific failure.
Full Explanation
The official NLP task families do not share one number. Token classification, text classification, summarization, and question answering each need a matching metric family. A single mystery score, or BLEU forced onto token tags, hides task-specific failure. An NCA-GENM multimodal board is the wrong certification and the wrong scoreboard.