A community wind-ensemble helper scores high on a public summarization benchmark but fails on program-note house style. What should the team add?
Select an answer to reveal the explanation.
Short Explanation
Think of a helper that looks great on a public summarization quiz but still writes the wrong program-note voice. Add a task-specific set. A generic high score, no eval, or a content filter does not prove house style.
Full Explanation
A public benchmark does not automatically prove the business task works. A task-specific set is needed for house style. Trusting the generic score, dropping eval, or using a filter does not measure that style.