A clerk's office loves its note-tagging checkpoint and wants to reuse it unchanged to write inspection summaries. What should the experiment expect?
Select an answer to reveal the explanation.
Short Explanation
A note-tagging checkpoint is not automatically a summarization winner. The new task needs its own references and score. Reusing tag F1, shipping after any paragraph appears, or calling a task-swap always safe hides the real outcome.
Full Explanation
A checkpoint tuned for one NLP task is an experimental starting point, not a trophy on a different task. Tagging quality does not answer whether inspection summaries are useful or faithful. A fair summarization eval needs its own references or a human rubric, plus a score that matches generation. Reusing tag F1, or shipping after any paragraph appears, hides the real outcome.