A municipal water desk's automatic scores say a new summarizer is better; night operators say it drops boil-water alerts. What should the experiment do?
Select an answer to reveal the explanation.
Short Explanation
Automatic scores say the summarizer is better. Night operators say boil-water alerts vanished. Treat that as a missing failure mode and add checklist items or critical-span recall. Discarding the complaints, standing up NIM, or moving to a rack exam does not repair the scoreboard.
Full Explanation
If automatic metrics miss a critical failure mode, the experiment's scoreboard is incomplete. Operator reports of dropped boil-water alerts are a reason to add checklist items or critical-span recall, not a reason to ignore the users. Serving every draft or moving the work to a datacenter exam does not repair the missing metric. Associate evaluation follows the failure mode, then rescores.