A parcel-lookup helper has a high automatic QA score, but clerks say it names the wrong lot even when a few words overlap the gold span. What should the experiment add?
Select an answer to reveal the explanation.
Short Explanation
A high automatic QA score can still name the wrong lot when a few words overlap the gold span. Add a small human rubric on a sample: correct lot, and grounded in the passage. Fluency and a higher automatic cutoff do not replace that check.
Full Explanation
Automatic QA metrics can give partial credit when words overlap even if the lot is wrong. When operators disagree with the headline number, a small human rubric is part of the experiment. The rubric should check the correct lot and whether the answer is grounded in the passage. Fluency scores and a higher automatic cutoff do not replace that sample.