A rush-light shop intern scores caption–still rows with a spoken-word error count borrowed from the Riva booth. How should pairing quality for those rows be judged?
Select an answer to reveal the explanation.
Short Explanation
A speech scoreboard belongs on the audio shelf, not on the picture–caption shelf. CLIP cares whether the English line matches the still—not how cleanly someone spoke elsewhere.
Full Explanation
Image–text pairing quality is judged on the image and English text arms: does the caption describe the still? Spoken-word error rates are audio-corpus metrics from ASR evaluation and do not measure CLIP alignment. Mixing those scoreboards confuses modality-specific data checks.