An aquarium dive-bell office tags species names in a transcript, then a speech synthesizer reads the tags aloud to a gallery. Visitors hear the wrong Latin name even though the photo wall is correct. Where should the next single-stage experiment look first?
Select an answer to reveal the explanation.
Short Explanation
Wrong words out of the speaker can mean bad tags or a clumsy reader. Fix one stage at a time — highlighter versus voice — so you know which knob moved the error.
Full Explanation
A spoken error after token tagging can sit in the text-arm tags or in the later synthesizer. The next experiment should change only one stage so the failure can be localized. Blaming the correct photo wall, rewriting every speech stage at once, or shifting to facility cooling does not isolate the multimodal speech pipeline fault.