A ship-caulker records a long take that includes the phrases “oakum, then pitch.” One annotator marks when each phrase starts; another writes a single mood note on the whole file. Which annotation helps the speech recognizer learn from the clip?
Select an answer to reveal the explanation.
Short Explanation
Imagine sheet music with bar lines versus a sticky note that just says “energetic.” The recognizer needs words pinned to time, not a floating mood remark.
Full Explanation
Time-aligned transcript annotation ties spoken content to moments in the audio so ASR customization can learn from it. An unanchored mood note, an image caption, or a deploy setting does not provide that speech timing signal.