A locksmith bench records spoken key-cut requests and types what was said for each clip. What is the correct unit of ASR customization data?
Select an answer to reveal the explanation.
Short Explanation
ASR learning is like flash cards: one spoken clip on the front, the written line on the back. Unlabeled buzz, stills, or a vague mood note leave the recognizer without that paired lesson.
Full Explanation
Riva ASR customization expects aligned audio–transcript pairs as the basic training or adaptation unit. Unlabeled machine noise, image-only folders, and unanchored remarks do not give the model a clear speech-to-text mapping for each example.