A flute-maker intern wires Riva ASR directly to “generate the picture” with no English text prompt. What is the correct Domain 4 correction?
Select an answer to reveal the explanation.
Short Explanation
Speech goes to the talk booth; stills need written English. CLIP paints from a text prompt—typed or already transcribed—not from a raw speech service hook.
Full Explanation
The CLIP image path conditions on English text. Riva ASR belongs on the conversational speech path; raw speech should not be wired as the still generator. Avatar tools and Helm scaling do not fix a modality mismatch on the generate call.