A glasshouse desk recognizes English questions correctly, but the spoken reply comes out in a language visitors do not speak. What kind of experiment does that call for?
Select an answer to reveal the explanation.
Short Explanation
The ear hears English; the mouth answers in another tongue—that is a pipeline dress-code mismatch, not a new picture model. Align ASR and TTS language settings in a focused cross-stage trial.
Full Explanation
Conversational pipelines can fail when ASR and TTS language packs disagree even if recognition accuracy looks fine. The correct experiment checks and aligns language configuration across stages. Vision models, fabric tuning, and unrelated CLIP work do not address speech-language mismatch.