A harbor museum claims “clearer context helps everything” and runs a tighter sentence on the postcard generator and a tighter reply template on the Riva booth. How should metrics be logged?
Select an answer to reveal the explanation.
Short Explanation
Postcards and help-desk speech are different senses. Grade the card on image–text fit and the booth on transcription or listener clarity—do not mash them into one fake score.
Full Explanation
Parallel multimodal experiments need a metric per modality. Image–text alignment and speech transcription or listener checks match the postcard and Riva trials; a single blended number or infrastructure proxy hides modality-specific outcomes.