A paper-mill watermark lab already has a working still-image encoder. Staff want to try three transformer wordings for the accompanying note and leave the encoder untouched. Why is that a controlled multimodal experiment?
Select an answer to reveal the explanation.
Short Explanation
Keep one hand on the photo encoder and wiggle only the wording hand. That is how you learn whether better notes help without rebuilding the eyes of the system.
Full Explanation
A controlled multimodal experiment varies one modality arm at a time. Holding the still-image encoder fixed while trying three transformer wordings isolates the effect of text generation. That design is not a new vision architecture, does not discard the encoder, and does not require regenerating the image set.