A night-market voice stall files complaints that mix speech-to-text with a stall photo. Staff can try a pretrained classifier with no local labels, paste six examples into the instruction, or attach a small adapter on the text side. How should those options be treated?
Select an answer to reveal the explanation.
Short Explanation
Zero-shot, six sticky examples, or a tiny add-on module are three different dials on the same routing job. Try them as experiments — not as a product bake-off that ignores the stall photo pairing.
Full Explanation
Prompt-only classification, a few in-context labels, and a small text adapter are graded levers on the language arm of a speech-to-text plus photo routing task. Framing them as experiments supports comparison without forcing from-scratch vision training or conflating them with image regeneration.