A wind-farm blade-photo desk has inspector notes and stills. Leadership wants “something with a pretrained language model.” If the real need is to pull the matching still for a typed fault question, which text-arm trial should they run?
Select an answer to reveal the explanation.
Short Explanation
If the job is “ask a question, get the right blade photo,” that is question-answering with a still at the end — not a chatty box, a vague summary, or a sorter nobody asked for.
Full Explanation
Among official NLP task types as the text arm of a photo+notes experiment, QA is the match when a typed question must surface the right still. Classification bins or summarization serve other decisions, and a generic chatbot that ignores the photo set fails the multimodal requirement.