A historic brickworks archive has thousands of unlabeled kiln diaries and a small labeled set of caption this photo pairs. Which application choice matches each data type?
Select an answer to reveal the explanation.
Short Explanation
Think of thousands of unlabeled kiln diaries versus a small stack of caption-this-photo pairs. Continued pre-training on the diaries, fine-tuning on the pairs. Do not swap those, and do not start from random weights.
Full Explanation
Continued pre-training fits unlabeled domain text. Fine-tuning fits labeled task pairs. Swapping those jobs, treating them as a dashboard, or starting from random weights is the wrong design decision.