A voter-registration desk can scrape two thousand noisy instruction pairs or carefully edit two hundred. A consultant says large language models just want scale. How should they design the supervised-adaptation data experiment?
Select an answer to reveal the explanation.
Short Explanation
Two thousand noisy instruction pairs versus two hundred carefully edited ones. Compare quality and de-duplicated coverage against raw count on a held-out set. Do not assume scale always wins, and do not hand this to an ingest-pipeline design or an NCA-GENM score.
Full Explanation
For supervised adaptation, quality and de-duplicated coverage can beat raw count. The right move is to measure both treatments on a held-out set instead of assuming scale always wins. That measurement is an experiment, not a later-domain ingest-pipeline design. A multimodal exam score does not settle the data-quality question.