A town council asks why a “large” language model needs huge text, many weights, and GPU time. What associate picture should they get?
Select an answer to reveal the explanation.
Short Explanation
“Large” here is a three-part picture: more weights, more diverse text, and enough GPU time. Together they buy broader language skill. One labeled row does not replace that pile, cooling capacity is not the definition, and nobody needs a scaling-law formula on this exam.
Full Explanation
Associate scale intuition is qualitative: data, parameter count, and compute together produce broader language skill. It is not “one row is enough,” not a cooling exam, and not a Professional cluster design.