A 311 desk can paste three graded tickets into the prompt or spend a night on a small supervised adaptation. How should they compare those two treatments?
Select an answer to reveal the explanation.
Short Explanation
Three graded tickets in the prompt versus a night of small supervised adaptation. Head-to-head on the same gold tickets, then score task quality. Naming the prompt pattern, skipping the prompt-only arm, or judging by prompt length is not the comparison.
Full Explanation
In-context examples and a small supervised adaptation are two treatments, not two lecture titles. A fair experiment uses the same gold tickets for both arms and reports the task score. Prompt-pattern names do not replace that comparison, and a weight update is not automatically the winner. The desk needs evidence on the shared holdout, not a longer prompt.