A public-library help desk lets one intern star thirty catalog-note summaries and calls the stack human evaluation. What must a real human-eval experiment include?
Select an answer to reveal the explanation.
Short Explanation
One intern starring thirty summaries is a convenience poll. A real human-eval experiment needs a written rubric, a planned sample, and at least a second rater or an agreement check. A thumbs-up button, a Triton hop, or a star-counting kernel is not that design.
Full Explanation
Human evaluation is itself an experiment. A written rubric states what good means, a sample plan states what is rated, and a second rater or an agreement check keeps a single informal like from becoming the official score. One intern starring thirty notes is a convenience poll, not a human-eval design. Serving endpoints or GPU plumbing is not a substitute for that design.