A county night-sky preserve has thousands of nightly answers and cannot hire a person for each one. Another model scores helpful, grounded, or rude against a rubric. Which evaluation pattern is that?
Select an answer to reveal the explanation.
Short Explanation
Think of thousands of nightly answers and no person for each one. Another model scores helpful, grounded, or rude against a rubric. That is LLM-as-a-judge, not a replacement for every high-risk human review.
Full Explanation
LLM-as-a-judge is the conceptual use of a model to score another model’s outputs. It does not replace every human review on high-risk content. ROUGE and from-scratch pre-training are different ideas.