Quiz 8 Question 6 of 20

A county courts system uses an LLM-as-a-judge approach, where a separate foundation model scores whether chatbot-generated legal-information responses are accurate and complete, to evaluate output at a scale human reviewers could not sustain. What is the most important limitation the team should account for when relying on this approach?

Select an answer to reveal the explanation.

Motivation