A county recycling-pickup tester compares GenAI-generated cases to an expert-written reference set and the specified requirements. Which quality metric is the tester primarily assessing?
Select an answer to reveal the explanation.
Short Explanation
Accuracy is the report-card grade against a trusted answer key—expert cases and the real requirements—not how confident the model sounded.
Full Explanation
Accuracy measures overall correctness versus a reference such as expert-written cases, specified requirements, or other agreed standards (for example, coverage of those requirements). Fluency, token speed, and vendor branding are not this metric.