The City of Maplewood's digital services team has deployed a Copilot Studio agent to answer citizen questions about permit applications, zoning regulations, and municipal services. The agent uses a curated knowledge base of official city documents. Because incorrect information could mislead residents and create legal liability, the agency director requires a formal approval process before the agent goes live. The evaluation lead, Simone, presents evaluation results showing the agent's average groundedness score. The director asks: 'What groundedness score indicates that the agent's responses are adequately supported by the official documents?' What is the generally accepted minimum groundedness score threshold that indicates acceptable agent quality in Copilot Studio AI evaluation?
Select an answer to reveal the explanation.
Short Explanation and Infographic
On the Copilot Studio 1-5 groundedness scale, a score of 3 is generally the floor for acceptable quality — think of it like a passing grade of 60% on an exam. Scores below 3 mean too many responses are drifting from the source material. Scores of 4 or 5 are ideal but may not be achievable for all knowledge base configurations. Correct answer: C.
Full explanation below image
Full Explanation
Copilot Studio's AI evaluation framework scores groundedness (and other metrics like relevance and coherence) on a scale of 1 to 5, where 1 represents the worst quality and 5 represents the best. Microsoft's evaluation framework guidance establishes that a score of 3 or higher generally indicates acceptable agent quality — responses are mostly grounded in the knowledge source with acceptable levels of unsupported content.
Option C is correct. A groundedness score of 3 or higher is the commonly referenced threshold for acceptable quality in Copilot Studio evaluations. This means that on average, most of the agent's responses are substantially supported by the knowledge source content. For high-stakes deployments like government services, teams often target 4 or higher as a goal, but 3 is the minimum acceptable baseline before production deployment.
Option A is incorrect because a groundedness score of 1 represents the lowest possible quality — nearly all responses are unsupported by source content, indicating severe hallucination. This would immediately disqualify an agent from production deployment, particularly in a government context where legal liability is cited.
Option B is incorrect because a groundedness score of 2 indicates below-average source support — the majority of responses have inadequate grounding. This is below the acceptable threshold and would indicate significant hallucination risk. A score of 2 should trigger immediate remediation of the knowledge base or agent configuration.
Option D is incorrect because requiring a perfect score of 5 for any production deployment is an unrealistically high bar. Perfect groundedness is difficult to achieve across all query types, especially with complex or ambiguous questions. Setting 5 as a mandatory minimum would prevent many capable and safe agents from deploying. The realistic target for high-quality agents is 4 or above, with 3 as the acceptable floor.
Exam tip: Memorize the scale: 1-2 = unacceptable (agent needs remediation), 3 = minimum acceptable threshold, 4-5 = good to excellent quality. On the AB-620 exam, threshold questions are often structured as 'what score indicates acceptable quality?' — the answer is almost always 3 or higher on the 1-5 scale.