test-and-manage-agents
Microsoft Certified: AI Agent Builder Associate · 71 questions
- Priya is an AI developer at Contoso Financial Services who has built a Copilot Studio agent that answers questions about investment products. Her manager asks her to set up a formal evaluation framework using a test set before the agent goes live. Priya needs to understand what a properly structured test set must contain to enable accurate AI evaluation scoring in Copilot Studio. Which combination of elements is required in each test set entry for evaluation to function correctly?
- Northshore Medical Group has deployed a Copilot Studio agent that helps clinicians look up treatment protocols from a curated SharePoint knowledge base. The compliance team is concerned that the agent might be making up information not present in the approved documents — a risk they call 'hallucination.' The AI team lead, Marcus, wants to choose the most appropriate evaluation method to surface this specific risk. Which evaluation method should Marcus prioritize to detect whether the agent's responses are fabricating information not found in the source documents?
- Apex Capital Management has deployed a Copilot Studio agent to assist wealth advisors with client portfolio questions. During evaluation, the team notices that the agent sometimes responds with technically accurate information about investment concepts — but the information does not actually address what the advisor asked. The agent answers a question about bond duration by explaining equity volatility. The evaluation team needs to select a metric that will flag this specific type of failure. Which evaluation metric directly measures whether the agent's response addresses the user's actual question?
- Stellartech Retail has deployed a Copilot Studio agent to handle return and exchange policy questions for store associates. During evaluation review, the team notices that some agent responses are on-topic and drawn from the correct policy documents — but the explanations jump between unrelated points, contradict themselves mid-paragraph, and leave associates more confused than before. An evaluator named Delia wants to flag this specific quality failure using the correct evaluation metric. Which evaluation metric is Delia looking for?
- Orion Commerce has built a Copilot Studio agent to handle customer inquiries for their e-commerce platform covering shipping, returns, product availability, and account issues. The AI team lead, Fatima, is designing the initial test set and wants to ensure it is representative enough to catch real-world agent failures before go-live. She asks a consultant to review her draft test set of 8 questions, all of which ask about shipping timelines in nearly identical phrasing. What should the consultant advise Fatima about test set design for a representative evaluation?
- Lakeview Insurance has built a Copilot Studio agent that handles claims status inquiries for policyholders. The agent team has prepared a test set of 35 question-answer-context triples. Now the team lead, Rodrigo, needs to run the evaluation and review the groundedness, relevance, and coherence scores to decide whether the agent is ready for production. In Copilot Studio, where should Rodrigo navigate to upload the test set, trigger the evaluation run, and view the resulting metric scores?
- Veridia Telecom has deployed a Copilot Studio agent to handle billing, technical support, and account management inquiries for residential customers. After two weeks in production, the operations team suspects that certain conversation topics are failing at a higher rate than others — specifically, customers seem to be escalating to human agents unusually often after asking about bill disputes. The product owner, Keiko, wants to confirm this hypothesis and identify which specific topics are performing poorly. Where in Copilot Studio should Keiko look to find topic-level performance data, escalation rates, and resolution metrics?
- LexAI Solutions is building a Copilot Studio agent for a law firm to help paralegals look up case law and statutory references. The agent quality lead, Tomas, is designing the test case library. He has written 40 test cases that all follow the same pattern: a well-formed legal query in formal language, matched against a clear, single-source answer. His colleague points out that real paralegals will also ask vague questions, use informal phrasing, mix multiple legal topics in one question, and occasionally try to get the agent to give legal advice it is not authorized to provide. What type of test cases is Tomas's library currently missing?
- Pinnacle Manufacturing has deployed a Copilot Studio agent to assist maintenance engineers with equipment troubleshooting. The agent uses a curated SharePoint library of OEM service manuals as its knowledge source. During evaluation, the QA engineer Lena observes two different failure modes: in some cases, the agent responds with accurate answers that don't actually address the engineer's question; in other cases, the agent's response addresses the question but includes claims that cannot be traced to any manual in the library. Lena needs to choose separate evaluation metrics to detect each failure mode independently. Which statement correctly identifies the key distinction between groundedness and relevance?
- The City of Maplewood's digital services team has deployed a Copilot Studio agent to answer citizen questions about permit applications, zoning regulations, and municipal services. The agent uses a curated knowledge base of official city documents. Because incorrect information could mislead residents and create legal liability, the agency director requires a formal approval process before the agent goes live. The evaluation lead, Simone, presents evaluation results showing the agent's average groundedness score. The director asks: 'What groundedness score indicates that the agent's responses are adequately supported by the official documents?' What is the generally accepted minimum groundedness score threshold that indicates acceptable agent quality in Copilot Studio AI evaluation?
- CloudSync Solutions operates a Copilot Studio agent for internal IT helpdesk support. The agent's knowledge base is a SharePoint site that is updated weekly with new runbooks, policy changes, and software release notes. After the initial evaluation passed with strong groundedness and relevance scores, the IT director declared the agent production-ready. Three months later, users report that the agent gives outdated or contradictory answers — likely because the knowledge base content has drifted significantly since the initial evaluation. The agent team lead, Hira, is tasked with preventing this from recurring. What approach should Hira implement to maintain ongoing evaluation quality as the knowledge base evolves?
- A Copilot Studio administrator at Contoso wants to review how users have been interacting with an agent over the past 30 days — specifically how many conversations were handled, which topics were triggered most, and how many conversations ended without resolving the user's issue. Where should the administrator look for this data?
- A Copilot Studio agent at Fabrikam is intermittently failing when it calls an external HR API. The developer needs to find which specific API calls are failing, what error codes are returned, and at what time the failures occur. Which monitoring tool provides this level of detail?
- The operations team at Lucerne Publishing wants to receive an automated alert when their Copilot Studio agent's escalation rate exceeds 30% in any given hour. How should this alert be configured?
- After running a batch evaluation of a Copilot Studio agent's generative AI responses against a golden dataset of 100 expected Q&A pairs, the evaluation report shows: Precision = 0.91, Recall = 0.74, F1 = 0.82. The agent frequently fails to retrieve relevant answers even though when it does answer, it is usually correct. Which metric most directly reflects this 'misses answers' problem?
- A Copilot Studio agent at Bellows College has been in production for 60 days. The analytics show an average CSAT score of 2.8 out of 5. The team wants to identify which specific topics are driving low satisfaction scores. What should they do?
- The operations team at Tailspin Toys notices that 40% of their Copilot Studio agent sessions end without the user's issue being resolved. The team wants to understand whether users are abandoning conversations early or whether the agent is failing to complete topics. Which metric helps distinguish between user abandonment and agent failure?
- A quality assurance engineer at Datum Corporation is reviewing the batch evaluation results for a Copilot Studio agent's generative AI responses. The evaluation report includes metrics: Coherence = 4.2/5, Groundedness = 2.8/5, Relevance = 3.9/5. Which metric indicates that the agent is generating responses that are not sufficiently supported by the knowledge source content?
- A customer at Woodgrove Bank reported that the bank's Copilot Studio agent gave incorrect information during their conversation yesterday afternoon. The support team wants to review the exact conversation flow, the variables that were set, and which topics were triggered during that specific session. Where can they access this information?
- The Copilot Studio development team at Contoso has made significant changes to their agent's generative AI configuration and knowledge sources. Before promoting the new version to production, they want to objectively compare the response quality of the new version against the current production version. What is the recommended approach?
- A Copilot Studio agent at Pacific Airlines is connected to Azure Application Insights. The operations team has created an Azure Monitor alert rule that fires when the API error rate exceeds 5% in any 15-minute window. The alert fires, but no one receives a notification. What is most likely misconfigured?
- The Copilot Studio Analytics for Litware Inc.'s agent shows that the topic 'Request Office Supply' has a trigger rate of 0.3% over 30 days, while 'IT Password Reset' has a trigger rate of 42%. The product team wants to use this data to optimize the agent. What is the most appropriate action to take based on these trigger rates?
- Your organization wants to implement a formal CI/CD process for Copilot Studio agents, promoting them through Development → Test → Production environments with controlled approvals. Which Microsoft tool is designed specifically for this Power Platform deployment automation?
- Your organization requires that all Copilot Studio agent deployments to the Production environment receive explicit approval from a designated stakeholder before proceeding. How do you configure this in Power Platform Pipelines?
- A Copilot Studio maker is starting a new agent project that will eventually be deployed to production through formal pipelines. What is the recommended first step from an ALM perspective before building any agent topics or knowledge sources?
- An administrator is about to import a Copilot Studio agent solution into the production environment. They have both a managed and an unmanaged version of the solution. Which version should be deployed to production and why?
- After deploying a Copilot Studio agent solution via Power Platform Pipelines to the test environment, the pipeline run completes with a warning about connection references. What post-deployment action is required to make the agent fully functional in test?
- After deploying a Copilot Studio agent to production, a product manager asks how to determine whether the agent is successfully resolving user inquiries or frequently escalating to human support. Which built-in feature provides this operational intelligence?
- During development, a maker wants to test a specific topic in isolation to verify that its logic branches and variable assignments work correctly, without going through the normal trigger phrase recognition. How can they test the topic directly in Copilot Studio?
- Before deploying a Copilot Studio agent solution from development to test via Power Platform Pipelines, a team lead wants to validate the solution for common issues like missing dependencies, deprecated components, and performance anti-patterns. Which tool performs this pre-deployment validation?
- A Power Platform Pipelines deployment of an updated Copilot Studio agent to production causes unexpected agent behavior. The previous version was working correctly. What is the recommended recovery approach using Power Platform's built-in capabilities?
- During a test conversation in the Copilot Studio Test console, a maker notices that a condition node is taking the wrong branch. They suspect a variable has an incorrect value. What Test console feature helps them diagnose this without adding extra 'send message' nodes to display variable values?
- Your organization wants to automate Power Platform Pipelines deployments using a service account that is not tied to any individual employee's identity, ensuring deployments continue even when team members leave. Which identity type should be used for the pipeline service account?
- A support lead reviews a Copilot Studio agent that handles password reset guidance. They need a high-level view of how many sessions were resolved by the agent versus escalated or abandoned last week. Which analytics area is the best first place to review session outcomes?
- Litware uses Dev, Test, and Production Power Platform environments. Their Copilot Studio agent and its dependent flows live in a solution. They want makers to promote the solution with approvals and minimal custom scripting. Which approach best fits?
- Your team packages a Copilot Studio agent that calls a SharePoint knowledge site URL and a SQL API base URL that differ per environment. You also use a Power Automate flow with a SharePoint connection. Which TWO solution practices should you apply so Dev → Test → Prod promotion works cleanly? (Choose two.)
- Before promoting a solution containing a Copilot Studio agent to Test, the ALM lead wants static analysis for common Power Platform issues. Which tool should they run?
- Your organization finalizes a Copilot Studio agent for Production. What is the recommended solution type to import into Production?
- Leadership wants customer satisfaction scores for a Copilot Studio customer service agent. CSAT tiles in analytics are empty. What must be true for CSAT data to appear?
- Ops wants deeper diagnostics of custom events when users complete a multi-step enrollment topic, beyond the standard Copilot Studio analytics tiles. Which approach fits?
- After a pipeline deploys a solution to Test, a flow action used by the agent fails with a connection error even though it worked in Dev. What is the most likely ALM misconfiguration?
- Analytics show one topic with high trigger rate but low resolution and high escalation. Which TWO actions are the best next steps? (Choose two.)
- An environment variable ApiBaseUrl is in the agent solution with Dev default https://api-dev.contoso.com. How should Test use https://api-test.contoso.com without editing topic nodes each deploy?
- While testing, a Condition node never takes the expected branch even though the user answered the question. What Test pane practice helps diagnose fastest?
- Enterprise ALM requires that Production deployments of the agent solution be approved by the ops lead and executed without using a personal maker account password in automation. Which TWO elements support this? (Choose two.)
- Before releasing a knowledge-heavy agent, the QA lead wants evidence answers stick to sources rather than inventing claims. Which evaluation focus is most relevant?
- Weekly analytics show escalation rate jumped from 8% to 28% after a knowledge source update. What is the best immediate operational response?
- A Contoso customer support agent is generating responses that sometimes include product pricing not found in the company knowledge base. The QA lead wants to implement a systematic metric to detect when the agent fabricates information not supported by its retrieved source documents. Which evaluation metric should the QA lead configure in Azure AI Foundry to address this concern?
- An AI agent built for an HR department is flagged because it frequently answers questions about company benefits by providing accurate historical data from its knowledge base, but users report the answers do not address what they actually asked. For example, a user asking about dental coverage receives a detailed response about vision benefits instead. Which Azure AI Foundry evaluation metric is most appropriate for diagnosing this agent quality problem?
- Fabrikam's development team has built a Copilot Studio agent and needs to move it through Development, Test, and Production environments in a controlled, repeatable way. The team wants each stage transition to require an explicit approval before the solution is deployed to the next environment. Which Power Platform feature should the team configure to meet this requirement?
- Northwind Traders has configured a Power Platform Pipeline with three stages: Dev, UAT, and Production. The pipeline administrator wants to ensure that the Head of IT must personally review and approve any deployment to the Production stage before the solution is released. The Head of IT should receive an email notification when a deployment is waiting. How should the pipeline administrator configure this requirement?
- The operations manager at AdventureWorks wants to review the exact messages exchanged between customers and their Copilot Studio booking agent over the past week, including which topics were triggered, how the conversations ended, and whether any conversations were escalated to a human. Where in Copilot Studio should the manager navigate to access this information?
- A business analyst at Contoso is preparing a monthly report on their Copilot Studio IT helpdesk agent's effectiveness. The analyst needs to report on total conversations, the percentage of conversations the agent resolved without human intervention, and the percentage that ended without any resolution. Which Copilot Studio Analytics view provides these three KPIs in a single dashboard?
- A Copilot Studio maker is debugging a topic called 'Order Status' that is supposed to call a Power Automate flow and return a tracking number. During testing, the agent gives a generic 'I could not retrieve that information' message instead of the tracking number. The maker wants to see the exact path the conversation took through the topic nodes, which branch conditions were evaluated, and what value was stored in the tracking number variable. What should the maker do in the Test pane to diagnose this issue?
- A maker at Tailspin Toys has finished building a Copilot Studio agent in the Development environment. The agent includes a custom topic, a connection to a SharePoint knowledge source, a Power Automate flow, and an environment variable for the SharePoint site URL. The team needs to move all of these components to the Test environment as a single deployable unit. What should the maker do before running the Power Platform Pipeline?
- Alpine Ski House is establishing governance for their Copilot Studio agent development lifecycle. The organization wants to ensure that: (1) makers can freely build and modify the agent in one environment, (2) testers can validate agent behavior without risk of changes from makers, and (3) end users interact only with a stable, approved version. Which environment architecture satisfies all three requirements?
- A customer experience manager at Woodgrove Bank reviews their Copilot Studio loan inquiry agent monthly. Last month's data shows the agent had a resolution rate of 68%, but user satisfaction scores are significantly lower than expected. The manager suspects users are being routed to human agents too often for questions the agent should handle. Which two Copilot Studio Analytics metrics should the manager examine together to confirm this hypothesis?
- A quality engineering team at Contoso needs to validate their Copilot Studio agent against 200 predefined test utterances every time a new version is deployed to the Test environment. Running each utterance manually through the Test pane is too time-consuming. The team wants an automated, repeatable testing approach that can flag regressions when topic behavior changes between deployments. Which approach best meets this requirement?
- Britta is an ALM administrator promoting a Copilot Studio agent from the development environment to production. The team uses managed solutions for production deployments. After importing the managed solution, a developer in production reports they cannot edit the agent's system topic 'Greeting' directly in the production environment. Is this the expected behavior, and why?
- Contoso's Power Platform pipeline is configured with three stages: Development → Test → Production. The platform administrator wants to require explicit human approval before any solution can be deployed to the Production stage. Which pipeline feature should the administrator configure to enforce this requirement?
- A Copilot Studio agent handles multi-step insurance claim submissions. Your monitoring team reports that 40% of sessions that begin the claim process never reach the final confirmation step. Application Insights is connected to the agent. You want to identify exactly which conversational step users abandon most often and visualize the drop-off sequence. Which Application Insights capability should you configure?
- Your organization uses a three-environment ALM pipeline (Development → UAT → Production) for a Copilot Studio agent. A developer exports the solution from Development as an unmanaged solution and imports it directly into Production to apply an urgent hotfix. Three days later, the scheduled managed solution deployment from UAT fails with a conflict error, and several generative answers topics that were working in UAT are now missing in Production. What is the root cause of this failure?
- A financial services company is deploying a Copilot Studio agent that uses generative AI to answer customer questions about investment products. The compliance team requires that the agent must never generate responses that could be construed as personalized financial advice, must block all responses containing specific competitor product names, and must log every instance where a content filter triggers. Which combination of controls best satisfies all three requirements?
- A Copilot Studio agent uses Azure AI Search as a knowledge source to answer questions about company HR policies. HR publishes updated policy documents to SharePoint every Friday. Users consistently report that the agent provides outdated answers on Monday mornings, referencing policies that were changed the previous Friday. The Azure AI Search index has a SharePoint Online indexer configured. What is the most likely cause and the correct resolution?
- A legal department deploys a Copilot Studio agent backed by Azure AI Search over a corpus of 500-page legal contracts. Users report that when they ask about specific clause obligations (e.g., 'What are the termination notice requirements in the Master Services Agreement?'), the agent returns partial clause text that cuts off mid-sentence and lacks the preceding context that defines key terms. The knowledge source uses the default fixed-size chunking with 512-token chunks and no overlap. Which chunking strategy change would most directly resolve the incomplete context problem?
- A security architect is reviewing a proposed Copilot Studio computer-use agent that will automate data entry into a legacy desktop application that has no API. The agent will take screenshots, interpret UI elements, and click/type to perform data entry on behalf of users. The architect raises three concerns: (1) the agent could be manipulated by content displayed on-screen to perform unintended actions, (2) the agent runs under a service account with broad permissions, and (3) audit logs only capture the agent's intent, not its actual screen actions. Which concern correctly identifies a recognized security risk specific to computer-use AI agents?
- A project manager presents four proposed computer-use agent scenarios to an AI architecture review board. The board must identify which scenario is most appropriate for a computer-use agent and which scenario is most problematic from a reliability and governance standpoint. Which evaluation is correct?
- A Copilot Studio agent in a production environment uses a Power Automate cloud flow that combines the SharePoint connector (to read policy documents) and a custom HTTP connector (to call an external vendor API for pricing data). After the IT department applies a new tenant-level DLP policy that places SharePoint in the 'Business' data group and HTTP in the 'Non-Business' data group, the cloud flow begins failing at runtime. Which is the correct explanation of why the flow fails and what must be done to resolve it?
- An enterprise is designing a multi-agent system in Copilot Studio where a primary 'Triage Agent' receives all user requests and routes them to one of three specialized sub-agents: a 'Policy Agent' (knowledge-based), a 'Transaction Agent' (API-connected), and a 'Escalation Agent' (human handoff). From a testing and management perspective, which design decision will create the greatest operational challenge in production?
- A healthcare organization's Copilot Studio agent assists clinicians with medication reference queries. Following a regulatory audit, the compliance team must demonstrate that every agent interaction can be traced to a specific user identity, that the exact AI-generated response delivered to each clinician is retained for seven years, and that any instance where the agent's safety filters blocked a response is logged. The current configuration uses standard Copilot Studio telemetry sent to Application Insights. Which gap must be addressed to meet all three audit requirements?
- An organization maintains a fleet of 12 Copilot Studio agents across three business units (HR, Finance, IT). Each business unit has its own Development and Production environment pair. A new tenant-wide security requirement mandates that all agents must incorporate a centrally managed 'Security Disclaimer' topic that shows a compliance message at the start of every conversation. This disclaimer topic must be versioned centrally and updated simultaneously across all 12 agents when the legal team revises the message. Which ALM architecture best satisfies this requirement while minimizing deployment risk?