Managing Risks of Generative AI in Software Testing
CT-GenAI · 75 questions
- A recycling-center scale-house tester asks an LLM for cases covering posted surcharge rules. The reply confidently includes a full pack for a “hazardous-oil surcharge” that never appears in the posted rules. Which GenAI defect does this best illustrate?
- A municipal water-meter portal draft from an LLM lists an acceptance criterion that “a renter may reset a landlord’s leak alert.” That rule is nowhere in the user story. Which defect is this?
- A county poll-book checkout script from an LLM looks tidy but calls screens and APIs the election app does not expose, so it will not run. Which GenAI defect does this represent?
- A community-college registration planner receives this from an LLM: “If the waitlist is full, night labs close; night labs are closed; therefore the waitlist is full.” Which defect type is this?
- A city cemetery-plot system asks an LLM to sequence staff effort. The reply treats “engraving takes two days” as a reason to skip the legal-hold check that must run first. Which defect is shown?
- A municipal dog-license desk is told the LLM “thinks like a clerk” about late-fee ladders. Why do math-like and multi-step test tasks often go wrong?
- A public-beach locker kiosk prompt asks an LLM for guest names for test data. The reply returns only English-language given names even though the city serves several language communities. Which defect does this illustrate?
- A fire-station equipment-checkout pack from an LLM (1) invents a hose not in inventory, (2) issues a radio before the required drill sign-in, and (3) writes locker IDs only in one alphabet. Which mapping is correct?
- A volunteer-fire shift app generates a case that “overtime starts after six hours,” but the posted SOP says eight. The tester compares the pack to the SOP and flags the extra rule. What detection approach is this?
- A municipal compost-pickup pack invents a “holiday skip if snow exceeds two inches” rule that looks plausible to a new tester. The solid-waste supervisor says that rule was never adopted. Which detection method does this exemplify?
- In one generated senior-center meal-delivery pack, one case says the cutoff is 10:00 and another says 11:30. Both sound locally plausible. What should the tester conclude and do?
- A city-zoo membership script from an LLM says: open join page → tap Pay → then enter card → then tap Submit, attempting payment before a card exists. How should the tester classify and detect this?
- A public-radio pledge script is syntactically tidy; when run against the pledge form it clicks a control that is not on the page. Which detection approach revealed the defect?
- A municipal court-scheduling strategy requires weekday, evening, and interpreter-assisted bookings, but the LLM’s synthetic data is weekday-English only. How should the tester identify this problem?
- A county 311 intake pack from an LLM is all happy-path functional clicks; load, accessibility, and privacy cases never appear even though the charter lists them. What defect does this show?
- A public-school bus-routing pack will feed a safety-critical stop list. A colleague wants the same light skim used on a cafeteria-menu wording draft. What should the tester select?
- A community-garden plot lottery LLM keeps inventing a “senior override” because the prompt omitted the posted eligibility table. Adding that table as input data stops the invention. Which mitigation is this?
- A municipal ice-rink booking pack botches a three-rule fee ladder in one shot. The tester splits the work into “list rules → apply one booking → check the total” and reviews each reply before the next. Which mitigation is named?
- A community-band rehearsal-hall brief arrives as a scanned poster with overlapping notes; the LLM mixes call times. The tester re-supplies the same facts as a simple table. Which mitigation does this illustrate?
- A municipal wedding-license desk uses a general chat model for a multi-step waiting-period calculation and keeps getting the order wrong. Switching to a model suited to that logical task is proposed. Which Chapter 3 mitigation is this?
- A county flood-alert pack from one model invents a “text all residents at 3 a.m.” acceptance criterion; a second model stays inside the posted policy. The tester compares both replies before promoting any case. Which mitigation is this?
- A municipal lost-and-found chatbot is prompted only with “write tests” and no desk rules. The pack invents shelves and reverses claim order. What does this situation illustrate about defect likelihood?
- A community-tool-library lead wants this chapter’s risk item to teach embedding stores and a tuning job so invented loan rules stop. What is the correct Chapter 3 stance?
- A municipal bulk-trash scheduler sends the same prompt twice and receives two different pickup-route cases. What causes this non-deterministic behavior?
- A county animal-control script pack changes verbs on every run. The tester is advised to reduce randomness so wording stays more consistent. Which action matches that advice?
- A municipal after-school sign-out generator becomes repetitive once the team lowers temperature for more stable wording. What trade-off does a low temperature setting introduce?
- A city taxi-medallion desk needs stable automation locators from an LLM, not brainstormed edge-case names. Which temperature choice best fits that reproducibility-first task?
- A community-kitchen rental tester needs yesterday’s GenAI data pack again to compare two prompt versions fairly. Which control reuses the same pseudo-random sampling sequence in implementations that support it?
- A sidewalk-cafe-permit lead claims that setting a random seed “turns the LLM into a deterministic rule engine.” What does a seed actually provide?
- A county deed-recording pair sees two GenAI runs diverge and debates whether to change temperature or set a seed. How do these two mitigations differ?
- A public-art mural-permit pack still drifts on long scripts even after the team uses a low temperature and a seed. What should the team remember about reproducibility and verification?
- A snow-emergency parking tester pastes a live tow list with names and plates into a chatbot; a later summary repeats a plate that should never have left the yard. Which privacy risk does this illustrate?
- A community-choir ticket desk learns the chatbot vendor keeps prompts to “improve the model,” and the desk cannot say who sees last night’s donor list. Which privacy concern does this describe?
- A municipal birth-certificate counter wants a public chatbot to draft cases from real applicant rows that include names, dates of birth, and addresses. What compliance risk should the tester raise?
- A county jury-summons tool embeds an LLM beside the juror file used in test workflows. Beyond ordinary application bugs, what security risk should testers recognize?
- A public-boat-ramp kiosk team is briefed that someone might try to alter the test assistant’s behavior or pull sensitive slip-holder data through crafted input. Which risk class does this briefing describe?
- A municipal parking-boot desk is warned that planted files in the test-data drop could push the assistant into wrong conclusions about who is booted. Which risk does this warning name?
- A city bus-transfer tester sorts two briefings: (1) yesterday’s prompt still sits on a vendor disk, and (2) an outsider might reach the test-tool account. How should these concerns be mapped?
- A community-orchard harvest kiosk security brief asks testers to name the four official GenAI security vectors in syllabus v1.1. Which set is correct?
- A fire-hydrant-use-permit team is told that context manipulation targets confidential training data, for example by overloading the context window until stray training snippets appear. What is the correct tester stance?
- A county septic-inspection study sheet still labels a §3.2.2 row as “data exfiltration.” Which v1.1 name should the candidate select for that row on the exam?
- A dark-sky-park reservation assistant starts emitting fragments that look like another city’s logs, not the posted park rules. What is the appropriate tester response to this possible context-manipulation leakage symptom?
- A municipal leaf-collection brief defines a GenAI security vector as introducing data that disrupts the AI’s output — for example, an image that lures the model into another context and provokes hallucinated acceptance criteria. Which vector is that?
- A city food-truck-permit tester attaches a vendor sketch; the generated acceptance criteria include “skip health-sticker if the truck is electric,” which is not in the ordinance. Security later says the sketch was tainted input. What should the tester conclude?
- A workshop kiln-booking lead hears that people may submit fake evaluations when rating an AI-generated test report. Which official security vector names manipulating training or rating data this way?
- A municipal storm-drain report assistant keeps marking “blocked inlet on a school walk” as low severity after a week of unexplained five-star ratings on sloppy drafts. What should the tester do?
- A public-observatory booking generator is supposed to emit a UI script. Review finds an unexpected outbound “diagnostics” channel that nobody requested. Which security vector has the reviewer likely found?
- A municipal yard-waste tester wants to execute a freshly generated keyword script in the shared garage because it compiled cleanly. Given malicious code generation as a named risk, what is the correct stance?
- A city crossing-guard scheduler is given two risk cards: (A) the model may leak training snippets after its window is abused; (B) a planted image may push invented acceptance criteria. How should the cards be mapped?
- A community-beehive-registration pair must label two events: (1) someone stuffed fake “looks good” scores on generated inspection reports; (2) a generated hive-sensor script grew an unexpected remote call. How should these events be mapped?
- A municipal dump-sticker designer pastes the city’s unpublished fee table and a vendor’s unreleased locator map into a public chatbot to draft test cases. Which privacy or security concern does this scenario best illustrate?
- A county plat-map desk learns that someone is submitting bogus star-ratings on the AI-written weekly test report to skew what the team trusts. Which named vulnerability does this scenario match?
- A public-sauna reservation tester is handed a huge untrusted paste plus a mystery PNG and told to “just generate the pack” with an LLM. What is the most appropriate tester gate?
- A municipal pothole-report tester is about to paste a full citizen export—phones and emails included—into an LLM just to draft three cases. Which mitigation principle should guide the prompt?
- A city tree-removal permit desk replaces real owner names and parcel IDs with tokens before asking an LLM for boundary-related test data. Which mitigation strategy is being applied?
- A community-darkroom booking log with member IDs will travel to an LLM-powered test tool. Which mitigation best addresses protection of that data in motion and at rest?
- A municipal parking-garage overnight crew keeps pasting gate-camera stills into a public chatbot because “nobody said not to.” Which organizational mitigation directly addresses this pattern?
- A county health-inspection pack includes a generated script and a severity table from an LLM. Which mitigation is essential before the pack is used?
- A public-archive retrieval team handles sealed personnel files and wants stronger GenAI privacy and security controls for test tasks. Which pair of complementary mitigations best fits the syllabus guidance?
- A municipal tree-planting-permit program will keep an LLM in the test toolchain for years. Which set of ongoing controls best matches the syllabus mitigations?
- A municipal ice-cream-cart permit desk compares a one-line text rewrite with a long multi-file analysis in an LLM-powered test tool. What primarily drives the difference in energy consumption?
- A municipal ski-tow ticket team wants a model to paint a fake lodge UI for every case instead of only writing the steps in text. Which energy comparison from the syllabus best applies?
- A county library-of-things tester shrugs that “one more regenerate” is nothing, then the squad does it on every case every sprint. Which environmental point does this illustrate?
- A public-pier fishing-license kiosk regenerates the same happy-path pack six times “for luck” with no prompt change. Which official environmental practice should the tester apply?
- A municipal campground lead refuses to care about GenAI energy use until given an exact kilogram of CO₂ per test case. Which syllabus stance best responds?
- A city fountain-show booking chatbot is used all afternoon from testers’ laptops for GenAI-assisted case drafting. Where does that usage increase load and energy use?
- A municipal greenhouse-plot tester can either attach twenty decorative site photos or paste the three posted rules as text when asking an LLM for cases. Which environmental choice is better when the text is sufficient?
- A county weigh-station draft item begins: “Open the accredited-course energy simulator and enter token counts.” Why should that stem be rejected for the written CT-GenAI bank?
- A city hall-rental test lead asks which named instrument specifies requirements for managing AI systems in the organization and promotes consistent, reliable GenAI-in-testing practice. Which answer is correct?
- A municipal boat-slip waitlist team wants the named framework for AI systems using machine learning that stresses lifecycle, data quality, transparency, and safety for GenAI used in testing. Which instrument is that?
- A sidewalk-snow-shovel registry will use GenAI on a high-visibility city service. Which named instrument is the regulation that classifies applications by risk level and mandates transparency, accountability, and bias mitigation for GenAI used in testing?
- A county 4-H fair-entry desk wants the named framework—not a statute—that offers guidelines for managing AI risks with a focus on fairness, transparency, and security, and that supports preventing biased test results. Which instrument is it?
- A municipal outdoor-movie-permit board shows four cards: ISO/IEC 42001:2023, ISO/IEC 23053:2022, the EU AI Act, and NIST AI RMF 1.0. Which labeling of instrument types is correct?
- A homestead-exemption desk must match four GenAI-in-testing needs: (1) manage AI systems for consistent practice; (2) ML lifecycle, data quality, and safety; (3) legal risk classes plus accountability and bias duties; (4) fairness, transparency, and security guidelines that help avoid biased test results. Which mapping is correct?
- A county clerk wants a bank item to quote EU AI Act article numbers and every NIST function in detail. What should the exam writer recall about scope?