Testing AI-Based Systems
CT-AI · 52 questions
- A district-heating leak detector ships with a deep network whose weights and biases are frozen at go-live and can change only if the team retrains. A locker-room climate agent keeps changing setpoints from a live reward after opening night. Which system is locked, which is adaptive, and which is easier to test?
- A sidewalk-heave classifier is locked in the field, then the vendor ships a new fitted snapshot every quarter. Testers still have last quarter’s expected-result file. How should they treat the update?
- A public-pool turbidity classifier is locked, yet the same frame scored on two parallel graphics nodes can differ by a few thousandths. Stakeholders call that a defect in the lock. What should the tester recognize?
- Testers run a long weekend of simulated demand through a bike-share rebalance agent that learns from a reward. Monday’s routing no longer matches Friday’s. What testability problem is that?
- After a storm week a food-bank restock agent invents a transfer habit testers never enumerated. The board asks for the pre-written cases that cover it. What should the tester say?
- A streetlight-outage agent adapts after a heat wave. Testers keep an automated suite aimed at core routing and safety stops and rerun it whenever the agent changes a lot, plus a watch that the miss rate has not crossed a stated band. What pattern should the tester pick?
- A town-clerk minutes summarizer does not change weights during a sitting, but the vendor drops a new snapshot every month. Where should the tester place that GenAI pattern?
- Before opening, testers put an ice-rink thickness advisor in a closed hall, shift the fake weather, and watch whether the learning rule adapts in a safe direction. What should the tester know about that rehearsal?
- An orchard frost-alert once calls a freeze a thaw. Growers want the release failed on that single night. What should the tester explain?
- A community-clinic no-show predictor is given the identical booking twice and returns two nearby scores. A conventional scheduler would have repeated exactly. Why do testers look at distributions and confidence instead of a single expected bit?
- A recycling-sort model was fitted on weekday household bags. Weekend festival bags have a different mix. What sample should testers use for a statistically meaningful verdict?
- An elevator-fault predictor is highly confident on a rare hydraulic sound and still wrong. The board wants a number they can defend. What should the tester point to?
- A watershed flood-stage classifier must show a safety threshold with high confidence to a regional board. Three vivid examples will not do. What kind of evidence should the tester use?
- A community-solar output model uses stochastic elements in the net and also emits probabilistic scores learned from years of sky logs. Either source can make two runs differ. Why is a statistical approach needed?
- A school-bus late-arrival estimator is scored like a conventional timer: each run is pass or fail. Why is that the wrong verdict for this system?
- A mountain-rescue radio-triage model emits an urgency score. Dispatchers cannot name one correct number for a garbled call. Testers agree a band and a tolerance. Which oracle workaround is that?
- A botanical-garden pest identifier started as an exploratory pilot. Requirements are still thin and keep moving. Why is the oracle hard?
- An archives handwriting transcriber is run on eighteenth-century court rolls. Checking every token by hand is not practical. What oracle challenge is that?
- A museum caption writer is judged by visitor taste. Two curators call the same caption a pass and a fail. What oracle problem is that?
- A self-learning public-pool chemical advisor keeps updating after new readings. Last month’s expected-result cards no longer match what “correct” means. What oracle problem is that?
- Port container-damage vision scores wander when the bay lighting, chill-store temperature, or camera-link lag change. Testers write those environment values into the test setup so repeats are comparable. Which oracle aid is that?
- A cooperative has two grain-moisture estimators and no lab gold label for every wagon. Which oracle workaround should the tester choose?
- Wildlife night-time audio has almost no labels. Testers train a proxy on a small labeled subset and use it to score the unlabeled pile. What oracle is that, and what risk remains?
- Testers of a public-radio playlist writer send station briefs and then judge clarity, freshness, and house-rule fit of the produced hour. They never open the weights. Which GenAI test approach is that?
- A county permit-triage assistant can be reached from a clerk console or an API, with an optional house brief and a user packet that can hold hundreds of pages. Which GenAI testing problem is that?
- A community-college placement-essay commenter is given the same essay twice: once with a high temperature and a generous token cap, once with a low temperature and a tiny cap. The two comments are not the same product. What must testers control?
- A regional-rail delay explainer keeps prior turns. A tester’s fifth question is answered as if the first rumor were still true. Which GenAI factor changed the output?
- A building-permit drawing-completeness assistant is reviewed by humans against criteria written in the requirements (missing stamp, unreadable scale, absent egress note). There is no single expected paragraph. How is the result decided?
- Generated garden wayfinding maps are scored by a vision checker. What should the tester allow, and what risk remains?
- A town-minutes summarizer is compared on a published language-understanding suite so two vendor snapshots can be lined up. Testers also log processor, graphics, memory, network, and response time on inference. What should stay in the GenAI plan?
- A clinic-advice assistant is about to leave internal QA. A lead is told to start red teaming. What is the first official step?
- Library staff want red teaming run on the lobby kiosk after lunch. Where should the tester require access?
- A permit assistant can be probed for security holes (named types only: indirect prompt injection, hidden content in retrieved documents) or for unsafe answers to an ordinary resident question. How should a tester tell those two red-team aims apart?
- A minutes summarizer has finished internal quality checks and is a week from public launch. When should red teaming run?
- Testers hold long multi-turn sessions with a rec-center coach-bot, using open exploration and a short checklist. They then group the failures and turn them into a dataset the builders can use to harden the system. Which later official red-team steps are those?
- A county assistant’s input space is huge. A small expert desk cannot cover it. Which coverage tactic should the tester choose from the official set?
- A minutes summarizer already posts a strong score on a static language suite. A sponsor wants to skip red teaming. What should the tester say?
- After a pre-launch red-team week, operations ask what still happens on the live assistant. How should the tester contrast the two teams?
- A dairy somatic-cell predictor still needs component, integration, system, and acceptance work. Which two extra ML-specific levels should the tester also name?
- Testers argue whether input data testing is only about the historic fitting set for a dairy predictor. What should the tester say it covers?
- A lead asks what ML model testing is for on the dairy predictor. What should the tester say it concerns?
- A rink kiosk has a clerk screen, a data pipe, and a radio link that are not models. Which testing still applies to those pieces?
- Testers must show that cleaned frames arrive in the shape the thickness model expects and that scores reach the kiosk screen. The same shop also buys a weather score as a service. Where do those checks sit?
- The thickness model is embedded in the full kiosk build; a smaller compressed copy is also tried. What should system testing check?
- One team checks interfaces and data exchange with an outside weather bureau in an environment that looks like the rink. Another team decides whether a bought scoring service meets the shop’s ML performance bar. How should those two be split?
- A water-utility board asks why the MLS test plan is risk-based instead of testing everything the same. What should the tester say?
- A solar-output project picked a poor algorithm, a thin evaluation approach, and a development toolkit with a known security hole. Which risk lane is that?
- A transit late-bus model was fitted on one depot’s winter logs, the live pipe drops night rows, and weekend routes were never in the pile. Which risk lane is that?
- A pest identifier misses its agreed score, looks overfitted on greenhouse shots, and flips on tiny input tweaks. Which risk lane is that?
- A second facilitator wants the MLS risk list grouped by ISO/IEC 25059 names (correctness, robustness, societal and ethical mitigation, and the rest) instead of by workflow lane. What should the tester accept?
- A lead has three named forms on the table: data-pipeline testing, adversarial testing, and a review of algorithm or model suitability. How should those attach to risk lanes without teaching later-chapter methods?
- A sponsor hears about input data testing and model testing and wants to drop component, integration, system, and acceptance work. What should the tester say?