Model Testing for Machine Learning Systems
CT-AI · 67 questions
- A community blood-bank inventory model keeps under-ordering for one neighborhood clinic. A tester is handed the Chapter 6 risk table. Which named mitigation should come first?
- A community-loom pattern ranker starts promoting designs that copy a living artisan’s signature motif. The risk is an unethical model, not a math miss. Which table mitigation should the tester pick?
- A municipal leaf-collection router is scored on “bags collected.” Crews start looping the same easy blocks. Which table mitigation matches a model that exhibits reward hacking?
- A greenhouse-vent model keeps humidity inside the stated band but slams louvers so often the motors fail. Accuracy still looks fine. Which table mitigation should the tester pick?
- A cider-press yield model has no affordable gold gallons figure for each run. The table lists the test-oracle problem. Which named mitigations should the tester pick?
- A tide-pool species ID model fails when a plate is only slightly scuffed. The risk is lack of AI robustness to unexpected inputs. Which table mitigations should the tester pick?
- A maple-sap yield classifier misses the agreed recall on an independent pile. Which table mitigation matches failure to achieve required model performance measures?
- A snow-fence placement brief says “keep drifts off the road” and nothing else. The table lists poor system requirements. Which mitigation should the tester pick first?
- A bike-share rebalancing model ships with a one-line note and no intended use, measured score, or interface. Which table mitigation matches poor model documentation?
- A clock-tower chime predictor was swapped overnight. Four desks report four different pains: the kiosk dies on first request; the live hit-rate is worse; the new model disagrees with last week’s on the same rows; last winter’s accepted layouts now fail. How should those four update risks be matched?
- A municipal grit-bin refill team is told to “write some docs.” Two named frameworks sit in the syllabus: a concise model overview (intended uses, evaluation, ethics) and a standardized dataset description (motivation, composition, collection, uses). Which review artifacts should the tester name?
- A cheese-cave humidity model is machine-fit, opaque, data-dependent, and swapped every thaw. A regulator asks why testers review a card instead of “just reading the code.” What should the tester explain?
- A bandstand crowd-noise card lists two different intended users, an old accuracy, and no hardware note. What should a documentation review hunt for?
- A wind-turbine icing card describes the algorithm and the training pile but is silent on whether the test pile was independent and whether adversarial or functional activities were run. Which documentation checklist section should the tester flag?
- A fountain algae model’s card never states the deploy environment, monitoring alerts, retraining or rollback plan, or how the model will be retired. Why are those contents in-scope for the review?
- A community apiary winter-loss model is treated as high-risk by the city’s AI rule. Go-live is blocked until a documentation audit passes. What extra purpose does that review serve?
- A ferry boarding-wait model is right on the first Friday boat. A sponsor wants a pass sticker. What should the tester require instead?
- An avalanche-beacon “burial likely” model’s brief only says “be accurate.” Testers cannot size the pile. What must acceptance criteria state?
- A library hold-time model looks strong on the slice used to nudge hyperparameters. Which pile may be used for the statistical ML functional-performance test?
- A beehive mite-risk model is scored on last spring’s calm yards. Summer yards are hotter and louder. What must be true of the test pile?
- A compost-odor model’s run finishes. A slide says “87% and we ship.” How should the tester rewrite the report?
- A choir-attendance model is being scored against a stated margin and confidence. After a large prefix of the planned pile, the running score is stably inside the band. What official idea may testers use, without cherry-picking a lucky streak?
- A street-tree failure model sits on a path over a playground. The safety desk wants a stricter reliability claim than a typical accuracy band. What kind of statistical criterion may testers apply?
- A parking-occupancy team can afford a bigger test pile. One camp wants to keep the confidence level and shrink the margin. The other wants to keep the margin and raise the confidence. What do those two official trade-offs mean?
- An ice-fishing hut-permit photo model is shown a plate that looks unchanged to the clerk but is now scored as the wrong hut class. What should the tester name that successful input and the activity that found it?
- A mushroom-cap ID kiosk fails on a naturally scuffed cap and, separately, a researcher asks whether someone could aim such a change at the model. What distinction should the tester keep?
- A hydrant-flow camera model is about to go live. A tester argues for an adversarial pass. What is the official purpose of that pass?
- Testers cannot see inside a pottery-glaze photo model. They raise a similar model whose internals they do know, search that stand-in, then replay the interesting inputs on the original, assuming shared decision boundaries. What approach is that?
- A swimming-lane occupancy camera is a closed box. The desk has no stand-in model. What brute-force black-box approach should the tester recognize?
- A salt-barn bag-count model’s architecture, parameters, and training story are available to the test desk. What should white-box adversarial testing mean here?
- A telescope cloud-cover desk can either sit with a few carefully chosen sky plates or run a tool that produces a large number of variations. How should those two ways of performing adversarial testing be contrasted?
- An oyster-reef cover model has one survey photo that already passed. Testers will change that source using a stated relation and check how the output must move. What technique is that?
- A drawbridge gap-time model is opaque and too costly to simulate by hand. Relational properties would still give confidence. Why should testers choose metamorphic testing rather than a traditional expected-result oracle?
- A playground surface-wear model scored a source photo as “resurface soon,” and that case passed. Testers take a second photo of the same panel under the same light a minute later. What follow-up case and consistency expectation should they derive?
- An alpaca-fleece grader outputs a micron score. A source case passed. Holding other fields fixed, a higher grease-content reading should not produce a finer (lower-micron) grade. What follow-up should testers derive?
- A storm-drain clog model passed on a grate photo. Testers rotate that photo by a small amount that does not hide the debris. What follow-up and invariance expectation should they derive?
- A carousel-load model passed on a quiet Tuesday car. One relation adds an empty duplicate sensor row. Another adds a documented extra rider mass. What should testers derive for those two follow-ups?
- A stained-glass color-match model has never produced a passing source case for a pane family. A writer still wants follow-ups. What limitation should the tester recall?
- A pier-piling tilt model has a passed slack-tide reading. Physics of the berth: a documented higher water level, other fields fixed, should not reduce predicted tilt. How should testers derive the follow-up?
- A sled-hill packing-density desk invented three metamorphic relations over lunch. How should those relations themselves be validated before they mint follow-up cases?
- A kite-festival wind model passes every follow-up from one incomplete relation and still misses an absolute speed error. What should the tester treat that as?
- A skating-pond thickness model was fit on mid-winter sensor rows. Late-season rows are warmer and noisier. What should the tester name that change?
- A CSA-box “will they skip this week?” model still sees the same signup fields, but after a new pickup rule the old “loyal” pattern now predicts the wrong skip. What should the tester name that change?
- After a plow-contract change, a grit-bin model sees new truck IDs in the feed. After a new ice-treatment chemical, the same temperature-and-traffic row should now be called “treat” instead of “hold.” How should those two changes be labeled?
- A dog-park peak-occupancy model is live. Rangers file a current head-count. Testers compare that feedback to the model’s output and to a threshold. What activity is that?
- A marionette-show length advisor asks patrons to rate the suggested runtime after the curtain. How should testers treat that rating in dynamic drift testing?
- A community-garden watering model gets no thumbs-up. Testers instead watch whether beds the model left dry actually dried out. How should testers treat that observed outcome?
- A trolley-bunching model is live, and nobody is labeling every stop. Testers compare the statistical shape of incoming features and of predicted outputs with the shapes seen at fit time, using a named distribution test. What activity is that?
- A darkroom exposure advisor can memorize the training strips, miss the pattern on every strip, or land in between. What three outcomes should testers name, and when does the detection work sit?
- A cemetery plot-demand model recites last decade’s odd memorial-week spikes and then fails on next month’s unseen weeks. What should testers name that behavior?
- A loom-thread tension model looks strong on the validation slice. Testers run an independent test pile that includes uncommon weave patterns the fit never saw. The test score is much worse than validation. What should that gap signal?
- A glass-anneal schedule model is a shallow tree, and the training pile never recorded kiln-zone, which the glass shop says drives the outcome. Both training and validation scores stay poor. What should testers name that?
- A sled-dog rest-stop model posts consistently low accuracy, precision, recall, or F1 on both the training set and the validation set. How should testers read that pair?
- A lighthouse foghorn-timing desk plots learning curves. Training error and validation error stay high and close, with little improvement as epochs pass. What does that picture indicate?
- An orchard-blight model is still in generation. Scores are poor on the fit pile and the tune pile. A colleague calls it “drift.” Another wants a Chapter 3 speech about a thin three-way split. How should testers keep this item?
- A boardwalk crowd model has variant A (current) and variant B (a candidate). Both see the same style of gate counts. Testers will need several runs and a statistical comparison. What approach decides which variant is better?
- A quilt-block ranking desk has no gold “best layout” for every block. They already have last season’s ranker. How should testers treat A/B testing here?
- A garden-watering model was retuned. Agreed ML functional-performance metrics exist. What should A/B testing check before the update replaces the previous variant?
- A trolley-bunching advisor changes itself overnight. Automated checks compare characteristics after the change with those before. When should testers accept the change?
- A dog-park occupancy desk wants a fair call between two models. Which statistical comparisons does the syllabus name for A/B testing of ML systems?
- A marionette-lighting advisor has two implementations. One camp wants to know which scores better on agreed metrics. The other wants a pseudo-oracle that hunts disagreements as defects. How should those jobs be assigned?
- A grit-bin refill desk has no gold “hours-to-empty” for every bin. They run the candidate model and an alternative version on the same inputs and compare outputs. What should testers name that activity and the alternative?
- A cheese-cave humidity pair was built by the same desk on the same open toolkit. Both miss the same midnight spike and “agree.” What must be true of the pseudo-oracle?
- A drawbridge opening model is new. The harbor still has a conventional coded timetable that solves the same “may we open?” question. What may testers use as a back-to-back pseudo-oracle?
- A wind-turbine icing model must answer in a tight field window. Testers raise a slower lab twin that only has to match functional outputs. What is enough for that back-to-back check?
- A lighthouse visibility model is being moved from the development closet to the pier kiosk. Testers feed both environments the same frames. Which check should they pick?
- An orchard-blight pair is scored two ways. A/B would ask which variant wins on F1. Testers instead want disagreements on odd leaf edges as defect clues. Which technique should they pick?