Quality Characteristics for AI-Based Systems
CT-AI · 23 questions
- A water-treatment lab scores its algae-bloom predictor two ways: one checklist about the built product and another about use in the plant. How should a tester place ISO/IEC 25059?
- A tile-glaze line used to treat every wrong shade call as a failed release. The new vision grader is probabilistic, so the quality plan now names an acceptable band of wrong calls instead of demanding zero misses. Which 25059 characteristic is that?
- A greenhouse climate planner is already in production. When a storm front changes humidity, it retargets vent and mist setpoints on its own, without a code drop. Which 25059 characteristic is that?
- A ski-patrol dispatcher can hit a “hold all” key and take route assignment away from the advisor while the next sled is still being staged. A consultant calls that “usability.” Which 25059 characteristic is that?
- A university archives desk must tell researchers which captioning model is live, where its training corpus came from, and where the model card lives. The same information is also meant to support a user’s sense that the service is understandable. Which 25059 characteristic is that?
- A recycling-belt sorter still meets its agreed sort-error band when a share of frames are overexposed, a few labels in the feed were skewed, a jam shakes the camera, or an operator feeds a bag the wrong way. Which 25059 characteristic is that?
- A hydroelectric gate advisor is about to recommend opening a spillway. The night operator must be able to stop that act in time so the river below is not put at risk. Security, not the control-room “ease of use” checklist, owns that freeze. Which 25059 characteristic is that?
- A municipal hiring screen must not treat named groups differently on the agreed fairness metric, must drop fields that identify applicants, and must leave the final offer with a human panel. Which 25059 characteristic is that bundle?
- Two freeze buttons sit on the same quarry-loader console. One lets a supervisor take steering back so the shift can finish the load. The other exists only so an operator can abort a move that would endanger a walker in the pit. How should those two be classified?
- A rare-book lab asks why the test plan now talks about error bands, post-go-live retargeting, seize times, model-card visibility, dirty-input endurance, hazard freezes, and a fairness gap. What should the tester see?
- A mine-conveyor stop is supposed to “keep hands off the belt.” The conventional interlock had numbered requirements down to code. The vision add-on is specified mostly as “learn from last year’s incident clips.” What safety challenge should the tester explain?
- A bakery oven-guard model, given the same thermal frame twice, can emit two slightly different “pull tray / leave tray” commands because of small input jitter or internal randomness. A safety assessor wants a guarantee of one exact command. What obstacle should the tester explain?
- A wildlife-crossing barrier was safety-tested in March. By August it has been updating itself from new trail-camera nights, and the March test no longer describes what it does. What should the tester explain?
- A ski-lift hold model dropped a chair line and nobody can reconstruct the reason from the weights. An explainability add-on might sketch a local reason, but it is not widely available and the vendor warns it can slow the hold decision. What safety problem should the tester explain?
- A medical-device brake that uses a vision cue is treated as a safety component under a risk-based statute and therefore draws extra development and test duties. The plant’s older functional-safety standard never names AI; a sibling standard still forbids it in the safety function. What regulatory picture should the tester explain?
- A paper-mill trip system has a conventional pressure interlock beside a new acoustic “web-break” model. The assessor lists five extra AI headaches: fuzzy goals in the sound archive, non-repeatable calls, a model that keeps learning, a decision no one can narrate, and rules that changed twice this year. What should the tester recognize?
- A cider-mill bruise grader cannot be accepted with “never wrong.” Legal and ops instead want a numeric error band, a timed seize, and a fairness gap, each with a threshold. What typical shape of AI acceptance is that?
- A herbarium’s leaf-ID helper may mis-label at most 4% of a held-out tray of rare sheets, and recall on those rare sheets must stay at or above an agreed floor. What kind of acceptance criteria are those numbers?
- A greenhouse vent planner must retarget humidity setpoints within eight minutes after a 20-point outside-humidity jump, without a technician pushing a new build. What acceptance criterion is that?
- A radio-archive loudness advisor must yield to a board-op “take the fader” key within two seconds, and must fully mute itself if the key is held and no human ack arrives. What acceptance criterion is that?
- Every scored insurance-claim file must carry the live model’s unique version id, a link to its model card, and the vintage of the training corpus, matching the firm’s disclosure rule. What acceptance criterion is that?
- A sawmill knot classifier must keep its agreed error band when 12% of frames are motion-blurred, when a share of labels in a dirty feed are skewed, and when the mill lights flicker for half a minute — it may degrade to a slower, coarser mode, but it must not crash or silently abandon the band. What acceptance criterion is that?
- Three acceptance lines sit on one dam-and-hiring program: any “open spillway” advice stays frozen 45 seconds so a licensed operator can cancel it; a hiring screen’s agreed fairness metric may not exceed a stated gap across named groups; the conventional over-pressure interlock must still trip even if the acoustic model is unsure. How should those three be mapped?