Meridian's evaluation team struggles to find enough real historical examples of rare, severe ramp-safety incidents to fully test the computer-vision model against. How might generative AI be used, consistent with CPMAI's view of generative AI as a project accelerator, to strengthen the Model Evaluation phase here?
Select an answer to reveal the explanation.
Short Explanation
This is a legitimate, practical use of generative AI at the evaluation stage — use it to generate realistic synthetic edge cases so you can actually test the rare scenarios you don't have enough real examples of, not to fake results.
Full Explanation
CPMAI treats generative AI as an accelerator woven across the six-phase lifecycle, and Model Evaluation is a legitimate place for it: when real historical examples of rare, severe events are scarce — as they typically are for safety incidents by their very nature — generative AI can help produce realistic synthetic test scenarios (varied lighting conditions, unusual debris shapes, uncommon ramp-activity patterns) that supplement limited real data and give the evaluation broader edge-case coverage than the historical record alone provides. Saying generative AI has no role in evaluation contradicts CPMAI's explicit framing of GenAI as a cross-phase accelerator, not a tool confined to early phases. Using generative AI to fabricate fake incident reports and present them as genuine historical events is a serious integrity violation — synthetic test data must be clearly labeled and used transparently to broaden test coverage, never disguised as real events to inflate perceived evaluation completeness. Replacing the entire evaluation and Go/No-Go process with an AI-written report abandons actual model testing altogether, which defeats the entire purpose of the Model Evaluation phase and would be reckless for a safety-relevant system. The exam point: generative AI's legitimate accelerator role in evaluation is generating supplementary synthetic test material, transparently labeled as such — not fabricating evidence or replacing genuine testing.