An organization's AI red team discovers that a deployed LLM can be made to ignore its system prompt and generate harmful outputs by prefixing requests with a specific multi-token sequence. What risk management response is MOST appropriate?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Here's the deal — a is correct because a discovered jailbreak that allows harmful outputs is an active safety vulnerability requiring immediate response — the model should be taken offline to prevent exploitation, and the incident should be processed through formal AI incident management including executive notification, root cause analysis, and remediation. B delays response to an active vulnerability.
Full explanation below image
Full Explanation
A is correct because a discovered jailbreak that allows harmful outputs is an active safety vulnerability requiring immediate response — the model should be taken offline to prevent exploitation, and the incident should be processed through formal AI incident management including executive notification, root cause analysis, and remediation. B delays response to an active vulnerability. C bypasses formal incident management and executive accountability. D could enable widespread exploitation before a fix is deployed.