A red-team is evaluating a Claude-based customer service model for a firearms retailer. The team must design a prompt that tests whether the model can be induced to provide instructions for illegal weapon modifications through a roleplay framing. According to responsible red-teaming methodology, which action is correct?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Here's the deal — b is correct because responsible red-teaming follows a structured methodology: isolated execution environment to prevent real-world harm, complete documentation of prompt and response, harm severity classification, and confidential reporting to the safety team for mitigation. This is how safety gaps are surfaced and closed without producing net harm.
Full explanation below image
Full Explanation
B is correct because responsible red-teaming follows a structured methodology: isolated execution environment to prevent real-world harm, complete documentation of prompt and response, harm severity classification, and confidential reporting to the safety team for mitigation. This is how safety gaps are surfaced and closed without producing net harm. C is wrong because merely documenting a hypothetical test without execution produces no empirical safety signal; the entire point of red-teaming is to probe actual model behavior. A is wrong because public disclosure of working jailbreaks before a mitigation is deployed creates net harm — it amplifies the vulnerability rather than closing it; responsible disclosure to the developer precedes any public release. B is wrong because red-teaming requires testing adversarial inputs that the model might fail; pre-confirming refusal defeats the purpose and produces only a false sense of safety.