What is 'red-teaming' in the context of AI safety?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Red-teaming AI means adversarial testing — experts try to break the model, find jailbreaks, elicit harmful outputs, and discover failure modes before bad actors do.
Full explanation below image
Full Explanation
Red-teaming is borrowed from cybersecurity: an adversarial team (red team) actively tries to defeat the safety measures of a system to find vulnerabilities before they can be exploited in deployment. For AI models, red-teaming involves: attempting jailbreaks and prompt injections, testing for harmful content elicitation, probing for dangerous capability uplift, testing edge cases and unusual scenarios, and checking behavior under adversarial conditions. Anthropic conducts extensive red-teaming before model releases as part of its safety evaluation process. Option A is absurd. Option C misappropriates competitive gaming language. Option D is nonsensical.