Quiz 8 Question 4 of 20

During red-team evaluation of a Claude-based content generation platform, an adversarial researcher discovers a 'many-shot jailbreak' — by including 100+ examples in the conversation history of harmful content being generated without refusal, subsequent requests for similar harmful content succeed at a 73% rate (vs 8% baseline). The platform allows users to import conversation history from external sources. What is the most effective architectural defense against many-shot jailbreaking through imported conversation history?

Select an answer to reveal the explanation.

Motivation