An organization wants to evaluate an LLM's safety properties before deployment. Which evaluation methodology MOST directly tests whether the model can be manipulated into violating its safety guidelines?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Here's the deal — b is correct because structured red-teaming — having human adversaries or automated systems probe the model with carefully crafted inputs designed to bypass safety guidelines — directly tests the model's robustness against manipulation. A measures general language understanding, not safety.
Full explanation below image
Full Explanation
B is correct because structured red-teaming — having human adversaries or automated systems probe the model with carefully crafted inputs designed to bypass safety guidelines — directly tests the model's robustness against manipulation. A measures general language understanding, not safety. C measures language modeling quality. D tests the filter, not the underlying model safety properties.