Quiz 10 Question 6 of 20

A Constitutional AI critique-revision loop is being implemented for a content moderation helper model. The critique prompt instructs a Claude instance to identify any way the response could cause harm. During testing, the critique model flags nearly every response as potentially harmful, including a response explaining how hand sanitizer works. The revision model then produces overly hedged, unhelpful outputs. What is the most likely cause of this pattern and how should it be addressed?

Select an answer to reveal the explanation.

Motivation