Anthropic describes a tension between Claude being safe and Claude being helpful. How does Anthropic primarily approach this tension?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Anthropic explicitly says that being unhelpful isn't 'safe by default' — refusing too much is its own kind of failure. The goal is genuinely helpful AND genuinely safe.
Full explanation below image
Full Explanation
Anthropic's philosophy recognizes that both excessive restriction and excessive permissiveness are failure modes. An overly cautious Claude that refuses benign requests fails users and undermines trust in AI. An insufficiently cautious Claude that helps with harmful requests causes real harm. The goal is accurate judgment — being helpful with the vast majority of legitimate requests while declining genuinely harmful ones. Option A describes over-refusal, which Anthropic explicitly identifies as a problem. Option C describes unsafe behavior. Option D is wrong — Anthropic sets a default balance that operators can adjust within bounds.