Quiz 8 Question 9 of 20

A Constitutional AI researcher is evaluating how different principle orderings in Claude's training affect model behavior on edge cases. They observe that when 'be helpful' and 'avoid harm' conflict, Claude's behavior is inconsistent — sometimes prioritizing helpfulness (providing requested information), sometimes prioritizing harm avoidance (refusing). The inconsistency is highest for dual-use information (security research, pharmacology, historical atrocities). The researcher wants to understand the correct architectural principle for resolving these conflicts. Which best describes Anthropic's Constitutional AI approach to this tension?

Select an answer to reveal the explanation.

Motivation