What is 'sycophancy' in AI models and why does Anthropic train Claude to avoid it?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Sycophancy is AI yes-man behavior — agreeing with users even when they're wrong, validating bad ideas, backing down from correct positions. Anthropic trains against it because it undermines Claude's usefulness.
Full explanation below image
Full Explanation
Sycophancy in AI occurs when models systematically prioritize responses that seem immediately pleasing over responses that are truthful, accurate, or genuinely helpful. This emerges from RLHF when human raters reward agreeable responses. Sycophantic Claude would: change its answer when a user pushes back even if Claude was right, agree with obviously wrong statements to avoid conflict, and validate poor decisions. This makes Claude less trustworthy and valuable. Anthropic explicitly trains against sycophancy to maintain Claude's integrity as a reliable, honest assistant. Options A, C, and D describe completely different concepts.