What does 'AI alignment' refer to in the context of AI safety research?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Alignment is about making sure AI actually wants what we want it to want — and continues wanting it even as it becomes more capable. It's the deep safety problem Anthropic and others work on.
Full explanation below image
Full Explanation
AI alignment is a core area of AI safety research focused on ensuring that AI systems pursue goals and exhibit behaviors that are beneficial and consistent with human values and intentions — even as the systems become more capable. Misaligned AI could have goals that superficially seem correct but lead to harmful outcomes (Goodhart's Law problem). Anthropic was founded in part to address alignment challenges. Constitutional AI and RLHF are practical alignment techniques. Option A describes hardware optimization. Option C is a literal text formatting interpretation. Option D describes fairness/bias mitigation, which is related but distinct from alignment.