How does Constitutional AI relate to RLHF in Anthropic's approach?
Select an answer to reveal the explanation.
Short Explanation and Infographic
CAI is RLHF with a written constitution — principles guide self-critique so you need fewer human preference labels for safety.
Full explanation below image
Full Explanation
Anthropic's Constitutional AI trains models to critique and revise outputs against explicit principles, complementing or reducing pure human preference labeling used in classic RLHF. It does not eliminate all human involvement, is not mere marketing synonymy, and is not image-only.