Quiz 8 Question 7 of 20

An AI safety team is designing a content moderation system for a major social platform. Claude must classify user posts into: Safe, Borderline, Harmful, Severely-Harmful. The 'Borderline' category is the primary accuracy challenge — posts that are provocative but legal, potentially triggering but not policy-violating. Currently, Claude classifies 8% of posts as Borderline. Human reviewers classify 3% as Borderline, and 5% of Claude's Borderline classifications are later escalated to Harmful. The team wants to calibrate Claude's Borderline classification to match human reviewer standards. Which is the most rigorous calibration approach?

Select an answer to reveal the explanation.

Motivation