An AI model used for automated content moderation on a social media platform has a false positive rate for hate speech detection that is three times higher for posts written in African American Vernacular English (AAVE) compared to Standard American English. What type of bias does this illustrate?
Select an answer to reveal the explanation.
Short Explanation and Infographic
Here's the deal — b is correct because this is linguistic bias (a form of measurement bias): the model was trained predominantly on Standard American English and has not learned to correctly distinguish AAVE linguistic patterns from hate speech markers. This causes systematic over-flagging of legitimate AAVE speech, disproportionately impacting Black users.
Full explanation below image
Full Explanation
B is correct because this is linguistic bias (a form of measurement bias): the model was trained predominantly on Standard American English and has not learned to correctly distinguish AAVE linguistic patterns from hate speech markers. This causes systematic over-flagging of legitimate AAVE speech, disproportionately impacting Black users. This is a documented real-world AI fairness problem. Temporal bias (A) would suggest AAVE emerged recently, which is not accurate. Adversarial bias (C) would require intentional manipulation. Sampling bias from overrepresentation of AAVE-flagged content (D) would not explain why legitimate AAVE is falsely flagged.