Quiz 8 Question 12 of 20

A team builds an agent that generates inline code comments. They evaluate it using BLEU score against a reference set of human-written comments, achieving a score of 0.72. When they deploy the agent, developers report the generated comments are technically accurate but unhelpful — they often restate what the code does rather than explaining why. The BLEU score does not capture this quality dimension. What does this reveal about the team's evaluation approach?

Select an answer to reveal the explanation.

Motivation