A housing-authority casework review prompt originally said "extract the fields accurately." Extraction quality stayed inconsistent across caseworkers' scanned intake forms. Replacing that line with an explicit, itemized rubric -- e.g., "a household-size field is correct only if it matches a number appearing on the form; do not infer from unit type" -- produced measurably fewer disputed extractions. What does this demonstrate?
Select an answer to reveal the explanation.
Short Explanation
Telling a model to “be accurate” is like telling a caseworker to “do good work.” Spell out the rule, that household size must match a number on the form, and both sides grade against one yardstick.
Full Explanation
A quality adjective is not an acceptance criterion. “Accurate” names a goal without saying what would count as meeting it, which leaves both the model and the human reviewer to supply their own threshold. When two parties supply different thresholds, the result is a dispute rather than a detectable error, which is precisely what the housing authority was seeing across scanned intake forms.
An itemized rubric converts a subjective goal into an operational test. A rule such as “a household-size field is correct only if it matches a number appearing on the form; do not infer from unit type” is checkable in seconds by the model as it extracts and by the reviewer as they audit. Disputed extractions fell because caseworkers and Claude were finally grading against one explicit criterion instead of two private ones.
Attributing the gain to length mistakes a correlate for the cause, since a short but specific rule such as “never infer” outperforms a long vague one; treating a JSON Schema as sufficient confuses structure with semantics, because a schema constrains field types and presence but cannot judge whether a value was correctly derived from the source document; and appealing to training exposure cannot explain a controlled before-and-after change that tracks a single prompt edit.
Exam caveat: a rubric improves agreement only on the criteria it states, so any field it omits silently reverts to the model's own interpretation. Operational check: score a held-out set of scanned forms against the rubric with two independent human reviewers, and treat disagreement between them as a missing rubric line rather than as a model failure.