How should a city monitor error rates even among high-confidence permit extractions?
Select an answer to reveal the explanation.
Short Explanation
Even the "sure" pile needs a random dipstick. Stratified sampling catches sneaky errors in high-confidence extractions.
Full Explanation
Cities must measure error rates even among high-confidence permit extractions by using stratified random sampling of those high-confidence outputs on an ongoing basis. Confidence is a routing signal, not proof of correctness; silent errors in the "sure" pile still become wrong fees, wrong parcels, or wrong applicant data.
Stratified random sampling works because it continuously estimates residual error where automation is most trusted, across document types and fields, instead of assuming the confident subset is clean. Results feed calibration, thresholds, and review-volume decisions.
Assuming high-confidence outputs never contain errors fails because calibration is imperfect and municipal forms shift over time. Reviewing only the lowest-confidence fields and ignoring the rest fails because it never measures the production path that ships without clerks. Sampling only once at launch and never again fails because models, prompts, scanners, and form templates drift after go-live.
Exam caveat: sampling is for measurement and governance—not a substitute for routing low-confidence or contradictory fields to humans. Operational check: schedule ongoing stratified draws from high-confidence permit extractions, score them against clerk ground truth, and publish segment error rates to decide whether review can shrink.