A 311 request classifier returns structured JSON with a category field and a confidence field (0-1). The team adds a rule: any result with confidence < 0.6 is automatically routed to a human dispatcher instead of auto-filed. What is this design pattern primarily addressing?
Select an answer to reveal the explanation.
Short Explanation
It is a triage nurse flagging “not sure, get a doctor to look” instead of guessing every case. The model's own uncertainty picks which 311 classifications a dispatcher sees, so ambiguous requests do not get quietly mis-filed.
Full Explanation
Not every automated answer deserves the same trust, and a system that treats them identically spends its scarce human attention at random. Emitting a confidence value alongside the classification and wiring a threshold on top of it makes the system's own uncertainty the thing that decides where a person looks.
This is confidence calibration paired with escalation. The confidence field is part of the structured output, so the routing rule is a deterministic downstream check rather than another model judgment, and the 0.6 threshold concentrates dispatcher review on the requests most likely to be wrong instead of on a random sample or on the entire queue.
Reading it as tool distribution confuses a workflow routing decision with tool selection, since nothing about a confidence threshold implies a different MCP server; invoking Message Batches cost optimization misplaces the concern entirely, because batching is about asynchronous bulk-processing economics rather than accuracy-based routing; and attributing it to the lost-in-the-middle effect misapplies a long-context attention phenomenon to a pattern with nothing to do with document length, since a short complaint can produce a low-confidence classification just as easily as a long one.
Exam caveat: self-reported confidence is only useful once it has been shown to be calibrated, because a model that says 0.9 on work it gets wrong a third of the time makes the threshold theater. Operational check: hold out a labelled sample of resident requests, measure accuracy within each confidence band, and set the threshold from that curve rather than from a round number.