A public-health benefits research agent produces a summary of eligibility rules and lists several supporting claims, some pulled confidently from a verified state statute and others inferred loosely from an outdated blog post. What should the agent's output preserve so a human reviewer can calibrate trust appropriately?
Select an answer to reveal the explanation.
Short Explanation
Footnote each line rather than stamping the whole memo. A reviewer needs to see which sentence came from the statute and which came from a stale blog post.
Full Explanation
Trust in an agent-produced summary is not a single quantity. A public-health eligibility brief assembled from a verified state statute and an outdated blog post contains claims of genuinely different reliability, and any representation that collapses them into one number or one confident voice destroys the information a reviewer needs most.
Attaching provenance and a confidence indicator to each individual claim keeps reliability at the granularity where it actually varies. The reviewer can accept the statute-backed eligibility thresholds, flag the blog-derived inference for verification, and act on the well-supported parts without waiting on the shaky ones. Per-claim attribution also makes the brief auditable later, when someone asks where a number came from.
A single overall confidence score averages strong and weak claims into a figure that describes neither, forcing uniform trust or uniform doubt; dropping source attribution for smoother prose removes the very handle a reviewer uses to verify anything, trading auditability for style; asserting equal confidence for every claim actively misstates the evidence, which is worse than silence because it reads as a considered judgment.
Exam caveat: provenance is traceability, not authority—a citation to a weak source is still provenance, and the confidence indicator is what tells the reviewer the source is weak. Operational check: sample claims from the finished brief and confirm each resolves to a named source and a stated confidence, and that at least one low-confidence claim is visibly marked rather than quietly upgraded.