A public-health agent has access to a delete_patient_record tool that permanently removes a record from the system. What tool design pattern is most appropriate to prevent an accidental or misinterpreted call from causing irreversible harm?
Select an answer to reveal the explanation.
Short Explanation
You do not wire the button that permanently shreds patient records to fire the instant someone glances at it. You add a second key that only a person can turn.
Full Explanation
Safeguards should be sized to reversibility. Permanently deleting a patient record cannot be undone by a later correction, so the design question is not how to detect a mistake afterward but how to make a single agent decision insufficient to cause one.
Gating the destructive call behind an explicit human-approval step—a confirm parameter satisfied by a human-in-the-loop check—inserts a structural pause between intent and effect. The reviewing person carries case context the model may lack, and because the pause lives in the execution path itself it applies whether the call arose from a sound plan or a misread instruction.
A longer, more detailed description may improve comprehension but remains advisory text that cannot stop a call once issued; logging the deletion for later supervisor review is purely reactive, since the record is already gone by the time anyone reads the log; removing the description makes the tool harder to use correctly when deletion is legitimate, obscuring capability without constraining it.
Exam caveat: the gate belongs on the execution boundary or a hook wrapping it, not in prompt language—a confirm flag the model can set for itself is not human approval. Operational check: issue a delete call in a test environment and verify nothing is removed until a separate human action is recorded, and that the approval record names who approved which record.