A parks department agent processes 400 facility-inspection records. On the first run it fails after roughly 260 records. How should the job be designed so that re-running it is safe?
Select an answer to reveal the explanation.
Short Explanation
Long jobs fail, so design for the second run. Checkpoint each inspection record as it finishes and the re-run opens the file, sees 260 done, and starts at 261.
Full Explanation
Idempotence is the property that makes retry a safe default: running the job twice produces the same result as running it once. For a 400-record inspection pass that dies partway, idempotence is not a nicety—it is the difference between a cheap resume and paying for the whole batch again after every failure.
Writing a checkpoint as each record completes gives the job durable per-unit state, so a re-run reads the checkpoint, skips the roughly 260 finished records, and continues from where it stopped. Because the write happens per record rather than as one batch at the end, a crash preserves everything finished up to that instant, and the checkpoint doubles as a progress view that makes the partial results usable in their own right.
Restarting from the first record every run redoes completed work, scales badly as the record count grows, and risks duplicate output when the writes append rather than overwrite; raising the turn budget lowers the odds of one failure mode while doing nothing about tool errors, interruptions, or rate limits, and leaves the recovery story unchanged; having an operator hand-split the remainder puts a human in the retry loop for bookkeeping the job should own, and manual range-splitting is exactly where off-by-one gaps and duplicates originate.
Exam caveat: a checkpoint written before the record's real side effect completes can mark work done that never landed, so order the write after the effect or make the effect itself idempotent. Operational check: kill the job mid-run, restart it, and confirm the processed total reaches 400 with no record handled twice.