A water utility's outage assistant repeats the same 300-word block — its persona, tone rules, and escalation policy — at the top of every user message in a long conversation. Where does this content belong in the Messages API?
Select an answer to reveal the explanation.
Short Explanation
The system prompt is the job description you hand someone on day one, not a note you re-staple to every ticket. Persona and escalation policy belong there, stated once.
Full Explanation
The Messages API separates the system prompt from the messages array so that role, tone, and standing policy can be stated once and apply across an entire request. A 300-word block repeated at the top of every user turn is that separation being ignored, and the cost compounds with every turn of a long outage conversation.
The top-level system parameter carries persona, tone rules, and escalation policy for the whole request without restatement. Each user turn then contains only the resident's actual question, standing rules stay distinct from conversational content, and a stable system block makes an ideal prompt-caching prefix, so the utility stops paying full price for the same words on every turn.
Putting the rules in the first user message alone buries policy inside conversational content, where it competes with the resident's question and weakens as the transcript grows; placing them in an assistant message misrepresents authored policy as something the model already said, a weaker channel for standing instructions; and a tool definition's description scopes one tool's use rather than conversation-wide persona.
Exam caveat: a longer system prompt is not automatically better, because standing rules still compete for attention, and a constraint that genuinely applies to one turn belongs in that turn. Operational check: move the block into system, replay a long outage transcript, and compare both input token counts and adherence to the escalation policy on the final turns against the old arrangement.