A permit assistant can be probed for security holes (named types only: indirect prompt injection, hidden content in retrieved documents) or for unsafe answers to an ordinary resident question. How should a tester tell those two red-team aims apart?
Select an answer to reveal the explanation.
Short Explanation
Security red teaming targets malicious inputs and named AI-specific risks such as indirect prompt injection or hidden content in retrieved documents. Safety red teaming targets harmful outputs from ordinary use. Those two aims are not the same. Do not invent payloads.
Full Explanation
Distinguish security (malicious inputs / named risks) from safety (ordinary-use harmful outputs). Collapsing the aims, writing a payload, or dropping security red teaming leaves the LO.