Human approval needs meaningful information
A confirmation dialogue that does not say what is about to happen is not oversight, it is a button.
"Allow this action?" with no detail produces reflexive approval within a day. If a person is the control, they need to see what the agent intends, on what, and why, in language that lets them disagree. Otherwise you have added a delay to the process and a name to the audit log, without adding any actual judgement.
More on Prompt injection
- Prompt secrecy does not solve prompt injectionSecret rules, open letterbox
- Output encoding still matters around LLMsBlunt it before it lands
- Retrieval content needs trust boundariesMaterial, never orders
- A model can leak data present in its contextIt can say what it can see
- Guardrails are layers, not one magic classifierMind the gaps, plural
- Telling the model to ignore attacks is not a hard boundaryPaint is not a barrier
