A model can leak data present in its context
Anything placed in the context window can come back out.
If a system puts one customer's record into the prompt to answer a question, that record is available to the model for the rest of that exchange, and models can be talked into repeating things. This is the mechanism behind a whole class of incidents where users see somebody else's data: nothing was breached, the data was simply handed to a system that then said it out loud.
More on Prompt injection
- Prompt secrecy does not solve prompt injectionSecret rules, open letterbox
- Output encoding still matters around LLMsBlunt it before it lands
- Retrieval content needs trust boundariesMaterial, never orders
- Guardrails are layers, not one magic classifierMind the gaps, plural
- Telling the model to ignore attacks is not a hard boundaryPaint is not a barrier
- Human approval needs meaningful informationApprove what, exactly
