Retrieval content needs trust boundaries
The documents a system retrieves are an input, and inputs from places you do not control deserve suspicion.
Many retrieval systems index whatever is in a shared drive, a ticketing system or a wiki, all of which are writable by a wide range of people and sometimes by customers. Anybody who can write into that pile can influence what the model says, and in an agent, what it does. Where the corpus comes from is a security decision, not a data engineering one.
More on Prompt injection
- Prompt secrecy does not solve prompt injectionSecret rules, open letterbox
- Output encoding still matters around LLMsBlunt it before it lands
- A model can leak data present in its contextIt can say what it can see
- Guardrails are layers, not one magic classifierMind the gaps, plural
- Telling the model to ignore attacks is not a hard boundaryPaint is not a barrier
- Human approval needs meaningful informationApprove what, exactly
