Retrieval content needs trust boundaries

The documents a system retrieves are an input, and inputs from places you do not control deserve suspicion.

Many retrieval systems index whatever is in a shared drive, a ticketing system or a wiki, all of which are writable by a wide range of people and sometimes by customers. Anybody who can write into that pile can influence what the model says, and in an agent, what it does. Where the corpus comes from is a security decision, not a data engineering one.

More on Prompt injection