Prompt secrecy does not solve prompt injection
Hiding the system prompt does not stop anybody attacking it.
System prompts leak routinely, and injection does not require knowing what the instructions say in the first place. Somebody does not have to read your prompt to write text that overrides it. Secrecy is at best a small obstacle, and relying on it means the actual defences, permissions and boundaries, never get built.
More on Prompt injection
- Output encoding still matters around LLMsBlunt it before it lands
- Retrieval content needs trust boundariesMaterial, never orders
- A model can leak data present in its contextIt can say what it can see
- Guardrails are layers, not one magic classifierMind the gaps, plural
- Telling the model to ignore attacks is not a hard boundaryPaint is not a barrier
- Human approval needs meaningful informationApprove what, exactly
