System prompts are instructions, not a security boundary
The system prompt is the standing instruction at the top of the conversation: who you are, what you will not do.
It is worth writing carefully and it is not a security control. It sits in the same context window as everything else the model reads, competing with the user's text and with anything the system fetched, and it can be argued with, overwritten or leaked. Treating it as a boundary is like writing "staff only" on a door with no lock. If something genuinely must not happen, it needs enforcing where the action occurs, not requested at the start of a conversation.
More on AI and LLMs
- An LLM predicts plausible continuations, not verified truthIt only checks the shape
- Context is temporary working material, not permanent knowledgeThe board gets wiped
- Retrieval adds documents, not guaranteed correctnessThe filter slot is empty
- Temperature changes variation, not factualityThe dial only sets the spread
- LLM output is untrusted input downstreamIt comes in round the back
- Fine-tuning changes behaviour, not every limitationNew type, same roller
