LLM output is untrusted input downstream
Whatever a model produces should be treated as user input by whatever consumes it next.
If the output goes into a web page, it can carry a script. Into a database query, it can carry injection. Into a shell command, worse. The usual mistake is treating model output as trusted because it came from your own system, when in fact it was shaped by whatever the model read, which may well have been written by somebody else.
More on AI and LLMs
- An LLM predicts plausible continuations, not verified truthIt only checks the shape
- Context is temporary working material, not permanent knowledgeThe board gets wiped
- Retrieval adds documents, not guaranteed correctnessThe filter slot is empty
- Temperature changes variation, not factualityThe dial only sets the spread
- System prompts are instructions, not a security boundaryA sign, with no fence
- Fine-tuning changes behaviour, not every limitationNew type, same roller
