Guardrails are layers, not one magic classifier
No single filter catches everything, so guardrails work as a series of imperfect layers rather than one reliable check.
Input filtering, output filtering, permission limits, rate limits, human approval on consequential steps, monitoring afterwards. Each is defeatable on its own; together they raise the cost. The failure mode is buying one product, believing the problem is solved, and discovering the gap when something routes around it.
More on Prompt injection
- Prompt secrecy does not solve prompt injectionSecret rules, open letterbox
- Output encoding still matters around LLMsBlunt it before it lands
- Retrieval content needs trust boundariesMaterial, never orders
- A model can leak data present in its contextIt can say what it can see
- Telling the model to ignore attacks is not a hard boundaryPaint is not a barrier
- Human approval needs meaningful informationApprove what, exactly
