Guardrails are layers, not one magic classifier

No single filter catches everything, so guardrails work as a series of imperfect layers rather than one reliable check.

Input filtering, output filtering, permission limits, rate limits, human approval on consequential steps, monitoring afterwards. Each is defeatable on its own; together they raise the cost. The failure mode is buying one product, believing the problem is solved, and discovering the gap when something routes around it.

More on Prompt injection