Tool results are untrusted inputs too
Everything an agent fetches is input from somewhere else, including the output of its own tools.
A search result, an API response, the contents of a document it just read. Any of those can contain text that reads like an instruction, and the model has no dependable way to tell the difference. Treating tool results as trusted because they came from your own plumbing is how indirect prompt injection gets in through the back.
More on AI agents
- Tool descriptions are part of the control surfaceThe label is the lever
- Agent memory can preserve poisoned contextIt stays in the water
- Agent identity should be separate from user identityTwo necks, one badge
- Delegated agents create authority chainsThe thread stays attached
- Agent loops need budgets and stopping conditionsSomething has to say enough
- High-impact tools need stronger confirmationMatch the catch to the consequence
