AI security
Models, agents and the pipelines around them. New surface, mostly old failure modes.
52 sketches
AI agentsIt acts, then you find out
AI data leakageNobody attacked anything
AI guardrailsThe filter argues, the reach does not
AI hallucinationsA confident answer nobody checked
AI identityEvery action needs a name behind it
AI monitoringThree pens lifted off the paper
AI red teamingPassing once is not passing
AI-enabled social engineeringCheck on a line you chose
AI-generated codeWritten faster than it can be read
Agent permissionsBlast radius is what it can reach
Indirect prompt injectionThe instruction nobody typed
Model inversionThe data leaves its shape behind
Model supply chainsNew parts, the same questions
Model theftIt copies, and the original stays
Prompt injectionRead as an order, not as text
Shadow AIThe path around the desk
Training data poisoningNothing inside the machine was touched
An LLM predicts plausible continuations, not verified truthIt only checks the shape
Context is temporary working material, not permanent knowledgeThe board gets wiped
Retrieval adds documents, not guaranteed correctnessThe filter slot is empty
Temperature changes variation, not factualityThe dial only sets the spread
System prompts are instructions, not a security boundaryA sign, with no fence
LLM output is untrusted input downstreamIt comes in round the back
Fine-tuning changes behaviour, not every limitationNew type, same roller
Benchmarks measure selected tasks, not universal capabilityThe score covers this much
Access controls protect the service, not every answerChecked at the door only
AI risk changes when output becomes actionWhen the answer gets a lever
Prompt secrecy does not solve prompt injectionSecret rules, open letterbox
Output encoding still matters around LLMsBlunt it before it lands
Retrieval content needs trust boundariesMaterial, never orders
A model can leak data present in its contextIt can say what it can see
Guardrails are layers, not one magic classifierMind the gaps, plural
Telling the model to ignore attacks is not a hard boundaryPaint is not a barrier
Human approval needs meaningful informationApprove what, exactly
Tool descriptions are part of the control surfaceThe label is the lever
Agent memory can preserve poisoned contextIt stays in the water
Agent identity should be separate from user identityTwo necks, one badge
Delegated agents create authority chainsThe thread stays attached
Agent loops need budgets and stopping conditionsSomething has to say enough
High-impact tools need stronger confirmationMatch the catch to the consequence
Tool results are untrusted inputs tooNo sieve on the back route
Agent observability needs action-level tracesAsk for the receipt
Agent rollback is not always possibleThe winch only reaches so far
A model file is executable trust in another formIt looks like data until you open it
Dataset provenance matters for security and governanceWhere did this batch come from?
Model version changes can be security changesOne plate swapped inside
Third-party AI APIs extend the data boundaryThe fence moves with the call
Evaluation data can leak into trainingIt has already seen the exam
Fine-tuning credentials are production credentialsThe bench feeds the floor
Open models shift responsibility toward the operatorThe engine comes with the engine room
Model registries are release infrastructureThe shelf that ships
AI inventories need models, prompts, tools and dataFour parts, one entry
