Training data poisoning

You can change what a model says without touching a line of its code, by getting to the material it learns from or looks things up in.

Poison the training data and the behaviour is baked in. Poison the documents it retrieves at the time of asking and you influence the answer directly, which for most organisations is the more realistic concern, because that pile of documents is usually theirs and usually much easier to write into than they assume. Anything the model reads is an input, and inputs from places you do not control deserve the same suspicion as any other.

Checked against the primary source.

More on AI security