AI red teaming
Testing an AI system is not like testing ordinary software, because the thing you are testing does not answer the same way twice.
A conventional test that passes has established something. An adversarial prompt that fails today may succeed on the fifth attempt, after a rewording, or after the model is updated. So a single pass proves much less than it would elsewhere, and the useful posture is continuous testing rather than a certificate from a point in time.
Checked against the primary source.
