Definition

AI red teaming uses adversarial testing — often automated models in self-play against defender models — to discover prompt injection, jailbreaks, and agent exploits before deployment, then fold findings into training and safeguards.

Key Points

Sources