Red-teaming
Adversarial testing of an AI system by a team simulating attackers or misusers to find failure modes before they become incidents.
Adversarial testing of an AI system by a team simulating attackers or misusers to find failure modes before they become incidents.