← Research papers
2025arXivunread

Red Teaming AI Red Teaming

Subhabrata MajumdarBrian PendletonAbhishek Gupta
Publisher pagePDF
Open graph

Citations

0

Open access

No

Source

arxiv

OpenAlex

Not enriched

arXiv

2507.05538

Abstract

Red teaming has evolved from its origins in military applications to become a widely adopted methodology in cybersecurity and AI. In this paper, we take a critical look at the practice of AI red teaming. We argue that despite its current popularity in AI governance, there exists a significant gap between red teaming's original intent as a critical thinking exercise and its narrow focus on discovering model-level flaws in the context of generative AI. Current AI red teaming efforts focus predominantly on individual model vulnerabilities while overlooking the broader sociotechnical systems and emergent behaviors that arise from complex interactions between models, users, and environments. To address this deficiency, we propose a comprehensive framework operationalizing red teaming in AI systems at two levels: macro-level system red teaming spanning the entire AI development lifecycle, and micro-level model red teaming. Drawing on cybersecurity experience and systems theory, we further propose a set of six recommendations. In these, we emphasize that effective AI red teaming requires multifunctional teams that examine emergent risks, systemic vulnerabilities, and the interplay between technical and social factors.

Collections

Add to collection

Paper intelligence

Research analysis

Confidence 85%

20 source chunks

Summary

Red teaming has evolved from its origins in military applications to become a widely adopted methodology in cybersecurity and AI. In this paper, we take a critical look at the practice of AI red teaming. We argue that despite its current popularity in AI governance, there exists a significant gap between red teaming's original intent as a critical thinking exercise and its narrow focus on discovering model-level flaws in the context of generative AI. Current AI red teaming efforts focus predominantly on individual model vulnerabilities while overlooking the broader sociotechnical systems and emergent behaviors that arise from complex interactions between models, users, and environments. To address this deficiency, we propose a comprehensive framework operationalizing red teaming in AI systems at two levels: macro-level system red teaming spanning the entire AI development lifecycle, and micro-level model red teaming.

Plain-language summary

Red teaming has evolved from its origins in military applications to become a widely adopted methodology in cybersecurity and AI. In this paper, we take a critical look at the practice of AI red teaming. We argue that despite its current popularity in AI governance, there exists a significant gap between red teaming's original intent as a critical thinking exercise and its narrow focus on discovering model-level flaws in the context of generative AI. Current AI red teaming efforts focus predominantly on individual model vulnerabilities while overlooking the broader sociotechnical systems and emergent behaviors that arise from complex interactions between models, users, and environments. To address this deficiency, we propose a comprehensive framework operationalizing red teaming in AI systems at two levels: macro-level system red teaming spanning the entire AI development lifecycle, and micro-level model red teaming.

Research problem

We argue that despite its current popularity in AI governance, there exists a significant gap between red teaming's original intent as a critical thinking exercise and its narrow focus on discovering model-level flaws in the context of generative AI. We argue that despite its current popularity in AI governance, there exists a significant gap between red teaming’s original intent as a critical thinking exercise and its narrow focus on discovering model-level flaws in the context of generative AI. While that is not necessarily a bad thing—because red teaming is definitely a part of solving the AI security and safety problem—it is only a small part of a more complete set of tools that should be used.

Methodology

Red teaming has evolved from its origins in military applications to become a widely adopted methodology in cybersecurity and AI. We argue that despite its current popularity in AI governance, there exists a significant gap between red teaming's original intent as a critical thinking exercise and its narrow focus on discovering model-level flaws in the context of generative AI. Current AI red teaming efforts focus predominantly on individual model vulnerabilities while overlooking the broader sociotechnical systems and emergent behaviors that arise from complex interactions between models, users, and environments. To address this deficiency, we propose a comprehensive framework operationalizing red teaming in AI systems at two levels: macro-level system red teaming spanning the entire AI development lifecycle, and micro-level model red teaming. Rudd Abstract Red teaming has evolved from its origins in military applications to become a widely adopted methodology in cybersecurity and AI.

Main findings

We argue that despite its current popularity in AI governance, there exists a significant gap between red teaming's original intent as a critical thinking exercise and its narrow focus on discovering model-level flaws in the context of generative AI. We argue that despite its current popularity in AI governance, there exists a significant gap between red teaming’s original intent as a critical thinking exercise and its narrow focus on discovering model-level flaws in the context of generative AI. Introduction In recent past, the term red teaming has gained significant attention in diverse conversations around AI as a potential solution to find and address security, safety, and reliability concerns in generative AI (genAI) systems. This missing link negates the tremendous benefits that are not realized in bridging communities of practice in AI and cybersecurity, which could ultimately lead us to finding ways to achieve robust and effective AI governance and AI solutions that adhere more closely to the espoused principles of responsible AI. It should bring together multi- functional and cross-functional teams around the common goal of ultimately improving a product or user experience.

Key contributions

Contribution 1

To address this deficiency, we propose a comprehensive framework operationalizing red teaming in AI systems at two levels: macro-level system red teaming spanning the entire AI development lifecycle, and micro-level model red teaming.

Contribution 2

To address this defi- ciency, we propose a comprehensive framework operationalizing red teaming in AI systems at two levels: macro-level system red teaming spanning the entire AI development lifecy- cle, and micro-level model red teaming.

Contribution 3

Adoption in Cybersecurity The National Security Agency (NSA) first recognized the need for proactive cybersecurity measures in the 1980s, and pioneered the concept of “red teams” tasked with assessing the security of classified systems (TechRound, 2023).

Contribution 4

We propose that red teaming of AI systems be operationalized at two complementary levels: macro-level (or system) red teaming that spans the entire AI development lifecycle, and micro-level (or model) red teaming that focuses on the model powering the AI system.

Contribution 5

Macro-level (System) Red Teaming Just like technical debt in ML systems can arise from system components other than code (Sculley et al., 2015), ML (and AI) failures can stem from decisions made long before the first line of code is written (Figure 1).

Contribution 6

Inception The inception stage is critical, given this is where stakeholders first envision an AI solution to address a particular challenge.

Contribution 7

Maintenance Red teams must evaluate monitoring and alerting systems as the first line of defense against diverse failure modes.

Limitations

  • As the field matures, researchers increasingly recognize that red teaming alone cannot solve all the challenges in AI risk assessment (Ahmad et al., 2025).
  • Critically, they should help everyone be on the same page to answer the question “What does good look like?” To this end, stakeholders must articulate not just what they want the system to do, but what they absolutely cannot allow it to do.
  • Data distribution shifts over time can adversely impact system performance in ways that may not be immediately apparent.
  • Secondly, the integration of cybersecurity red teaming practices with AI-specific concerns assumes transferability that may not hold, given AI systems’ probabilistic behav- iors and sociotechnical complexities.

Future work

  • Future work should propose frameworks that can map cross-component interactions, track vulnerabilities across system boundaries, and automatically identify emergent risks arising from component interactions.

Evaluation metrics

accuracy

Keywords

artificial intelligencemachine learninglarge language modelAI safetyexplainabilityrisk managementhealthcare AIprivacyfairnessrobustness

No graph connections yet.

Sync citations or add papers to shared collections to build this network.

Knowledge graph

Citation network

Explore references, papers that cite this work and related papers in your Codex library.

References

0

No references have been linked yet.

Cited by

0

No saved paper is currently linked as citing this work.

Related papers

0

Add papers to shared collections or enrich their topics to find related work.

Research workspace

Attach the paper PDF, extract its text, classify its contents and create semantic embeddings.

Attach PDF

Upload the research paper so Codex can extract, chunk and search its contents.

Paper resources

Red Teaming AI Red Teaming.pdf

Status: completed20 chunks60604 charactersapplication/pdf