Red teaming AI systems: finding the weak spots on purpose

Imagine a customer service chatbot that has passed weeks of testing: it answers questions about delivery times, returns and invoices correctly. Then someone asks it to ignore its instructions and reveal them, and it does. The chatbot had been tested for everything it was supposed to do. Nobody had tried to make it misbehave, and that deliberate attempt is what red teaming adds.

The US National Institute of Standards and Technology (NIST) defines AI red-teaming as “a structured testing exercise used to probe an AI system to find flaws and vulnerabilities such as inaccurate, harmful, or discriminatory outputs, often in a controlled environment and in collaboration with system developers”.

Three ways in

Entry pointWhat a red team tries
The conversationgetting the model to drop its rules or disclose its system instructions and whatever its connected documents contain
The content it readsplanting instructions in pages, emails or files the system will process (prompt injection)
The tools it can usein agentic systems, making it send messages or change records on an attacker’s behalf

Beyond attacks, a red team also looks for harm that needs no attacker at all: confidently wrong answers in specialist questions, and outputs that disadvantage groups of people.

Who should test

NIST distinguishes red teams made up of members of the general public, of domain experts, of a combination of both, and of humans working together with AI-based tools. Few organisations can field all four. A realistic start is two or three colleagues from different departments, a collection of known attack prompts run automatically, and a shared list in which every finding is recorded well enough to reproduce it.

What the EU AI Act requires

The EU AI Act makes adversarial testing an explicit obligation for providers of general-purpose AI models with systemic risk. Article 55(1)(a) requires them to “perform model evaluation in accordance with standardised protocols and tools reflecting the state of the art, including conducting and documenting adversarial testing of the model with a view to identifying and mitigating systemic risks”.

Providers of high-risk AI systems have a related duty under Article 15(5): where appropriate, their technical solutions must include measures to prevent, detect and respond to attacks with inputs designed to cause the model to make a mistake. An organisation that merely deploys a chatbot on its website is not addressed by either provision, but its customers are affected either way when the bot misbehaves. Which obligations apply to your case is part of our EU AI Act compliance work.

To see what hidden instructions in web pages look like, the Prompt Injection Scanner checks a public URL for attempts to manipulate AI crawlers.

Related terms

TermWhat it means
Prompt injectionText an attacker places where an AI system will read it, written to be obeyed.
Adversarial exampleAn input crafted to make a model produce a wrong or unintended result.
Model Context ProtocolThe protocol that connects AI applications to tools; each connected tool is another route an attack can take.

Sources

Legal status: 4 September 2026. These notes are an editorial summary and not legal advice; they do not replace a lawyer’s review of your particular case.

This page expands an entry from the minoka AI glossary, which covers many more terms in brief.

Frequently Asked Questions

Frequently Asked Questions

Who should do red teaming?

Ideally people who did not build the system: colleagues from other departments, domain experts and, for larger systems, external testers, supported by automated tools that run known attack prompts.

How often should an AI system be red-teamed?

Before launch and again after every significant change, such as a new model, new connected data or new tools, because each change can open weaknesses that earlier tests did not cover.

Which EU AI Act provisions deal with adversarial testing?

Article 55 requires adversarial testing from providers of general-purpose AI models with systemic risk. Article 15 requires providers of high-risk AI systems to include, where appropriate, measures against attacks with manipulated inputs.

Can red teaming be automated?

Partly. Automated tools can run large numbers of attack prompts, but human testers remain important for finding new attack patterns and judging how serious a finding is.

Preferred source

Prefer minoka.de on Google

If you add minoka.de as a preferred source, Google shows articles from this site more often in Top Stories, AI Overviews and AI Mode. One click, a Google account, revocable at any time in your source settings.

Prefer on GoogleOpens the source settings at Google