Breadcrumb

  1. Red Hat Interactive Learning Portal
  2. OpenShift AI learning
  3. Red team an AI model with NVIDIA garak
  4. Understanding AI red teaming and garak

Red team an AI model with NVIDIA garak

Learn how to use the open-source scanner garak to test AI models for vulnerabilities. This step-by-step tutorial guides you through installing garak, targeting a model like Granite 3.1 8B, running a prompt injection probe, and analyzing the results to determine if further safety guardrails are required.

Generative AI models introduce non-deterministic behaviors, safety vulnerabilities, and attack surfaces that traditional software testing cannot catch. In this lesson, you will learn the foundational concepts of AI red teaming, explore how NVIDIA's open source garak scanner automates vulnerability probing, and see how to catch security flaws before models reach production.

Prerequisites:

In this lesson, you will:

  • Understand what AI red teaming is and why it differs from traditional testing. 
  • Learn the unique security challenges AI models present.
  • Explore garak’s architecture and how it finds vulnerabilities.
  • Discover when to use garak versus other security testing approaches.

What is Red Teaming?

Red teaming is a structured, adversarial security exercise where you deliberately try to break or exploit a target to find weaknesses in your applications, systems, or AI models before real attackers do.  It’s a practice that is borrowed from military and cybersecurity, assembling a team whose entire job is to think like the enemy and attack your own defenses, answering the question, "how does it break?". 

In the context of AI, red teaming means to systematically test a model with adversarial inputs, ones that attackers, or even a well-meaning but creative user, might send into production. 

One common approach is automated probing. Probing is when a library of pre-made attacks (probes) is run against a target LLM endpoint and the responses are parsed to see which probes caused the model to fail. This can include:

  • Hallucinations: when an AI system produces outputs that seem realistic but are factually incorrect, irrelevant, or made up.
  • Prompt injection: a type of cyberattack where malicious inputs are disguised as legitimate prompts, causing generative AI systems to leak confidential data, spread misinformation, or execute unauthorized actions. 
  • Toxic output: an instance of an AI model producing hateful, abusive, or profane or obscene content.
  • Jailbreaks: when vulnerabilities are exploited to bypass ethical guidelines, including prompt injection attacks.

Red teaming can be done manually, with human testers crafting adversarial prompts by hand, or be automated with tools that scale the process with hundreds of known attack patterns. Both are useful, since human testers come up with creative, context-specific attacks, and automated red teaming covers a large array of known vulnerabilities.

Why Red Team AI models?

Compared to traditional software, AI models present unique security challenges. This includes:

  • Unpredictable behaviors: Unpredictable and non-deterministic behaviors means that the same system can produce different outputs even with the exact same input. This results in insufficient traditional testing to catch all potential failures.
  • Hidden safety and bias risks: Proactive testing helps find deep-seated bias, toxic responses, and hallucinations before users interact with the system in production.
  • Security vulnerabilities: AI systems are vulnerable to techniques like the ones listed above, so red teaming helps find the weaknesses before attackers exploit them. 

What is garak?

Garak is an open source LLM vulnerability scanner created by NVIDIA and built to find security flaws and behavioral risks in AI systems. It probes the model for prompt injection, jailbreaks, data leakage, toxicity, and hallucinations. 

Garak is regarded as one of the easiest ways to find whether or not a model is able to defend against known attacks; it is not intended to find new or novel attacks. Garak also has CI-friendly output, meaning that you can seamlessly add the output to an automated pipeline.

How does garak work?

Garak uses a probe-detector architecture:

  • Generators: connect garak to AI model’s API.
  • Probes: pre-built attack patterns to test specific vulnerabilities. 
  • Detectors: analyze the responses to see if the attack succeeded. 
  • Harnesses: orchestrate testing process.
  • Evaluator: generates results with pass/fail outcomes.

In a run, garak loads the target model, runs probes through the harness, checks the responses with detectors, then compiles everything into a report.

When is garak used?

Garak is best for finding situations where a model may respond with unwanted outputs. It is not meant to find whether or not a model is safe, meaning that it is not meant to assess for social biases or how likely a system is to produce toxic output. For example, garak doesn’t check if a chatbot occasionally uses offensive language when answering normal questions. It checks whether an attacker can bypass content filters by disguising their malicious inputs through prompt injection or exploiting the model’s completion behavior. 

Before we begin with this learning path’s steps, it’s important to first understand crash testing AI models, including what we hope to accomplish with red teaming. In this lesson, we’ve included a video that will introduce you to these key concepts.

Crash test model deployments with NVIDIA garak

Learn how to crash test an AI model before it hits production by probing it the way real users and attackers will. You’ll see what can go wrong when a model ships without safety training, then walk through running NVIDIA’s open source garak scanner locally against GPT-2. This video breaks down how garak works and the next steps for production readiness.

Please see transcript below: 

0:06 - Why AI models need crash testing

You wouldn’t buy a car with a horrible crash test rating, much less one that has never been crash tested at all. You want to know how it handles a collision before you’re the one in the driver’s seat. 

Your AI models deserve the same treatment. Before they hit production, you need to know what happens when someone tries to break them. Today, I’m going to show you how to crash test an AI model using garak. 

0:27 - The risks of untested AI models

First, let’s see what happens when a model goes to production without any testing. I just sent a jailbreak prompt to GPT-2, a model with no safety fine-tuning, and it did exactly what I asked. No pushback, no refusal; it just complied. A model with proper safety training should refuse this outright. GPT-2 doesn’t even try. 

This is one of the 4 ways that an untested model can fail: hallucination, prompt injection, jailbreaks and toxic output. But the point is the same for all four. Any one of these is cheaper to catch before launch than after. 

1:04 - Crash test with NVIDIA garak

So, how do you catch these failures before they reach production? You crash-test the model. 

Garak is an open source LLM vulnerability scanner built by NVIDIA. It runs adversarial probes against your model, the same kinds of attacks a real user or attacker would try, and reports back what broke. 

I’m pointing garak at GPT-2 on Hugging Face and telling it to run DAN 11.0. It’s a well-known jailbreak technique that tries to trick the model into ignoring its safety guidelines. DAN stands for “Do Anything Now”. It’s a long, elaborate prompt that basically tells the model to pretend it has no rules. 

You can see it downloading and loading the model locally. This is all running on my machine. Nothing is being sent to an external API. 

Now it’s queuing up to probe and sending the attack prompts. Garak sends the same prompt multiple times, five by default, to see how consistently the model responds. That matters because a model might refuse once and comply the next time with the exact same input. 

You don’t want to know if it can fail… you want to know how often it fails.

Once all five attempts come back, garak runs detectors against each response. Think of probes as the attacks and detectors as the judges. The probes throw the punches and the detectors score whether the model stayed standing or went down. 

Each detector is looking for something specific, like whether the model adopted a jailbreak persona or whether it failed to refuse a harmful request at all. And DAN 11.0 is just one probe. Garak has probes for prompt injection, toxic output, Personally Identifiable Information (PII) leakage, and hallucination. You can run as many as you need to get a full picture of where your model is vulnerable. 

2:41 - Analyzing scan results and detector scoring

And here are the results. Two detectors ran against the model’s responses. The first one, the DAN detector, checks whether the model actually adopted the jailbreak persona, whether it started responding as if it had no rules. Two out of the five responses were clean. Three weren’t. That’s a 60% attack success rate. So, three out of five times the model played along with the jailbreak. 

The second [detector], mitigation bypass, is asking a different question. It’s not checking whether the model adopted the persona. It’s checking whether the model ever pushed back at all. Did it ever say that it can’t do that or give any kind of refusal? Zero out of five; not once! The model never even attempted to refuse the request. That’s a 100% attack success rate, and garak flags that as an immediate risk. 

So to put that together, the model went along with the jailbreak persona more often than not. And it always failed to refuse the request in the first place. It didn’t even recognize that it should say no. That’s the difference between a model that gets tricked and a model that has no defenses to begin with. 

3:48 - Pipeline automation with EvalHub

Garak also has an HTML report that you can share with your team for a more visual look at their results. That’s why garak gets wired into your continuous integration and continuous deployment (CI/CD) pipeline through EvalHub, the evaluation orchestration service for models on Red Hat OpenShift AI. 

You define a benchmark collection that includes garak as a provider. Set your pass/fail threshold and call EvalHub’s post/evaluations endpoint from your pipeline. Garak runs as a Kubernetes job against your model’s live endpoint, right alongside any other benchmarks in the same collection instead of as a separate, manual step. 

Run that scan as a gate in your deployment pipeline, not as a one-time launch task. Every time that a model change, prompt change, or retrieval source change is about to go into production, the pipeline runs garak first. If a small tweak quietly breaks something that was previously safe, you find out in CI, not from a customer. 

4:41 - Runtime protection with NeMo Guardrails

The second layer is runtime protection. There’s a reason safety reports don’t test that your SUV would avoid a UFO falling from above. Some incidents can’t be predicted. Testing before deployment catches known failure modes, but it doesn’t stop something new from happening in production. 

That’s where the NeMo Guardrails orchestrator comes in. It’s built on the open source NVIDIA NeMo Guardrails project and included with Red Hat OpenShift AI. It sits in front of your deployed model and screens inputs and outputs as they pass through. Using detectors, you can configure and tune for a use case without retraining the model itself. It exposes endpoints that can validate a message against your configured rails without even generating a response, or run input rails on the incoming message, generate the response, and check it through output rails before it ever reaches the user. 

Pre-deployment scanning with garak catches what you already know to test for. The NeMo Guardrails orchestrator catches what happens live. 

5:35 - Summary: three layers of safety

So here’s the full picture. Scan for known attack patterns with garak. Gate every production push on the result through EvalHub and add the NeMo Guardrails orchestrator as a runtime layer that watches live traffic. Not because every car will crash, but because you’d rather know the rating before you’re the one behind the wheel. 

Testing before deployment, automated scanning in a pipeline, and guardrails watching traffic: three layers, each one catching what the others can’t. And it all starts with that first scan, the one you just saw. 

In the next lesson, you’ll run your first garak scan!

Previous resource
Overview: Red team an AI model with NVIDIA garak
Next resource
Run your first garak scan