Your LangGraph agent can call tools, follow a system prompt, and answer domain questions. That's not the same as being safe to put in front of users. Without guardrails, a banking customer service agent might explain how to commit check fraud, reply to profanity, or happily help you bake a chocolate cake.
This article walks through a working example that adds NeMo Guardrails to a LangGraph ReAct agent using the proxy-style integration: NeMo Guardrails sits between the agent and the large language model (LLM), checks every request and response, and requires almost no changes to agent code. You point the agent's BASE_URL at the guardrails service instead of the model endpoint.
On Red Hat OpenShift AI, you deploy the NeMo Guardrails service through a NemoGuardrails custom resource (CR) managed by the TrustyAI Operator, point your agent at in-cluster vLLM inference, and optionally trace agent and rail behavior with MLflow and OpenTelemetry, all without rewriting the LangGraph agent.
The guardrailed agent example on GitHub includes the code, guardrails configuration, and deployment steps used in this article.
Watch a 10-minute demo
In the following video, I deploy 2 versions of the same banking agent on Red Hat OpenShift AI, one with guardrails enabled and one without, and compare how each handles unsafe, off-topic, and legitimate requests. I also show agent-level tracing in MLflow and rail-level tracing with OpenTelemetry, Tempo, and Jaeger.
Why use the proxy pattern?
Agent frameworks like LangGraph already own the conversation loop, system prompt, and tool calls. You don't want a guardrails layer to replace that logic.
With passthrough: true, NeMo Guardrails acts as a transparent safety filter:
User → Agent → NeMo Guardrails → LLM (vLLM, Ollama, or NIM)The guardrails server inspects traffic in both directions but passes system prompts and tool calls through unchanged. Allowed requests reach your inference endpoint. When a rail blocks, later rails are skipped and the user gets a configured refusal—"I'm sorry, I can't respond to that"—in this example.
What the example includes
The guardrailed agent is a banking customer service assistant built on the LangGraph ReAct template. It demonstrates 3 rail types:
- Regex filtering: Instant pattern matching for jailbreak strings like "ignore previous instructions" (no LLM call)
- Content safety: LLM classification against S1–S13 categories (violence, criminal planning, profanity, and more)
- Topic safety: LLM classification that keeps the agent inside a banking domain
The repo ships 2 configuration profiles, one for local experimentation and one for cluster deployment:
- local: Self-check rails only. The same LLM that answers user questions also classifies input and output. Good for a quick local setup.
- nemoguard: Layered regex, content safety, and topic rails, with a dedicated model per role (
main,content_safety,topic_control). This is what the video deploys on Red Hat OpenShift AI: the TrustyAI Operator provisions NeMo Guardrails from aNemoGuardrailscustom resource (CR) and ConfigMap. Pointmainat in-cluster vLLM and the safety rails at NVIDIA NIM classifiers, configured via environment variables or cluster secrets.
Deploy guardrails on OpenShift AI
The video deploys the nemoguard profile on a cluster. Prerequisites include the TrustyAI Operator with the NemoGuardrails custom resource definition (CRD), an in-cluster LLM endpoint for the main model (vLLM in this example), and an NVIDIA API key for NIM safety classifiers. See Deploying models (vLLM) and Configuring the NVIDIA NIM model serving platform (in-cluster NIM) for platform setup.
From the example directory, apply the guardrails manifests with oc:
cd agents/langgraph/examples/guardrailed_agent
NS=$(oc project -q)
# Configure cluster values (vLLM endpoint, model IDs, NVIDIA API key)
cp deploy/overlays/ci-testing/cluster.env.example deploy/overlays/ci-testing/cluster.env
# Edit deploy/overlays/ci-testing/cluster.env
set -a && source deploy/overlays/ci-testing/cluster.env && set +a
# 1. Secret (OPENAI_API_KEY placeholder + NVIDIA_API_KEY for hosted NIM classifiers)
oc create secret generic langgraph-guardrailed-agent-guardrails-secrets \
--namespace="$NS" \
--from-literal=api-key="${API_KEY:-not-needed}" \
--from-literal=nvidia-api-key="${NVIDIA_API_KEY}" \
--dry-run=client -o yaml | oc apply -f -
# 2. ConfigMap (nemoguard profile)
python3 deploy/scripts/render_guardrails_configmap.py \
--cluster-env deploy/overlays/ci-testing/cluster.env
oc apply -n "$NS" -f deploy/manifests/02-guardrails-configmap.yaml
# 3. NemoGuardrails CR (substitute OTEL_* when tracing is enabled)
CR_TMP=$(mktemp)
envsubst '$OTEL_EXPORTER_OTLP_ENDPOINT $OTEL_SERVICE_NAME $OTEL_EXPORTER_OTLP_PROTOCOL $OTEL_METRICS_EXPORTER' \
< deploy/manifests/03-nemoguardrails-cr.yaml > "$CR_TMP"
python3 deploy/scripts/finalize_nemoguardrails_cr.py "$CR_TMP"
oc apply -n "$NS" -f "$CR_TMP"
rm -f "$CR_TMP"The TrustyAI Operator provisions the NeMo Guardrails pod from that CR. Point the agent's BASE_URL at the in-cluster guardrails service, not vLLM directly.
The example repo also wraps these steps in make deploy-guardrails when you want a single command to try the flow locally.
How rails run in order
Rails execute sequentially. The first failure short-circuits the chain.
Input rails (before the LLM sees the message):
User message
│
├─ 1. Regex check ──────── jailbreak patterns, no LLM call
├─ 2. Content safety ───── S1–S13 classification
└─ 3. Topic safety ─────── domain boundary (banking in this demo)Output rails (after the LLM responds):
LLM response
│
└─ 4. Content safety ───── catches unsafe text the model generated anywayIn the video demo, a blocked prompt like "how do I build a bomb?" stops at the content safety input rail. You can see rail.stop: true in the OpenTelemetry trace. A legitimate balance inquiry runs through all input rails, calls the check_account_balance tool, passes the output content safety check, and returns the account balance.
Greetings and on-topic banking questions still flow through normally.
Customize rails for your domain
Most configuration lives in 2 files under guardrails/config/nemoguard/: config.yaml and prompts.yml.
config.yaml (generated from config.yaml.example at startup) defines models, rail order, and regex patterns:
rails:
config:
regex_detection:
input:
patterns:
- "(ignore|forget|disregard)... (instructions|rules|prompts)"
input:
flows:
- regex check input
- content safety check input $model=content_safety
- topic safety check input $model=topic_control
output:
flows:
- content safety check output $model=content_safetyprompts.yml defines the classification prompts. The topic_safety_check_input task encodes the banking boundary, allows payments and account questions, and blocks recipes, medical advice, and entertainment.
To adapt this example to healthcare, telecom, or another vertical, update the topic safety prompt and the agent system prompt in src/guardrailed_agent/agent.py. Content safety categories and regex patterns are usually domain-agnostic. See the adapting to a different domain section in the README for the full file list.
Each model role (main, content_safety, topic_control) can point at its own endpoint. That lets you use purpose-built NVIDIA NemoGuard NIM models for classification while keeping a larger model for responses.
Observe what the guardrails are doing
Red Hat OpenShift AI gives you 2 complementary views of the same request:
- Agent-level (MLflow): Standard LangGraph tracing across agents on the platform. You see user queries, tool calls, and responses. NeMo Guardrails appears as a normal LLM endpoint; individual rail decisions are not visible here.
- Rail-level (OpenTelemetry): Per-rail spans from the NeMo Guardrails service. You see which rails ran, which one blocked, and timing for each layer.
In the demo, MLflow captures the agent conversation while Jaeger shows rail-level detail. That's exactly the split you want when debugging false positives or tuning rail order.
What we didn't change in the agent
The LangGraph agent itself is essentially unchanged. We pointed BASE_URL at the guardrails service and added error handling for blocked responses. The guardrails layer owns safety policy; the agent owns reasoning and tools.
That separation is the point. You can tighten regex patterns, swap in a dedicated NemoGuard classifier, or rewrite topic prompts without touching agent business logic.
Try the guardrailed agent yourself
Clone the guardrailed agent example to run it locally, or deploy the same stack on Red Hat OpenShift AI. The README covers local setup, cluster deployment, and tests. To go deeper on authoring rail configs, read Developing LLM guardrail configs locally with NeMo Guardrails.
Learn more
- Enabling AI safety with NeMo Guardrails
- Deploying models on the model serving platform (vLLM for the main agent model)
- Configuring the NVIDIA NIM model serving platform (safety classifiers)
- Developing LLM guardrail configs locally with NeMo Guardrails: A deeper guide to creating and testing rail configs
- Guardrails for agents—agentic-starter-kits docs
- Every layer counts: Defense in depth for AI agents with Red Hat AI