Distributed training on OpenShift AI 3.4 with Kubeflow Trainer v2
Learn how to simplify distributed training with Kubeflow Trainer 2 on OpenShift AI by replacing framework-specific CRDs with a unified TrainJob resource.
Learn how to simplify distributed training with Kubeflow Trainer 2 on OpenShift AI by replacing framework-specific CRDs with a unified TrainJob resource.
Learn how to use garak, an open source LLM vulnerability scanner developed by NVIDIA, to perform red teaming on AI models.
Implement enterprise-wide automation with Pipeline Failure Analyzer: AI-powered CI/CD failure diagnosis.
Stop relying on accuracy for fraud detection. Build a custom-scored Python model sweep, automate LLM triage, and scale using Red Hat OpenShift AI.
Implement structured decision reads with DiffusionGemma on Red Hat AI to achieve low-latency, self-hosted AI for regulated industries.
Optimize LLM inference performance with a 71% latency reduction on NVIDIA H200 GPUs. Learn how to cut TTFT from 995ms to 287ms with smarter request routing.
Integrate Red Hat Satellite's MCP server with OpenShift Lightspeed for a single AI assistant with live infrastructure data and enterprise knowledge.
Add NeMo Guardrails to a LangGraph agent on Red Hat OpenShift AI without rewriting the graph. Jump to 5:28 to see the same in-cluster vLLM unguarded vs guarded: the unguarded agent answers a check-fraud prompt; NeMo Guardrails blocks it on content safety.
NeMo Guardrails is included in Red Hat OpenShift AI (upstream: NVIDIA NeMo Guardrails) and is deployed with a NeMo Guardrails custom resource managed by the TrustyAI Operator. This walkthrough covers the passthrough proxy, layered rails (regex → content safety → topic boundary → output safety), and opt-in tracing so you can see which rail fired.
Chapters
0:00 Introduction
0:17 What this agent does
1:02 Proxy architecture and rail order
2:25 Customizing rails and changing domain
4:23 Deploy locally and on OpenShift AI
4:58 Tracing overview
5:28 Unguarded vs guarded
8:04 Which rail fired
9:35 Wrap-up
Resources
Guardrailed Agent example:
https://github.com/red-hat-data-services/agentic-starter-kits/tree/main/agents/langgraph/examples/guardrailed_agent
Enable AI safety with NeMo Guardrails:
https://docs.redhat.com/en/documentation/red_hat_openshift_ai_self-managed/3.5/html/enabling_ai_safety_with_guardrails/enabling-ai-safety-with-nemo-guardrails_nemo-guardrails
Track AI inference costs by department using Red Hat OpenShift AI’s built-in MaaS gateway API keys, Perses dashboards, and MLflow tracing, with no manual instrumentation required.
Learn how to reduce Llama 3.1 8B Instruct model size and improve performance with W8A8 INT8 quantization.
Automate enterprise-wide RAG pipelines with Red Hat OpenShift AI's AutoRAG, speeding up AI optimization for your business.
Learn how to install and configure OpenCode, an open source AI coding assistant, for local development with Ollama, OpenVINO, or Red Hat AI.
Learn how to fine-tune your text-to-speech model for Turkish with Red Hat OpenShift AI and Kubeflow Trainer, reducing speech errors by over 90%.
Improve large language model inference speed with Speculators 0.6.0's FastMTP-style fine-tuning.
Cut Llama 3.1 8B VRAM by 46% without losing accuracy. Master the mechanics of INT8 W8A8 quantization, SmoothQuant, and GPTQ using llm-compressor.
Simplify Red Hat Developer Hub software template authoring with rhdh-templates, an agent skill for best practices and local validation.
Learn how to implement AI-driven AIOps workflows with Red Hat Ansible Automation Platform for efficient incident resolution
Automate firewall changes with Ansible Automation Platform, ServiceNow, and AI. Learn to build a governed, event-driven workflow with human approvals.
Learn how to configure admission fair sharing in Red Hat build of Kueue 1.4 on OpenShift to prevent job starvation and balance shared resources cluster quotas.
Learn how P-EAGLE in Speculators v0.6.0 uses parallel drafting to reduce LLM latency. Train and deploy custom draft models with this step-by-step guide.
Evaluate LLM guardrails with EvalHub: Improve AI safety with open-source platform
Evaluate AI agents on Red Hat OpenShift AI with IBM CLEAR and EvalHub. Learn how to automate analysis for recurring failure patterns in AI agent execution.
Learn how the vLLM LoRA dynamic loading flaw enables data theft and how to enforce zero trust defenses using Red Hat OpenShift AI and cluster management.
Learn how to reduce LLM memory footprint and costs with our guide on quantization schemes for inference engines
Learn how to automate RAG document processing with Red Hat OpenShift AI for production-ready parsing, embedding, and querying of PDF documents.