Benchmarking AI decision models against traditional guardrails
Compare Jev decision models with Red Hat guardrails to evaluate AI safety performance, cost, and latency for enterprise implementation.
Compare Jev decision models with Red Hat guardrails to evaluate AI safety performance, cost, and latency for enterprise implementation.
Learn how to test AI agent skills with skill-creator and promptfoo to prevent silent misfires and ensure reliable trigger behavior in CI pipelines.
Learn how to secure LLM supply chains by implementing behavioral probing and provenance verification to detect backdoors that evade standard model scanning.
Learn how to use garak, an open source LLM vulnerability scanner developed by NVIDIA, to perform red teaming on AI models.
Implement enterprise-wide automation with Pipeline Failure Analyzer: AI-powered CI/CD failure diagnosis.
Implement structured decision reads with DiffusionGemma on Red Hat AI to achieve low-latency, self-hosted AI for regulated industries.
Learn how to reduce Llama 3.1 8B Instruct model size and improve performance with W8A8 INT8 quantization.
Automate enterprise-wide RAG pipelines with Red Hat OpenShift AI's AutoRAG, speeding up AI optimization for your business.
Manage AI infrastructure costs and governance with Red Hat AI 3.4's Model-as-a-Service.
Learn how to install and configure OpenCode, an open source AI coding assistant, for local development with Ollama, OpenVINO, or Red Hat AI.
Improve large language model inference speed with Speculators 0.6.0's FastMTP-style fine-tuning.
Cut Llama 3.1 8B VRAM by 46% without losing accuracy. Master the mechanics of INT8 W8A8 quantization, SmoothQuant, and GPTQ using llm-compressor.
Learn how to configure admission fair sharing in Red Hat build of Kueue 1.4 on OpenShift to prevent job starvation and balance shared resources cluster quotas.
Learn how P-EAGLE in Speculators v0.6.0 uses parallel drafting to reduce LLM latency. Train and deploy custom draft models with this step-by-step guide.
Learn how to reduce LLM memory footprint and costs with our guide on quantization schemes for inference engines
Headed to Devoxx Belgium 2026? Visit the Red Hat Developer booth on-site to speak to our expert technologists.
Learn how to fine-tune large language models on Red Hat OpenShift AI with Ray and Training Hub for optimized SQL generation using the LoRA algorithm.
Learn how to optimize self-hosted LLM cost per token. Cut GPU spending and maximize real-world throughput with autoscaling, right-sizing, and vLLM tuning.
Learn how to “crash test” an AI model before it hits production by probing it the way real users and attackers will. You’ll see what can go wrong when a model ships without safety training, then walk through running NVIDIA’s open source Garak scanner locally against GPT-2. The episode breaks down how Garark works and the next steps for production readiness.
Explore LLM inference on Kubernetes using Red Hat AI on EKS. Trace requests from the Envoy gateway through the EPP scheduler down to individual vLLM pods.
Battle-test your AI agents before production. Discover how MiDojo's open source framework red-teams tool calls for real-world security and utility.
Explore Kubernetes resources for Red Hat AI Inference on Amazon EKS, enabling intelligent routing for your model serving.
Stop guessing RAG settings. Discover how AutoRAG uses fast evaluation sweeps to optimize chunking and retrieval precision for small LLMs on your data.
Improve model reliability at inference time with its_hub: Learn how to select accurate outputs without retraining.
Learn behavioral testing for agents, catching non-deterministic failures with golden queries and YAML files.