How llm-d routes model inference traffic on Amazon EKS
Explore LLM inference on Kubernetes using Red Hat AI on EKS. Trace requests from the Envoy gateway through the EPP scheduler down to individual vLLM pods.
Explore LLM inference on Kubernetes using Red Hat AI on EKS. Trace requests from the Envoy gateway through the EPP scheduler down to individual vLLM pods.
Optimize LLM deployment with Neural Navigator on Red Hat OpenShift AI, reducing cost overruns and latency spikes.
Confidently deploy LLMs with Red Hat support: Learn how to determine if your model is supported by Red Hat's vLLM community.
Battle-test your AI agents before production. Discover how MiDojo's open source framework red-teams tool calls for real-world security and utility.
See how a team upgraded Red Hat OpenShift AI 3.3.2 three to four times faster with an AI coding assistant, reducing engineering effort by 60%.
Reduce observability costs with Red Hat OpenShift AI summarizer, bridging the interpretation gap for cloud-native architectures.
Explore Kubernetes resources for Red Hat AI Inference on Amazon EKS, enabling intelligent routing for your model serving.
Stop guessing RAG settings. Discover how AutoRAG uses fast evaluation sweeps to optimize chunking and retrieval precision for small LLMs on your data.
Learn how to run isolated Llama 3.1 8B workloads on a single NVIDIA H100 GPU using OpenShift, Kubernetes dynamic resource allocation, and NVIDIA MIG.
Improve model reliability at inference time with its_hub: Learn how to select accurate outputs without retraining.
Optimize GPU efficiency with Red Hat OpenShift AI 3.4's flow control for llm-d, ensuring priority-based request queuing and fairness policies.
Learn behavioral testing for agents, catching non-deterministic failures with golden queries and YAML files.
Learn how to build a distributed RAG pipeline with Ray Data on OpenShift AI for high-performance parsing, embedding, and writing to a vector database.
Learn how to preserve existing model knowledge with Orthogonal Subspace Fine-Tuning (OSFT) for fine-tuning large language models in Training Hub.
Boost AI coding agent performance with AGENTS.md and Agent Skills. Learn how to standardize project context, structure tasks, and share skills across teams.
Dive into the Q2’26 edition of Camel integration quarterly digest, covering the
Boost AI and analytics workloads with Kove:SDM on Red Hat OpenShift, enabling applications to access memory resources beyond local node limits.
Combine OpenShift and OpenShell to improve AI agent security, protecting against data exfiltration and container escapes.
Headed to JavaZone 2026? Visit the Red Hat Developer booth on-site to speak to our expert technologists.
Learn how to parse SEC filings into JSON objects matching a given schema with agentic-graphrag-finance.
Learn how to connect GitHub and incident detection MCP servers to OpenShift Lightspeed. Automate your cluster troubleshooting and GitOps workflows using AI.
Operationalize AI agents with OpenShift and Kubernetes primitives for manageable, cost-optimized AIOps.
Learn how to build an open cloud native architecture for AI agents. This blueprint explains how to improve workload isolation and implement inference routing.
Learn how to set up local agentic AI computer use. Run quantized models like Qwen 3.6 with Hermes to automate desktop tasks on your own terms today.
Learn how OpenShift and OpenShell combine for dual protection of AI coding agents, reducing attack surface and securing your infrastructure.