Implement GPU-as-a-Service with Kueue and NVIDIA MIG
Learn how to implement GPU-as-a-Service on Red Hat OpenShift using Kueue, NVIDIA MIG, and a custom dashboard plug-in for self-service GPU resource booking.
Learn how to implement GPU-as-a-Service on Red Hat OpenShift using Kueue, NVIDIA MIG, and a custom dashboard plug-in for self-service GPU resource booking.
Explore the Data Governance Copilot architecture, integrating OpenShift AI with PG Airman MCP server for robust, agentic Postgres analytics.
Learn how to connect a modern Apache Iceberg lakehouse to LLM-hosted models using nothing but SQL on Red Hat OpenShift AI.
Learn how to deploy MemPalace as a production-ready Model Context Protocol server on OpenShift AI with HTTP/WebSocket transport and Kubernetes health probes.
Learn how to optimize EvalHub's benchmark evaluations by using Kueue for fair resource sharing, priority-based job scheduling, and automatic queueing.
Explore how modern agentic AI improves upon traditional text-to-SQL approaches.
Learn how Model-as-a-Service (MaaS) solves the problem of managing AI costs, security, and models for every developer in an organization.
Configure input guardrails for your Red Hat OpenShift AI voice agent. Discover how to deploy TrustyAI, handle system limitations, and trace with MLflow.
Learn how llm-d routes each inference request to the GPU that already has the relevant data cached, cutting down on time-to-first-token, and doubling throughput without changing hardware. Discover how Red Hat's stack packages this neatly into a single Kubernetes resource.
Learn about the Data Governance Copilot, designed to make PostgreSQL databases accessible to non-technical users via agentic, natural language interaction while maintaining data compliance.
Learn how to create a functional Red Hat pizza shop voice agent using Red Hat OpenShift AI, focusing on practical architecture choices and implementation lessons learned along the way.
Learn how to use Red Hat OpenShift AI's reusable components to build modular AI pipelines, speed up development, and focus on what differentiates your applications.
Learn how to deploy Hermes Agent, a self-improving AI agent with a learning loop, on OpenShift AI with GPU-accelerated vLLM model serving.
Learn how evaluation-driven development (EDD) turns AI optimization from an art into an engineering discipline with EvalHub.
Learn how we fine-tuned the vLLM Semantic Router's embedding model to reduce misrouting rates and improve routing accuracy in enterprise deployments.
Learn how to deploy and serve large language models (LLM) on Rebellions ATOM NPUs using Red Hat OpenShift AI and a certified vLLM container image on the Red Hat AI Inference Server. This post walks through the steps to set up the joint solution between Red Hat and Rebellions, including installing the Node Feature Discovery operator, the Rebellions NPU operator, creating the ATOM hardware profile in OpenShift AI, and creating the vLLM RBLN ServingRuntime.
A Llama Stack-dependent backend, or any rapidly-evolving upstream project faces a version-drift problem. Explore our no-cost solution that provides early warnings.
Learn how to transform a simple chatbot into an enterprise RAG application by applying metadata filtering, hybrid search, and neural reranking using the OGX framework in Red Hat OpenShift AI.
Learn how to extend your large language model's capabilities with Model Context Protocol (MCP) servers and skills.
Discover how Red Hat OpenShift AI 3.4's Models-as-a-Service (MaaS) capability streamlines AI inference by acting as an integrated AI gateway within the platform, providing centralized governance and routing requests to both self-hosted models and external providers.
Learn how to prevent silent failures in your production AI inference stack with end-to-end benchmarking.
Learn how to prevent GPU waste and financial loss by implementing just-in-time (JIT) checkpointing with Kubeflow Training SDK on OpenShift AI.
Learn about GPU compute kernels, their role in distributed AI inference, and the Hugging Face Kernel Hub.
Learn how our team implemented CI/CD pipelines for the it-self-service-agent AI quickstart and the benefits of using CI/CD for agentic systems.
Learn how Red Hat AI can help address the security challenges of AI agents in production, from semantic malware to container escapes.