Partition AMD Instinct GPU accelerators via device config manager (DCM) in Red Hat OpenShift
Make your AI infrastructure more efficient by partitioning AMD Instinct GPUs via
Make your AI infrastructure more efficient by partitioning AMD Instinct GPUs via
Learn how Model-as-a-Service (MaaS) solves the problem of managing AI costs, security, and models for every developer in an organization.
Learn how to use the EvalHub CLI to automate AI evaluations in your CI/CD pipelines. Install the SDK, configure profiles, and set up a production gate.
Configure input guardrails for your Red Hat OpenShift AI voice agent. Discover how to deploy TrustyAI, handle system limitations, and trace with MLflow.
Learn how to onboard a custom evaluation framework into EvalHub using one class, one method, and a container image. This guide covers the contract, data structures, and a complete minimal adapter.
Learn about the Data Governance Copilot, designed to make PostgreSQL databases accessible to non-technical users via agentic, natural language interaction while maintaining data compliance.
Learn how to create a functional Red Hat pizza shop voice agent using Red Hat OpenShift AI, focusing on practical architecture choices and implementation lessons learned along the way.
Learn how to implement true gang autoscaling on OpenShift using Red Hat build of Kueue and ProvisionRequest API. This approach ensures efficient and reliable scheduling of high-performance workloads like AI/ML training, HPC simulations, or large data processing.
Headed to WeAreDevelopers World Congress Europe 2026? Visit the Red Hat Developer booth on-site to speak to our expert technologists.
Learn how the Krkn scenario generator, an AI-assisted chaos engineering tool for Kubernetes, addresses the challenge of translating desired chaos tests.
Learn how to read an existing system collection, understand its threshold logic, and build your own collection that encodes your actual measurement strategy with thresholds that mean something.
Speculators v0.5.0 introduces DFlash support, enabling single-pass draft token generation with block diffusion for more efficient speculative decoding workflows. The release also adds unified online and offline training through vLLM’s native hidden states extraction system, improving training flexibility, version stability, and production readiness.
Learn how to use Red Hat OpenShift AI's reusable components to build modular AI pipelines, speed up development, and focus on what differentiates your applications.
Learn how to deploy Hermes Agent, a self-improving AI agent with a learning loop, on OpenShift AI with GPU-accelerated vLLM model serving.
Learn how evaluation-driven development (EDD) turns AI optimization from an art into an engineering discipline with EvalHub.
Learn how we fine-tuned the vLLM Semantic Router's embedding model to reduce misrouting rates and improve routing accuracy in enterprise deployments.
Learn about LogAn, an open source tool designed to overcome the limitations of using LLMs to analyze massive volumes of production logs.
Learn how to deploy and serve large language models (LLM) on Rebellions ATOM NPUs using Red Hat OpenShift AI and a certified vLLM container image on the Red Hat AI Inference Server. This post walks through the steps to set up the joint solution between Red Hat and Rebellions, including installing the Node Feature Discovery operator, the Rebellions NPU operator, creating the ATOM hardware profile in OpenShift AI, and creating the vLLM RBLN ServingRuntime.
A Llama Stack-dependent backend, or any rapidly-evolving upstream project faces a version-drift problem. Explore our no-cost solution that provides early warnings.
Learn how an expert red-teamed an infrastructure using Red Hat AI, OpenClaw, and abliterated models on Red Hat OpenShift on IBM Cloud.
Learn how to transform a simple chatbot into an enterprise RAG application by applying metadata filtering, hybrid search, and neural reranking using the OGX framework in Red Hat OpenShift AI.
Learn how to extend your large language model's capabilities with Model Context Protocol (MCP) servers and skills.
Discover how Red Hat OpenShift AI 3.4's Models-as-a-Service (MaaS) capability streamlines AI inference by acting as an integrated AI gateway within the platform, providing centralized governance and routing requests to both self-hosted models and external providers.
Learn how to prevent silent failures in your production AI inference stack with end-to-end benchmarking.
Learn how to prevent GPU waste and financial loss by implementing just-in-time (JIT) checkpointing with Kubeflow Training SDK on OpenShift AI.