The tokenomics of self-hosted LLMs
Learn how to optimize self-hosted LLM cost per token. Cut GPU spending and maximize real-world throughput with autoscaling, right-sizing, and vLLM tuning.
Learn how to optimize self-hosted LLM cost per token. Cut GPU spending and maximize real-world throughput with autoscaling, right-sizing, and vLLM tuning.
Secure Claude Code plug-ins: Learn how to protect your repository and implement review discipline for safer third-party code.
Improve CI/CD for Kubeflow Pipelines and Open Data Hub: Learn practical patterns for simplifying builds, fixing flaky tests, and more.
Explore LLM inference on Kubernetes using Red Hat AI on EKS. Trace requests from the Envoy gateway through the EPP scheduler down to individual vLLM pods.
Optimize LLM deployment with Neural Navigator on Red Hat OpenShift AI, reducing cost overruns and latency spikes.
Confidently deploy LLMs with Red Hat support: Learn how to determine if your model is supported by Red Hat's vLLM community.
See how a team upgraded Red Hat OpenShift AI 3.3.2 three to four times faster with an AI coding assistant, reducing engineering effort by 60%.
Reduce observability costs with Red Hat OpenShift AI summarizer, bridging the interpretation gap for cloud-native architectures.
Stop guessing RAG settings. Discover how AutoRAG uses fast evaluation sweeps to optimize chunking and retrieval precision for small LLMs on your data.
Optimize GPU efficiency with Red Hat OpenShift AI 3.4's flow control for llm-d, ensuring priority-based request queuing and fairness policies.
Learn how to build a distributed RAG pipeline with Ray Data on OpenShift AI for high-performance parsing, embedding, and writing to a vector database.
Learn how to preserve existing model knowledge with Orthogonal Subspace Fine-Tuning (OSFT) for fine-tuning large language models in Training Hub.
Operationalize AI agents with OpenShift and Kubernetes primitives for manageable, cost-optimized AIOps.
Learn how to build an open cloud native architecture for AI agents. This blueprint explains how to improve workload isolation and implement inference routing.
Learn how to set up local agentic AI computer use. Run quantized models like Qwen 3.6 with Hermes to automate desktop tasks on your own terms today.
Deploy a self-hosted AI coding assistant with vLLM and Red Hat OpenShift AI for privacy and operational independence.
Learn how to unlock observability for Models-as-a-Service in Red Hat OpenShift AI 3.4 with the new usage dashboard.
Learn about the llm-d batch gateway, a Kubernetes-native batch inference service that plugs into the same llm-d inference stack managed by Red Hat OpenShift AI.
Learn how to scale document processing with a guided example that combines Docling for structure-aware parsing, Ray Data for distributed streaming execution, and Red Hat OpenShift AI.
Discover how the MCP standardizes tool integration, how event streaming improves user experience, and how to safely deploy stochastic reasoning engines.
Learn how to implement GPU-as-a-Service on Red Hat OpenShift using Kueue, NVIDIA MIG, and a custom dashboard plug-in for self-service GPU resource booking.
Explore the Data Governance Copilot architecture, integrating OpenShift AI with PG Airman MCP server for robust, agentic Postgres analytics.
Learn how to connect a modern Apache Iceberg lakehouse to LLM-hosted models using nothing but SQL on Red Hat OpenShift AI.
Learn how to deploy MemPalace as a production-ready Model Context Protocol server on OpenShift AI with HTTP/WebSocket transport and Kubernetes health probes.