Optimizing distributed AI inference: Advanced deployment patterns
Learn about the three optimization levers for distributed AI inference: prefill/decode disaggregation, KV cache strategy, and speculative decoding.
Learn about the three optimization levers for distributed AI inference: prefill/decode disaggregation, KV cache strategy, and speculative decoding.
Learn how to connect the EvalHub runtime to internal or external model servers using service account tokens, API keys, or custom certificates.
This video demonstrates how to deploy Open Code, an open-source AI coding assistant, as a secure, multi-user web application on Red Hat OpenShift.
Learn about the five-dimensional design space in modern LLM serving, including tensor, pipeline, expert, data, and context parallelism.
Learn how to deploy MemPalace as a production-ready Model Context Protocol server on OpenShift AI with HTTP/WebSocket transport and Kubernetes health probes.
Learn how to optimize EvalHub's benchmark evaluations by using Kueue for fair resource sharing, priority-based job scheduling, and automatic queueing.
Discover the new features of Red Hat build of Apache Camel 4.18, including AI-driven semantic processing, Camel CLI Launcher, visual integration test flows, and more.
Discover how personal AI notebooks in Red Hat Developer Lightspeed can help developers find specific details in project documents quickly, grounded in context.
Look inside Red Hat AI Inference on Amazon EKS to understand its core architectural components and Kubernetes resources.
Discover how to use EvalHub and OCI persistence to make your AI evaluation results immutable, content-addressable, and fully auditable.
Explore how modern agentic AI improves upon traditional text-to-SQL approaches.
Explore the mechanics of gradient synchronization in PyTorch distributed training, focusing on MPI primitives like All-Reduce and core techniques like pipeline parallelism, tensor parallelism, and sharded data parallelism.
Learn when to use llama.cpp and vLLM for local inference of large language models (LLMs). Discover the key differences, benchmarks, and use cases for each engine.
Learn how speculative decoding can improve the performance of large language models (LLMs) in production by using a small, fast model to generate tokens speculatively and a large model to verify them.
Make your AI infrastructure more efficient by partitioning AMD Instinct GPUs via
Learn how Model-as-a-Service (MaaS) solves the problem of managing AI costs, security, and models for every developer in an organization.
Learn how to use the EvalHub CLI to automate AI evaluations in your CI/CD pipelines. Install the SDK, configure profiles, and set up a production gate.
Configure input guardrails for your Red Hat OpenShift AI voice agent. Discover how to deploy TrustyAI, handle system limitations, and trace with MLflow.
Learn how to onboard a custom evaluation framework into EvalHub using one class, one method, and a container image. This guide covers the contract, data structures, and a complete minimal adapter.
Learn about the Data Governance Copilot, designed to make PostgreSQL databases accessible to non-technical users via agentic, natural language interaction while maintaining data compliance.
Learn how to create a functional Red Hat pizza shop voice agent using Red Hat OpenShift AI, focusing on practical architecture choices and implementation lessons learned along the way.
Learn how to implement true gang autoscaling on OpenShift using Red Hat build of Kueue and ProvisionRequest API. This approach ensures efficient and reliable scheduling of high-performance workloads like AI/ML training, HPC simulations, or large data processing.
Headed to WeAreDevelopers World Congress Europe 2026? Visit the Red Hat Developer booth on-site to speak to our expert technologists.
Learn how the Krkn scenario generator, an AI-assisted chaos engineering tool for Kubernetes, addresses the challenge of translating desired chaos tests.
Learn how to read an existing system collection, understand its threshold logic, and build your own collection that encodes your actual measurement strategy with thresholds that mean something.