Red Hat OpenShift AI

Red Hat AI
Article

How AI observability works with MLflow

Cedric Clyburn

Learn how AI observability works with MLflow, an open source tool that can trace an agentic query and find the root cause of discrepancies between AI assistant responses and dashboard data. This post discusses the practical value of tracing an agent and what MLflow adds to an observability stack.

Featured image for Red Hat OpenShift AI.
Article

The tokenomics of self-hosted LLMs

Trevor Royer

Learn how to optimize self-hosted LLM cost per token. Cut GPU spending and maximize real-world throughput with autoscaling, right-sizing, and vLLM tuning.

Red Hat AI
Article

AutoRAG: Optimizing RAG for small models

Isaac Tigges

Stop guessing RAG settings. Discover how AutoRAG uses fast evaluation sweeps to optimize chunking and retrieval precision for small LLMs on your data.

Get started with vLLM feature share
Learning path

Get started with vLLM

Cedric Clyburn +1

Learn how to compress, serve, and benchmark LLMs with vLLM.