Red Hat OpenShift AI

Featured image for Red Hat OpenShift AI.
Article

The tokenomics of self-hosted LLMs

Trevor Royer

Learn how to optimize self-hosted LLM cost per token. Cut GPU spending and maximize real-world throughput with autoscaling, right-sizing, and vLLM tuning.

Red Hat AI
Article

AutoRAG: Optimizing RAG for small models

Isaac Tigges

Stop guessing RAG settings. Discover how AutoRAG uses fast evaluation sweeps to optimize chunking and retrieval precision for small LLMs on your data.

Get started with vLLM feature share
Learning path

Get started with vLLM

Cedric Clyburn +1

Learn how to compress, serve, and benchmark LLMs with vLLM.