AI inference

Featured image for vLLM interference article.
Article

What did AI cost you this quarter?

Markell Rawls +1

Track AI inference costs by department using Red Hat OpenShift AI’s built-in MaaS gateway API keys, Perses dashboards, and MLflow tracing, with no manual instrumentation required.

Featured image for open source.
Article

Use a local and open source code assistant

Seth Kenlon

Learn how to install and configure OpenCode, an open source AI coding assistant, for local development with Ollama, OpenVINO, or Red Hat AI.

Featured image for Red Hat OpenShift AI.
Article

Orchestrate production RAG with OpenShift AI

Ana Biazetti +1

Learn how to automate RAG document processing with Red Hat OpenShift AI for production-ready parsing, embedding, and querying of PDF documents.

Featured image for Red Hat OpenShift AI.
Article

The tokenomics of self-hosted LLMs

Trevor Royer

Learn how to optimize self-hosted LLM cost per token. Cut GPU spending and maximize real-world throughput with autoscaling, right-sizing, and vLLM tuning.

Red Hat AI
Article

AutoRAG: Optimizing RAG for small models

Isaac Tigges

Stop guessing RAG settings. Discover how AutoRAG uses fast evaluation sweeps to optimize chunking and retrieval precision for small LLMs on your data.