What GPU kernels mean for your distributed inference
Learn about GPU compute kernels, their role in distributed AI inference, and the Hugging Face Kernel Hub.
Learn about GPU compute kernels, their role in distributed AI inference, and the Hugging Face Kernel Hub.
Learn how our team implemented CI/CD pipelines for the it-self-service-agent AI quickstart and the benefits of using CI/CD for agentic systems.
Learn how Red Hat AI 3.4 uses EvalHub to orchestrate AI evaluations on Kubernetes. Scale frameworks like Garak and LightEval with built-in MLflow tracking.
Learn how to combine KServe and llm-d to optimize generative AI inference, improve performance, and reduce infrastructure costs. This article demonstrates the integration architecture and provides practical guidance for AI platform teams.
Users can deploy vLLM on a variety of hardware with a simple command. But a lot of work goes on below the surface to make the magic happen.