Generative AI on Kubernetes

Generative AI on Kubernetes by Roland Huss and Daniele Zonca is a comprehensive guide published by O'Reilly Media that addresses the challenges of running Generative AI workloads at scale on Kubernetes. The book covers the full lifecycle of deploying large language models (LLMs), from model serving and data management to GPU scheduling, observability, and model customization. Readers will gain practical knowledge of tools such as KServe, Ray Serve, KubeFlow, and vLLM, as well as architectural patterns for AI-driven applications including Retrieval-Augmented Generation (RAG). This title is available as an Early Release, giving readers access to the authors' raw and unedited content as it is written.

Preview thumbnail

In this guide, you'll explore

  • Comprehensive coverage of deploying and serving large language models (LLMs) on Kubernetes using tools like KServe, vLLM, Ray Serve, and KubeFlow
  • In-depth guidance on GPU discovery, workload scheduling, multi-GPU inference, and resource optimization for AI workloads
  • Practical observability strategies including metrics, tracing, logging, and model safety with hallucination detection and guardrails
  • Step-by-step instructions for model customization, fine-tuning jobs on Kubernetes, and running production-grade tuning pipelines
  • Architectural patterns for AI-driven applications including chat apps, backend AI services, and Retrieval-Augmented Generation (RAG)

Get Instant Access

By submitting this form, you agree to have your contact information, including email and phone, processed by ebulletins and the sponsors of this page for the purpose of following up on your professional interests.