Generative AI on Kubernetes by Roland Huss and Daniele Zonca is a comprehensive guide published by O'Reilly Media that addresses the challenges of running Generative AI workloads at scale on Kubernetes. The book covers the full lifecycle of deploying large language models (LLMs), from model serving and data management to GPU scheduling, observability, and model customization. Readers will gain practical knowledge of tools such as KServe, Ray Serve, KubeFlow, and vLLM, as well as architectural patterns for AI-driven applications including Retrieval-Augmented Generation (RAG). This title is available as an Early Release, giving readers access to the authors' raw and unedited content as it is written.