Course

Generative AI in Kubernetes

A hands-on workshop covering the state-of-the-art for running generative AI workloads on Kubernetes. Deploy LLMs with Ollama, use KServe for model serving, explore GPU scheduling with fake-gpu-operator and DRA, configure Gateway API Inference Extension, and learn about KAITO and llm-d. All exercises run on CPU — no GPU required.

Generative AI in Kubernetes (cover image)

About This Course

Workshop Overview

Duration: 90 minutes Level: Intermediate (comfortable with kubectl, YAML, Helm) Event: Devoxx Greece 2026

What You'll Build

Starting from a bare Kubernetes cluster, you'll build a complete AI inference platform:

LabTopic
Lab 1Cluster Setup + fake-gpu-operator
Lab 2Deploy an LLM (Ollama + TinyLlama)
Lab 3KServe InferenceService
Lab 4GPU Scheduling & DRA
Lab 5Gateway API Inference Extension
Lab 6KAITO Workspace
Lab 7llm-d Disaggregated Inference

Author

Alessandro Vozza — Cloud Native architect, Golden Kubestronaut, KubeCon speaker

About the Author

Alessandro Vozza

Alessandro Vozza

Find this author online