
Duration: 90 minutes Level: Intermediate (comfortable with kubectl, YAML, Helm) Event: Devoxx Greece 2026
Starting from a bare Kubernetes cluster, you'll build a complete AI inference platform:
| Lab | Topic |
|---|---|
| Lab 1 | Cluster Setup + fake-gpu-operator |
| Lab 2 | Deploy an LLM (Ollama + TinyLlama) |
| Lab 3 | KServe InferenceService |
| Lab 4 | GPU Scheduling & DRA |
| Lab 5 | Gateway API Inference Extension |
| Lab 6 | KAITO Workspace |
| Lab 7 | llm-d Disaggregated Inference |
Alessandro Vozza — Cloud Native architect, Golden Kubestronaut, KubeCon speaker