Lab 4: GPU Scheduling & DRA
Lab 4: GPU Scheduling & DRA
Goal
Explore how Kubernetes schedules GPU workloads using fake-gpu-operator, then see how DRA (Dynamic Resource Allocation) evolves the model with expressive, CEL-based device claims.
Part A: GPU Scheduling in Action
Each lab starts with a fresh playground, so we need to re-install fake-gpu-operator before pushing the scheduler.
1. Install fake-gpu-operator
helm repo add fake-gpu-operator https://fake-gpu-operator.storage.googleapis.com
helm repo update
helm install fake-gpu-operator fake-gpu-operator/fake-gpu-operator \
--namespace gpu-operator --create-namespace \
--set topology.nodes[0].gpuModel=A100 \
--set topology.nodes[0].gpuCount=4 \
--set 'topology.nodes[0].nodeLabels.nvidia\.com/gpu\.product=A100'
kubectl label node node-01 run.ai/simulated-gpu-node-pool=default --overwrite
sleep 15
2. Check GPU capacity
kubectl get node node-01 -o jsonpath='GPU capacity: {.status.capacity.nvidia\.com/gpu}'
echo ""
3. Deploy a GPU-hungry workload
Let's request more GPU pods than we have GPUs:
cat <<EOF | kubectl apply -f -
apiVersion: apps/v1
kind: Deployment
metadata:
name: gpu-greedy
spec:
replicas: 6
selector:
matchLabels:
app: gpu-greedy
template:
metadata:
labels:
app: gpu-greedy
spec:
nodeSelector:
run.ai/simulated-gpu-node-pool: default
containers:
- name: worker
image: busybox:1.36
command: ["sh", "-c", "echo 'GPU worker running' && sleep 300"]
resources:
limits:
nvidia.com/gpu: 1
EOF
4. Watch the scheduling
sleep 10
kubectl get pods -l app=gpu-greedy
Some pods run immediately. The rest stay Pending — the scheduler correctly enforces GPU limits, even with emulated hardware.
kubectl describe pod $(kubectl get pods -l app=gpu-greedy --field-selector=status.phase=Pending -o name | head -1) | grep -A3 Events
Look for: Insufficient nvidia.com/gpu
5. Clean up
kubectl delete deployment gpu-greedy
Part B: DRA — The Future of Device Allocation
The old device-plugin model is one-dimensional: "give me N GPUs." DRA replaces it with expressive claims.
6. Old way vs new way
Old (Device Plugins):
resources:
limits:
nvidia.com/gpu: 1 # "Give me a GPU. Any GPU."
New (DRA with CEL):
spec:
devices:
requests:
- name: gpu
exactly:
deviceClassName: gpu.nvidia.com
allocationMode: ExactCount
count: 1
selectors:
- cel:
expression: >
device.attributes["gpu.nvidia.com"].productName == "H100"
7. Check DRA API availability
kubectl api-resources | grep resource.k8s.io
8. Apply a DRA ResourceClaim
cat <<EOF | kubectl apply -f -
apiVersion: resource.k8s.io/v1
kind: ResourceClaim
metadata:
name: training-gpu
spec:
devices:
requests:
- name: gpu
exactly:
deviceClassName: gpu.nvidia.com
allocationMode: ExactCount
count: 1
selectors:
- cel:
expression: 'device.attributes["gpu.nvidia.com"].productName == "H100"'
EOF
kubectl get resourceclaim training-gpu -o yaml
The claim stays Pending — no DRA driver is installed (that requires the real NVIDIA DRA driver, donated to CNCF at KubeCon EU 2026). The point is seeing the API.
9. Apply a topology-aware claim
cat <<EOF | kubectl apply -f -
apiVersion: resource.k8s.io/v1
kind: ResourceClaim
metadata:
name: distributed-training
spec:
devices:
requests:
- name: gpu-pair
exactly:
deviceClassName: gpu.nvidia.com
allocationMode: ExactCount
count: 2
constraints:
- requests: ["gpu-pair"]
matchAttribute: "gpu.nvidia.com/numa-node"
EOF
This says: "give me 2 GPUs on the same NUMA node" — something device-plugins could never express.
What to Notice
- Part A proved GPU scheduling is a resource accounting problem — the scheduler doesn't need real hardware
- Part B showed DRA's expressive power — CEL selectors, topology constraints, device classes
- DRA is the first MUST requirement in the CNCF AI Conformance Program
- NVIDIA donated the GPU DRA driver to CNCF (March 2026) — this is going mainstream
Discussion
- Why did the community move from device-plugins to DRA?
- How does DRA enable multi-tenant GPU sharing vs MIG?
- What attributes would you select for in a production training job?
- Previous lesson
- Lab 3: KServe InferenceService
- Next lesson
- Lab 5: Gateway API Inference Extension