Challenge Easy

Node Is Ready, but Pods Won't Schedule

node-02 is back from a kernel upgrade but still shows Ready,SchedulingDisabled, and the Pods that can only run on it are stuck in Pending. Bring the node back into service and get the workload running, without touching the workload itself.

Scenario

node-02 is the only node in the cluster with a local SSD, and it carries the label storage=local-ssd. Last night the platform team took it out of service for a kernel upgrade. The upgrade is finished and the node is healthy again, but it still looks like this:

kubectl get nodes
NAME        STATUS                     ROLES           AGE     VERSION
cplane-01   Ready                      control-plane   2m12s   v1.36.4+k3s1
node-01     Ready                      <none>          2m6s    v1.36.4+k3s1
node-02     Ready,SchedulingDisabled   <none>          2m7s    v1.36.4+k3s1

In the meantime the ledger-api Deployment was rolled out in the payments namespace. It needs the local SSD, so its Pods can only run on node-02, and both replicas are stuck in Pending.


Task

Get both replicas of the ledger-api Deployment in the payments namespace Running on node-02.

Do not change the Deployment's configuration.

When you are done, node-02 must be fully back in service, with nothing left over from the maintenance window.

Important

Keep the ledger-api Deployment configuration exactly as it is.


Hint 1: Start from the Pods

A Pending Pod has not been placed on any node yet. The scheduler records why in the Pod's events:

kubectl get pods -n <namespace> -o wide
kubectl describe pod -n <namespace> <pod>

Read the FailedScheduling event at the bottom. It lists one reason for every node the scheduler rejected.

Hint 2: What SchedulingDisabled means

Ready and SchedulingDisabled come from two different places. Ready is a health condition the kubelet reports. SchedulingDisabled comes from a field in the Node object's spec that tells the scheduler to skip the node:

kubectl get node <node> -o jsonpath='{.spec.unschedulable}'

kubectl has a pair of commands that set and clear this field around node maintenance. Look under Cluster Management Commands in kubectl --help.

See: Safely Drain a Node

Hint 3: Still Pending after the node accepts Pods?

The scheduler stops at the first check a node fails, so a second problem on the same node only shows up after the first one is fixed. Describe the Pod again and compare the newest event with the earlier one.

The event says what kind of restriction is left, but not which one. The node's own description lists it. Look at the Taints line:

kubectl describe node <node>
Hint 4: Clearing what is left

A taint is removed with the same command that adds it, followed by a trailing minus:

kubectl taint node <node> <key>=<value>:<effect>-

You do not need to restart or recreate the Pods. The scheduler retries Pending Pods on its own as soon as the node changes.

See: Taints and Tolerations


⚒ Test Cases