Scenario
The Job batch-job in namespace batch was deployed but keeps failing
and is close to exhausting its retry limit. A few seconds after the pod
starts, it stops and the Job shows Error status.
Inspect the cluster and the manifest to find out why it is not completing, fix the bug in the manifest, and reapply it.
Task
Inspect the running Job and the manifest at /home/laborant/batch-job.yaml.
Find the bug, fix it in the manifest, then delete and reapply the Job.
Hint 1 — Inspect the Job and its pods
Check the current Job state:
kubectl get job batch-job -n batch
kubectl describe job batch-job -n batch
Then look at what the pod printed before it stopped:
kubectl logs -n batch -l job-name=batch-job
Documentation
Hint 2 — Understand what Kubernetes uses to judge success
When a container process finishes, it returns a numeric exit code to the operating system. Kubernetes uses this — not the output — to determine success or failure:
- Exit code
0— the container succeeded - Any other exit code — the container failed
Read the manifest and look at the command field carefully:
cat /home/laborant/batch-job.yaml
Documentation
Hint 3 — Reapply the Job after fixing the manifest
A Job's pod template is immutable after creation. Delete the existing Job before reapplying:
kubectl delete job batch-job -n batch
kubectl apply -f /home/laborant/batch-job.yaml
Wait for completion:
kubectl wait job/batch-job -n batch \
--for=condition=complete \
--timeout=60s
Confirm:
kubectl get job batch-job -n batch
Documentation