Challenge Medium

Roll Out Pod Security Standards Across Three Namespaces

The platform team is rolling out Pod Security Admission. A monitoring agent is already blocked, the storefront team needs a staged rollout, and payments must move to the strictest level without losing its running service.

Scenario

The platform team is rolling out Pod Security Admission, the built-in admission controller that applies the Kubernetes Pod Security Standards per namespace. Three namespaces are in scope this week:

NamespaceWorkloadWhat it is
monitoringDaemonSet node-agentNode metrics agent. It reads host processes, so it uses the host network and PID namespaces and mounts the host's /proc.
storefrontDeployment storefront-webPublic website, 2 replicas. Its team has not hardened the image yet.
paymentsDeployment payments-apiCard payments API, 2 replicas. It handles card data and has to meet the strictest standard.

A few minutes ago, the platform team added a Pod Security label to the monitoring namespace, and node-agent stopped starting Pods right away. It is down on every node. storefront and payments have no Pod Security labels yet.

Note

Work from cplane-01. The API server on that machine writes its audit log to /var/log/kubernetes/audit.log, which needs sudo to read.


Task

  1. Get the node-agent DaemonSet running on every node. Its host access is legitimate, so change the monitoring namespace to the level that allows it, and leave the DaemonSet as it is.
  2. Set up the storefront namespace for a staged rollout: enforce the baseline level now, and report anything that would break the restricted level, both to whoever applies a change and in the API server audit log. The storefront-web Deployment must keep running.
  3. Make the payments namespace enforce the restricted level, with the payments-api Deployment running 2 replicas that all meet that level.
Important

Set every policy with the Pod Security Admission labels on the namespace. For the payments-api Deployment, change its Pod spec, not its image, and do not add exemptions. Leave the node-agent DaemonSet and the storefront-web Deployment as they are.


Hint 1 | Finding out why a controller creates no Pods

When an admission controller rejects a Pod, the controller that tried to create it records the reason as an event:

kubectl describe daemonset <name> -n <namespace>
kubectl get events -n <namespace> --field-selector reason=FailedCreate
kubectl get namespace <namespace> --show-labels
Hint 2 | The shape of a Pod Security label

Each namespace label sets one mode to one level. A namespace can carry one label per mode:

pod-security.kubernetes.io/<MODE>=<LEVEL>
kubectl label namespace <namespace> <label>=<value> --overwrite

A server-side dry run shows which running Pods a new enforce level would reject, without changing anything:

kubectl label --dry-run=server --overwrite namespace <namespace> <label>=<value>

See: Pod Security Standards, Pod Security Admission

Hint 3 | When a running Deployment is not compliant

A new enforce level never touches Pods that are already running, so a Deployment can look healthy while every new Pod is rejected. Force a rollout and see what the ReplicaSet reports:

kubectl rollout restart deployment/<name> -n <namespace>
kubectl get events -n <namespace> --field-selector reason=FailedCreate

The rejection message names each field the Pod spec needs, at Pod or container level:

securityContext:
  <field>: <value>
Hint 4 | Seeing what audit mode recorded

Audit mode does not print anything. It adds an annotation to the matching entry in the API server audit log:

sudo grep -c <annotation> /var/log/kubernetes/audit.log

⚒ Test Cases