A Kubernetes pod spec can ask for almost anything: the host's network stack, the host's process table, a privileged container with every Linux capability, a mount of the node's root filesystem. Any of those turns a compromised application into a compromised node, and a compromised node usually means the cluster. The Pod Security Standards are the upstream project's answer to the question which of those requests should an ordinary workload be allowed to make? They define three profiles, Privileged, Baseline and Restricted, as lists of pod fields and the values each profile permits. Pod Security Admission is the built-in admission controller that checks pods against them, stable since Kubernetes v1.25, the release that also removed PodSecurityPolicy.
This article goes control by control through both enforcing profiles, explains exactly what the admission controller evaluates and when, walks a real namespace from no policy to Restricted without breaking it, and covers the failure modes that bite in production. Facts were checked against the Kubernetes documentation source on 2026-10-03. Controls have version notes, so check the documentation for your own minor version.
The admission path in one picture
Three profiles and why there are only three
The three profiles form a strict ladder. Privileged is unrestricted and exists for system components that genuinely need the host: CNI plugins, storage drivers, node agents. Baseline blocks known privilege escalations while staying compatible with most ordinary container images; the design goal is that a typical web service runs under it unchanged. Restricted adds current hardening practice at some cost in compatibility: the container must run as a non-root user, drop all capabilities, use a seccomp profile and use only a short list of volume types.
The documentation's FAQ explains why there is no profile between Privileged and Baseline: privileges above Baseline are application specific, so a shared middle profile would either be too loose for most workloads or too tight for the ones that need it. That gap is filled with namespace isolation, exemptions or a policy engine, covered below. Every control is evaluated across all containers in the pod, including init and ephemeral containers. If any one container fails, the whole pod fails.
Baseline, control by control
Baseline is a list of things a pod must not ask for. The table summarises each control as written in the standard.
| Control | Fields | Allowed |
|---|---|---|
| HostProcess | windowsOptions.hostProcess (pod and containers) | unset or false |
| Host namespaces | hostNetwork, hostPID, hostIPC | unset or false |
| Privileged containers | securityContext.privileged | unset or false |
| Capabilities | capabilities.add | only AUDIT_WRITE, CHOWN, DAC_OVERRIDE, FOWNER, FSETID, KILL, MKNOD, NET_BIND_SERVICE, SETFCAP, SETGID, SETPCAP, SETUID, SYS_CHROOT |
| HostPath volumes | spec.volumes[*].hostPath | unset |
| Host ports | ports[*].hostPort | unset or 0 (a known list is not supported by the built-in controller) |
| Host probes and lifecycle hooks (v1.34+) | httpGet.host and tcpSocket.host in probes and hooks | unset or empty string |
| AppArmor | appArmorProfile.type and the legacy annotation | unset, RuntimeDefault or Localhost |
| SELinux | seLinuxOptions.type; user and role | type unset, container_t, container_init_t, container_kvm_t or container_engine_t (1.31+); user and role unset |
| /proc mount type | securityContext.procMount | unset or Default |
| Seccomp | seccompProfile.type | anything except Unconfined |
| Sysctls | spec.securityContext.sysctls[*].name | a safe, namespaced subset such as kernel.shm_rmid_forced, net.ipv4.ip_local_port_range, net.ipv4.ip_unprivileged_port_start, net.ipv4.tcp_syncookies and net.ipv4.ping_group_range, plus later additions |
Two details are easy to miss. The capability list is an allow-list for additions: it permits a short list of lower-risk capabilities, so Baseline stops you from adding SYS_ADMIN or NET_ADMIN but does not make you drop anything. And Baseline's seccomp rule only forbids an explicit Unconfined; a pod that sets no profile at all passes Baseline.
Restricted, control by control
Restricted includes everything in Baseline and adds six requirements. Unlike Baseline, several of them require you to set a field, so a pod spec that says nothing fails.
| Control | Requirement |
|---|---|
| Volume types | every volume must be configMap, csi, downwardAPI, emptyDir, ephemeral, persistentVolumeClaim, projected or secret |
| Privilege escalation | allowPrivilegeEscalation: false on every container |
| Running as non-root | runAsNonRoot: true, on each container or once at pod level |
| Running as non-root user | runAsUser must not be 0 anywhere it is set |
| Seccomp | profile must be RuntimeDefault or Localhost; unset is not allowed unless the pod-level field covers it |
| Capabilities | drop must include ALL; add may contain only NET_BIND_SERVICE |
Since v1.25 the profile reads spec.os.name: when it is windows, the privilege escalation, seccomp and capabilities controls are not required, because those fields have no effect on Windows. On Linux pods that use user namespaces (hostUsers: false), the runAsNonRoot and runAsUser checks are relaxed, because root inside such a pod maps to an unprivileged user on the host.
How Pod Security Admission evaluates a request
Pod Security Admission is configured mostly through namespace labels. Each of three modes takes a level and, optionally, a version:
pod-security.kubernetes.io/<MODE>: <LEVEL> # MODE: enforce | audit | warn
pod-security.kubernetes.io/<MODE>-version: <VERSION> # e.g. v1.34, or latestenforce rejects a violating pod. audit allows it but adds an annotation to the API server's audit event. warn allows it and returns a warning to the client, which kubectl prints. The modes are independent, so a namespace can enforce Baseline while warning and auditing at Restricted, which is the standard way to stage a tightening.
The most important behaviour to understand is the split between pods and workloads. enforce applies only to Pod objects. A Deployment, Job or StatefulSet whose template violates the enforced level is accepted; the controller then fails to create its pods. warn and audit, by contrast, also evaluate workload resources with pod templates, precisely so you see the problem at apply time. So a team that only reads apply output sees a warning, while the actual failure appears later, on the ReplicaSet.
The version label pins the definition of the level. restricted at v1.30 means the controls as they stood in 1.30, so a control added in a later release does not start rejecting pods when you upgrade the control plane. latest follows the running API server. A sensible default is to pin enforce and let warn and audit track latest, which shows you what the next upgrade would reject before it does.
Some pod updates are not re-evaluated: metadata changes other than the seccomp and AppArmor annotations, and valid updates to activeDeadlineSeconds or tolerations. Labelling a namespace does not touch running pods either; they are only checked again when recreated. The API server exports pod_security_evaluations_total, pod_security_exemptions_total and pod_security_errors_total, which are worth graphing during a rollout.
Cluster defaults and exemptions
Labels cover individual namespaces. Cluster-wide defaults for unlabelled namespaces, and exemptions, live in an AdmissionConfiguration file passed to kube-apiserver with --admission-control-config-file. The v1 configuration API requires v1.25 or later:
apiVersion: apiserver.config.k8s.io/v1
kind: AdmissionConfiguration
plugins:
- name: PodSecurity
configuration:
apiVersion: pod-security.admission.config.k8s.io/v1
kind: PodSecurityConfiguration
defaults: # used when a namespace has no label for that mode
enforce: "baseline"
enforce-version: "v1.34"
audit: "restricted"
audit-version: "latest"
warn: "restricted"
warn-version: "latest"
exemptions:
usernames: [] # authenticated or impersonated usernames
runtimeClasses: [] # e.g. a gVisor or Kata RuntimeClass
namespaces: ["kube-system"]The upstream default is Privileged for every mode, which means a fresh namespace enforces nothing. Exempt requests skip all three modes, not only enforce. Do not exempt controller service accounts such as system:serviceaccount:kube-system:replicaset-controller: pods are created by that identity, so the exemption silently extends to everyone who can create a ReplicaSet or Deployment. Exempting an end user only helps when they create pods directly. On managed clusters you may not control the API server flags, so namespace labels are often your only lever.
A pod spec that passes Restricted
This pod spec passes Restricted. Each line is there for a specific control, and setting the security context at pod level keeps container blocks short:
apiVersion: apps/v1
kind: Deployment
metadata: {name: api, namespace: payments}
spec:
replicas: 3
selector: {matchLabels: {app: api}}
template:
metadata: {labels: {app: api}}
spec:
securityContext:
runAsNonRoot: true # Running as non-root
runAsUser: 10001 # non-zero UID (image must not need root)
seccompProfile:
type: RuntimeDefault # Restricted requires it to be set
containers:
- name: api
image: registry.example.com/api@sha256:...
ports: [{containerPort: 8080}] # no hostPort
securityContext:
allowPrivilegeEscalation: false
capabilities: {drop: ["ALL"]}
readOnlyRootFilesystem: true # good practice; not required by Restricted
volumeMounts: [{name: tmp, mountPath: /tmp}]
volumes:
- name: tmp
emptyDir: {}readOnlyRootFilesystem is included because it is good hardening, but note that no profile requires it. The image needs a numeric non-root user; if the Dockerfile says USER app by name, the kubelet cannot verify runAsNonRoot without a numeric UID and the container fails to start, so set runAsUser explicitly or use a numeric USER.
Worked example: a namespace from nothing to Restricted
Take a payments namespace with no labels: an API Deployment, a Redis StatefulSet and a log shipper DaemonSet that mounts /var/log from the host. The goal is Restricted for the application and nothing broken on the way.
Step 1: measure. A server-side dry run evaluates the namespace's running pods against a level and prints warnings, without changing anything:
kubectl label --dry-run=server --overwrite ns payments \
pod-security.kubernetes.io/enforce=restricted
# Warning: existing pods in namespace "payments" violate the new PodSecurity enforce level "restricted:latest"
# Warning: api-7d9c...: allowPrivilegeEscalation != false, unrestricted capabilities, runAsNonRoot != true, seccompProfile
# Warning: logship-x2k...: hostPath volumesThe output is illustrative; the real messages list each violating pod and control. Running the same command against --all namespaces gives a cluster-wide inventory.
Step 2: separate what cannot comply. The log shipper reads node files through hostPath, which even Baseline forbids. It belongs in its own namespace with Privileged, owned by the platform team, not in the application namespace. Redis needed only a security context with a numeric non-root UID, plus a writable volume for its data.
Step 3: warn and audit before enforcing. Label the namespace enforce=baseline (nothing violates it once the shipper has moved), warn=restricted and audit=restricted. Developers now see warnings on every apply, and the audit log shows violations from CI and controllers. Fix the specs until warnings stop and the audit annotations disappear for a week.
Step 4: enforce, pinned. Set enforce=restricted with enforce-version pinned to the current minor version, keep warn and audit on latest, then roll every workload once so the running pods are proven against the policy rather than grandfathered in.
Failure modes
- Deployment applied, zero pods. Enforce rejects pods, not the Deployment, so
kubectl get deployshows 0/3 ready and the ReplicaSet's events show FailedCreate with the violated controls. Look atkubectl describe rs, not the Deployment. - The time bomb. Labelling a namespace does not evict running pods. Violators keep running until a node drain, eviction or rollout recreates them, then fail, often during an unrelated incident. Always roll workloads after tightening.
- Upgrade breaks pods. With
enforce-version: latest, a control added in a new release, such as the v1.34 probe host field, starts applying the moment the control plane upgrades. Pin enforce. - Over-broad exemptions. Exempting a controller service account or a whole namespace that ordinary teams can deploy into removes the control entirely. Audit exemptions like firewall rules.
- Runs as non-root by name only. runAsNonRoot with a username-only image fails at container start, not admission, so the error appears in pod events as a CreateContainerConfigError.
- Ephemeral debug containers.
kubectl debugadds ephemeral containers that are checked too; a debug profile that needs privileges is rejected in a Restricted namespace, by design.
Trade-offs and where it stops
Pod Security Admission is deliberately small. It only validates; it never mutates, so it will not add a missing seccomp profile for you. It has three fixed levels and no per-workload exceptions inside a namespace. It knows nothing about images, registries, labels, resource limits or service account tokens. In exchange it is built in, needs no webhook to keep available, and fails closed inside the API server.
When you need more, layer rather than replace. ValidatingAdmissionPolicy (CEL expressions evaluated in the API server) or a policy engine such as Kyverno or OPA Gatekeeper can require signed images from approved registries, add defaults, or allow one named workload a single extra capability. Keep Pod Security Admission underneath as the floor, because a webhook outage or a policy typo should not open the cluster. RBAC still decides who may create pods in which namespace, and that matters as much as what the pod may ask for. See container security architecture for seccomp, user namespaces and sandboxed runtimes, RBAC and ABAC for the identity side, software supply chain security for image policy, and secret rotation in Kubernetes for the secrets your pods mount.
What to do next
- Run
kubectl label --dry-run=server --overwrite ns --all pod-security.kubernetes.io/enforce=baselineand then with restricted, and save both outputs as your inventory. - Move every workload that needs host access into dedicated Privileged namespaces owned by the platform team.
- Set cluster defaults in AdmissionConfiguration (or label every namespace if you cannot): enforce baseline, warn and audit restricted.
- Add the pod-level and container security context from this article to your Helm charts or templates, and make CI fail on PSA warnings.
- Promote namespaces to enforce restricted with a pinned version, then roll every workload.
- Graph
pod_security_evaluations_totaland alert on ReplicaSet FailedCreate events. - Before each control-plane upgrade, compare warn output at latest against the pinned version and bump the pin deliberately.