A container is not a small virtual machine. It is an ordinary Linux process that the kernel shows a restricted view of the system: its own process list, network stack and filesystem tree through namespaces, and limited CPU and memory through cgroups. Every container on a node shares one kernel. That single fact shapes container security. An attacker who gains code execution in a container is already running on the host kernel, and every control you add is an attempt to make that kernel refuse what the attacker asks for next.

This article builds a container security architecture layer by layer, from the image you build to the alerts you act on, with Kubernetes as the running example. For each layer it explains what the control does in the kernel or the control plane, shows a working configuration, and says what it does not stop. It ends with a worked breach, the failure modes seen in real clusters, and a checklist you can apply to one namespace this week.

Advertisement

The threat model

Start by naming what you defend against, because each control maps to one of these.

  • Compromised workload. A remote code execution bug in your application or a dependency gives an attacker a process inside the container. This is the common case and the one most of this architecture limits.
  • Container escape. From that process the attacker reaches the host, through a misconfiguration (a privileged container, a host path mount, the container runtime socket) or a kernel or runtime vulnerability. Real runtime bugs have done this, such as the runc binary overwrite in CVE-2019-5736 and the leaked file descriptor in CVE-2024-21626.
  • Lateral movement. Without escaping, the attacker uses the pod's network access and service account token to reach databases, the cloud metadata service or the Kubernetes API.
  • Malicious or vulnerable image. The code was bad before it ran. Signing, provenance and SBOMs address this and are covered in software supply chain security; this article assumes those exist and covers what happens at run time.

Notice that the image layer reduces the chance of the first item, but everything after the first item is decided by runtime configuration. A perfectly scanned image running privileged is still one bug from owning the node.

The architecture at a glance

1. Buildminimal, non-root image2. Registryscan, sign, pin by digest3. AdmissionPod Security, policy4. SchedulingRuntimeClass, node pool5. Node: what stands between the container process and the hostContainer processUID 10001, no shellcapabilities: drop ALLno_new_privs, read-only rootseccomp RuntimeDefaultsyscall filterAppArmor or SELinuxuser namespace (hostUsers false)Runtimerunc, or gVisor / Kata sandboxcgroups: CPU, memory, pidsnamespaces: pid, net, mntShared host kernela kernel bug crosses every layer above unless a sandbox runtime adds its own6. Networkdefault-deny policy, egress allowlist, no metadata7. Detection and responseeBPF runtime events, audit logs, kill and rebuild
Seven layers, from build to response. Layers 1 to 4 decide what is allowed to run; layer 5 decides what a running process can ask the kernel for; layers 6 and 7 limit and expose what a compromised process does next.

Layers are deliberately redundant. Dropping capabilities and applying seccomp overlap; so do a read-only root filesystem and a non-root user. Redundancy is the point: misconfigurations happen, and each layer should hold when its neighbour is missing.

Advertisement

Layers 1 and 2: an image with less in it

Every binary in the image is a tool for an attacker and a source of scanner findings. Build in one stage and ship only the artefact in a minimal base: a distroless or scratch image with no shell, no package manager and no compilers. Set a numeric non-root user in the image, so that Kubernetes can verify it without guessing what a user name maps to.

# syntax=docker/dockerfile:1
# pin by digest in real use: golang:1.23@sha256:...
FROM golang:1.23 AS build
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN CGO_ENABLED=0 go build -trimpath -ldflags="-s -w" -o /out/api ./cmd/api

# no shell, no package manager
FROM gcr.io/distroless/static-debian12:nonroot
COPY --from=build /out/api /api
# numeric, so runAsNonRoot can verify it
USER 65532:65532
EXPOSE 8080
ENTRYPOINT ["/api"]

Two habits matter as much as the base. Never copy secrets into a layer, because deleting them in a later layer leaves them in the earlier one; use build secrets or runtime injection. And reference images by digest, not by tag, so that what was scanned is what runs. A scanner reports known vulnerabilities in packages it can identify. It does not find your application's logic bugs, malicious code without a CVE, or risky runtime configuration, which is why the remaining layers exist.

Layer 3: admission with Pod Security Standards

Kubernetes defines three Pod Security Standards. Privileged allows everything and exists for system components. Baseline blocks known privilege escalations: privileged containers, host namespaces, host path volumes, most added capabilities and host ports. Restricted adds hardening: the pod must run as non-root, set allowPrivilegeEscalation to false, drop all capabilities (adding back only NET_BIND_SERVICE), use the RuntimeDefault or a Localhost seccomp profile, and use only a limited set of volume types.

The built-in Pod Security Admission controller enforces these per namespace through labels. Each label sets a mode: enforce rejects violating pods, audit records them in the audit log, and warn returns a warning to the client.

apiVersion: v1
kind: Namespace
metadata:
  name: payments
  labels:
    pod-security.kubernetes.io/enforce: restricted
    pod-security.kubernetes.io/enforce-version: latest
    pod-security.kubernetes.io/audit: restricted
    pod-security.kubernetes.io/warn: restricted

Roll out by setting warn and audit first, fixing what they report, then enforce. Pin enforce-version to a specific Kubernetes version if you want upgrades not to tighten the rules silently. Pod Security Admission checks pod specs only. For rules it cannot express, such as allowed registries, required digests or required labels, add a policy layer: the built-in ValidatingAdmissionPolicy with CEL expressions, or an engine such as Kyverno or OPA Gatekeeper.

Layer 5: what the process can ask the kernel for

Admission decides whether a pod spec is acceptable; these settings are what the spec asks the kernel to enforce. Here is a deployment that passes the restricted level and adds two further controls.

apiVersion: apps/v1
kind: Deployment
metadata: {name: api, namespace: payments}
spec:
  replicas: 3
  selector: {matchLabels: {app: api}}
  template:
    metadata: {labels: {app: api}}
    spec:
      automountServiceAccountToken: false   # mount a token only if the app calls the API
      hostUsers: false                      # user namespace: root in pod is not root on host
      securityContext:
        runAsNonRoot: true
        runAsUser: 65532
        seccompProfile: {type: RuntimeDefault}
      containers:
      - name: api
        image: registry.example.com/api@sha256:<digest>
        securityContext:
          allowPrivilegeEscalation: false   # sets no_new_privs
          readOnlyRootFilesystem: true
          capabilities: {drop: ["ALL"]}
        resources:
          requests: {cpu: 250m, memory: 256Mi}
          limits: {memory: 256Mi}
        volumeMounts: [{name: tmp, mountPath: /tmp}]
      volumes: [{name: tmp, emptyDir: {sizeLimit: 64Mi}}]
  • Capabilities. Linux splits root's power into capabilities such as NET_ADMIN, SYS_ADMIN and NET_RAW. Container runtimes grant a default subset; dropping ALL removes the ones your service almost certainly does not need, including NET_RAW, which allows crafting packets.
  • No new privileges. allowPrivilegeEscalation false sets the kernel's no_new_privs flag, so setuid binaries and file capabilities cannot raise privileges after start.
  • seccomp. A seccomp filter makes the kernel reject listed system calls before they run. RuntimeDefault is the container runtime's profile, which blocks rarely needed and historically dangerous calls such as those for loading kernel modules or changing the system clock. Kubernetes does not apply it unless the pod asks or the kubelet's seccompDefault setting is enabled, so set it explicitly.
  • AppArmor or SELinux. These Linux security modules restrict which files, mounts and operations a process can touch, independent of user IDs. Use the profile your distribution's runtime ships unless you have a reason to write your own.
  • Read-only root filesystem. An attacker cannot drop tools into the image's paths or modify its binaries. Give the app small writable emptyDir mounts for the paths it really writes, such as /tmp.
  • User namespaces. With hostUsers false, UID 0 inside the pod maps to an unprivileged UID range on the host, so a process that escapes as root lands as nobody in particular. The feature is beta and enabled by default since Kubernetes v1.33, and needs a supporting kernel, runtime and filesystem; check the official requirements before relying on it.

Layer 4: sandboxed runtimes for code you do not trust

All of layer 5 still leaves the shared kernel exposed to whatever system calls are permitted. For untrusted workloads, such as customer code, build jobs from forks or code written by an AI agent, put another kernel boundary in the way. gVisor runs containers on a user-space kernel that implements Linux system calls itself, so the host kernel sees a much smaller interface. Kata Containers runs each pod in a lightweight virtual machine with its own guest kernel. Both are selected per pod through a RuntimeClass, so one cluster can run trusted services on runc and untrusted jobs on a sandbox, usually on a separate node pool.

The cost is real. System-call-heavy and I/O-heavy workloads run slower, some kernel features and device access are unsupported, and VM-based sandboxes add memory per pod and start-up time. Measure your workload on the sandbox before committing. The agent sandboxing article shows how to combine a sandbox with egress and credential policy for model-generated code.

Layers 6 and 7: network and detection

Kubernetes allows all pod-to-pod traffic by default. Start every namespace with a policy that denies all ingress and egress, then allow exactly the flows the service needs, including DNS. Policies are only enforced if the cluster's network plugin implements them, so test with a pod that should be refused.

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata: {name: default-deny, namespace: payments}
spec:
  podSelector: {}
  policyTypes: [Ingress, Egress]
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata: {name: api-allow, namespace: payments}
spec:
  podSelector: {matchLabels: {app: api}}
  policyTypes: [Ingress, Egress]
  ingress:
  - from: [{namespaceSelector: {matchLabels: {kubernetes.io/metadata.name: gateway}}}]
    ports: [{port: 8080}]
  egress:
  - to: [{podSelector: {matchLabels: {app: postgres}}}]
    ports: [{port: 5432}]
  - to: [{namespaceSelector: {}, podSelector: {matchLabels: {k8s-app: kube-dns}}}]
    ports: [{port: 53, protocol: UDP}]

Block the cloud instance metadata endpoint from pods that do not need it, and give pods cloud identity through workload identity features rather than node credentials. Turn off service account token automounting by default; a stolen token with broad RBAC is often a faster path than a kernel exploit. These are the network-level expressions of zero trust: every flow is explicitly allowed, never assumed.

Prevention fails eventually, so watch for it. Runtime detection tools built on eBPF, such as Falco and Tetragon, observe process starts, file writes and connections inside containers. The highest-signal rules for hardened workloads are simple because normal behaviour is narrow: a shell starting in a container that has none, a write to a read-only path, an unexpected outbound connection, or a new binary executing. Respond by killing the pod and redeploying from the known-good digest, and preserve evidence first.

Worked example: one bug, two clusters

An attacker finds a deserialization bug in the payments API and gets code execution. Follow them through two clusters.

In the unhardened cluster the container runs as root, privileged, with the default service account token mounted and open egress. The attacker downloads a toolkit with the image's shell, reads the token and lists secrets the account can see, queries the metadata service for the node's cloud credentials, then mounts the host disk through the privileged device access and installs persistence on the node. Minutes, no kernel exploit needed.

In the hardened cluster the same bug gives a process running as UID 65532 inside a user namespace. There is no shell or download tool. The root filesystem is read-only and /tmp is 64 MiB. There is no service account token, the metadata endpoint is blocked, and egress allows only Postgres and DNS. The attacker can still misuse the application's own database access, which is why the app's database role must be least-privilege too, but any attempt to fetch tools, execute a dropped binary or open a new connection triggers detection. The bug is the same; the outcome is an alert instead of a node takeover.

Failure modes in real clusters

  • Exemptions that outlive their reason. A namespace set to privileged for one agent becomes the place everything that fails admission gets deployed. Review exemptions monthly.
  • Debug pods. Privileged troubleshooting pods left running are standing escape hatches. Use ephemeral debug containers with a time limit instead.
  • Runtime socket or host path mounts. Mounting the container runtime socket or a host directory hands over the node. Baseline forbids host paths; enforce it.
  • Read-only root breaks the app. Teams turn the control off rather than finding the one path the app writes. Add emptyDir mounts instead.
  • Mutable tags. A scanned tag is replaced by an unscanned image. Require digests at admission.
  • Scanner fatigue. Thousands of findings in unused packages bury the exploitable one. Shrink the image first, then triage what remains by reachability.

Trade-offs

ControlProtects againstCost
Restricted Pod SecurityMost misconfiguration escapesSome images need rebuilding to run as non-root
seccomp RuntimeDefaultMany kernel attack pathsRare breakage for apps using unusual system calls
User namespacesRoot-in-container becoming root-on-hostKernel and runtime requirements; some volume types
gVisor or KataKernel exploits from untrusted codePerformance overhead, compatibility gaps, extra memory
Default-deny networkLateral movement and exfiltrationEvery new dependency needs a policy change

If you are on a managed container platform rather than Kubernetes, many of these knobs are set for you, some are not exposed, and the provider's sandbox is your kernel boundary; cloud container services covers those platforms, and the Kubernetes introduction covers the objects used here.

What to do next

  1. Pick one namespace and set Pod Security warn and audit to restricted; list every violation it reports.
  2. Rebuild the worst offenders on a minimal base with a numeric non-root user, referenced by digest.
  3. Add seccompProfile RuntimeDefault, drop ALL capabilities, set allowPrivilegeEscalation false and readOnlyRootFilesystem true, then switch the namespace to enforce.
  4. Disable service account token automounting and apply a default-deny network policy with explicit allows, including DNS.
  5. Test hostUsers false on a non-critical workload once your cluster version and nodes meet the requirements.
  6. Move any untrusted-code workload to a sandboxed RuntimeClass on its own node pool.
  7. Deploy a runtime detection tool with a shell-in-container rule and rehearse the kill-and-redeploy response.
Key takeaway: Containers share the host kernel, so container security is the work of making that kernel refuse an attacker's next request. Ship minimal non-root images by digest, enforce the restricted Pod Security level, drop capabilities, apply seccomp and an LSM, make the root filesystem read-only, use user namespaces where supported, sandbox untrusted code with gVisor or Kata, deny network flows by default and watch for the narrow behaviour a hardened container should never show.