Docker is usually introduced as a lightweight virtual machine. That picture gets you through a first tutorial and then fails you at the first real problem: a container that ignores Ctrl-C, an image that leaks a secret you deleted, a port that is open to the internet despite your firewall, a process killed with exit code 137 for no visible reason. Each of those makes sense the moment you see what a container actually is.

A container is an ordinary Linux process. The kernel gives it a private view of the system through namespaces, caps the resources it may use through control groups, and starts it on a root filesystem assembled from read-only image layers. Docker is the tooling that builds those filesystems, ships them through registries and asks the kernel to start processes inside them. This page builds that model from the bottom up, then uses it to write a production-grade Dockerfile and run it. For what comes after one host, see the Kubernetes introduction.

Advertisement

A container is a process with a restricted view

Start a container and look at the host's process list: the container's main process is there, with a normal host PID, owned by the kernel like any other. What makes it a container is three kernel features applied together.

Namespaces decide what the process can see. A PID namespace makes it believe it is PID 1 and hides every other process. A mount namespace gives it its own filesystem tree. A network namespace gives it its own interfaces, routing table and port space, so two containers can both listen on port 8000. UTS, IPC, user and cgroup namespaces isolate the hostname, shared memory, user ids and the cgroup tree in the same way.

Control groups (cgroups, now version 2 on current distributions) decide how much it may use: memory ceilings, CPU shares and quotas, a maximum process count, and I/O weights. When a container exceeds its memory limit, the kernel's OOM killer ends a process inside it, which surfaces as exit code 137 (128 plus SIGKILL's signal number 9).

A layered root filesystem decides what files it starts with. The image's layers are stacked read-only, with a thin writable layer on top, typically using the kernel's overlay filesystem. Writes go to the top layer and disappear when the container is removed.

The consequence that matters most: every container on a host shares the host's kernel. There is no guest operating system. That is why containers start in milliseconds and why a kernel vulnerability, or a container granted too much privilege, is a host problem rather than a container problem.

docker CLIbuild, run, pushREST APIdockerdnetworks, volumesgRPCcontainerdimages, snapshotsshim (runc v2)one per containercreateruncOCI runtime, exitsclone + execContainer processPID 1 in its namespaceKernel: namespacespid net mnt uts ipc user cgroupKernel: cgroups v2memory, cpu, pids limitsRoot filesystemoverlay of image layers + rw layerRegistryOCI distributionpull
The Docker Engine stack. The CLI talks to dockerd, which delegates image and container lifecycle to containerd. A per-container shim invokes runc, which asks the kernel for namespaces, cgroups and the overlay root filesystem, starts the process and exits.

From docker run to a running process

The docker command is only a client. It sends REST calls to the daemon, dockerd, over a Unix socket. The daemon owns Docker-level concepts such as networks, volumes and the build API, and delegates the container lifecycle to containerd, which pulls and stores images and prepares snapshots of their filesystems. For each container, containerd starts a small shim process, and the shim calls runc, the reference implementation of the OCI runtime specification. runc reads a JSON bundle describing namespaces, cgroup limits, mounts and the command, sets everything up, starts the process and exits. The shim stays as the process's parent, which is why containers keep running when the daemon restarts, if live restore is enabled.

This layering is why the same image runs under Kubernetes without Docker installed: Kubernetes talks to containerd or CRI-O directly, and both end up calling an OCI runtime.

One security fact follows directly: the daemon runs as root and can start a container that mounts the host's root filesystem. Anyone who can talk to the Docker socket, including members of the docker group and any container you mount /var/run/docker.sock into, effectively has root on the host. Rootless mode, which runs the daemon and containers inside a user namespace, removes that equivalence at some cost in features and networking performance.

Advertisement

What an image is

An OCI image is three kinds of content-addressed objects. Each layer is a compressed tar archive of filesystem changes: files added or modified, and special whiteout entries marking deletions. The config is a JSON document with the default command, environment, working directory, user and the ordered list of layer digests. The manifest ties a config to its layers. For multi-architecture images, an index points at one manifest per platform, and the client picks the one matching its CPU.

Every object is named by the SHA-256 digest of its bytes. Tags such as postgres:16 are mutable pointers to a manifest digest, and they move when the publisher pushes a new build. If you need the same bytes tomorrow, refer to postgres:16@sha256:... with the full digest. Because layers are content-addressed, two images built on the same base share those layers on disk and in the registry, and a pull downloads only what is missing.

Whiteouts carry a trap. Deleting a file in a later layer hides it from the container but leaves it in the earlier layer, where anyone with the image can extract it. A secret copied in one instruction and removed in the next is still shipped. Secrets belong in build-time secret mounts or in the runtime environment, never in a layer.

Building images: instructions, layers and the cache

A Dockerfile is a recipe the builder executes top to bottom. Instructions that change the filesystem, mainly RUN, COPY and ADD, produce layers; the others set metadata. The builder, BuildKit on current Docker Engine releases, caches each step keyed on the instruction and its inputs. For COPY, the input is the content of the copied files. When a step's key changes, that step and every step after it rebuild.

So order instructions from least to most frequently changed. Here is a worked example for a Python API: dependencies are resolved in a separate build stage, and only the result is copied into a clean runtime stage.

# syntax=docker/dockerfile:1
FROM python:3.12-slim AS build
WORKDIR /app
# Dependencies first: this layer is reused until requirements.txt changes.
COPY requirements.txt .
RUN --mount=type=cache,target=/root/.cache/pip \
    pip wheel --wheel-dir /wheels -r requirements.txt

FROM python:3.12-slim
RUN useradd --create-home --uid 10001 app
WORKDIR /app
COPY --from=build /wheels /wheels
RUN pip install --no-index --find-links=/wheels /wheels/* && rm -rf /wheels
# Source last: editing code only rebuilds from here down.
COPY --chown=app:app src/ ./src/
USER app
EXPOSE 8000
# Exec form: uvicorn is PID 1 and receives SIGTERM directly.
ENTRYPOINT ["python", "-m", "uvicorn", "src.main:app", "--host", "0.0.0.0", "--port", "8000"]

Walk through what each choice buys. Copying requirements.txt alone before the source means an edit to application code reuses the dependency layer, so a rebuild takes seconds instead of minutes. The cache mount keeps pip's download cache across builds without putting it in a layer. The multi-stage build leaves compilers and build tools in the first stage, so the shipped image is smaller and carries fewer packages to patch. The numeric non-root user limits what a compromised process can do. The exec-form entrypoint makes the server itself PID 1, which matters for shutdown, covered below.

Two smaller habits complete the picture. A .dockerignore file keeps .git, local virtual environments and credentials out of the build context, which both speeds the build and stops COPY . . from baking secrets into a layer. And pinning the base image by digest, updated deliberately by a bot or a scheduled job, makes builds reproducible without letting them fall behind on security fixes.

Running containers: limits, networks and storage

Building is half of it. The run command decides what the process can reach and consume.

# Build with a tag, then inspect what you made
docker build -t orders-api:1.4.2 .
docker image history orders-api:1.4.2          # one row per layer, with sizes
docker image inspect orders-api:1.4.2 --format '{{.Id}} {{.Config.User}}'

# Run with explicit limits, a user-defined network and a named volume
docker network create shop
docker volume create pgdata
docker run -d --name db --network shop -v pgdata:/var/lib/postgresql/data \
  -e POSTGRES_PASSWORD_FILE=/run/secrets/pg -v "$PWD/pg.secret:/run/secrets/pg:ro" postgres:16
docker run -d --name api --network shop -p 127.0.0.1:8000:8000 \
  --memory 512m --cpus 1.0 --pids-limit 256 --read-only --tmpfs /tmp \
  orders-api:1.4.2

# Diagnose
docker logs -f api
docker exec -it api sh          # fails on images without a shell; use a debug sidecar instead
docker stats --no-stream
docker inspect api --format '{{.State.ExitCode}} {{.State.OOMKilled}}'

Limits. Without --memory, a container can use all the host's memory and take its neighbours down with it. Set memory, CPU and process limits on every long-running container, and check State.OOMKilled when one dies with 137.

Networks. Containers on a user-defined bridge network can reach each other by container name through Docker's embedded DNS; on the legacy default bridge they cannot. Publishing a port with -p 8000:8000 binds on every host interface, and on Linux Docker programs the packet filter directly, so host firewall front-ends such as ufw may not block it. Bind to 127.0.0.1 unless the port is meant to be public.

Storage. The writable layer is slow for heavy writes and dies with the container. Named volumes are managed by Docker and survive container replacement; bind mounts map a host path and are convenient in development but tie the container to the host's layout and permissions. Databases always get a volume.

PID 1, signals and clean shutdown

docker stop sends SIGTERM to the container's PID 1, waits a grace period (10 seconds by default) and then sends SIGKILL. Two things commonly break this. First, a shell-form command such as CMD python app.py runs under /bin/sh -c, so the shell is PID 1 and may not forward the signal to your program, which is then killed hard after the timeout, dropping in-flight requests. Use the exec form. Second, the kernel treats PID 1 specially: signals it has not installed a handler for are ignored, and it is expected to reap orphaned child processes. A program that spawns children should run under a minimal init, which docker run --init provides.

In the application, handle SIGTERM by stopping intake, finishing in-flight work within the grace period and exiting zero. That one handler is what makes rolling deploys lossless, here and later under an orchestrator.

Several containers: Compose

Real applications are several processes. Compose describes them, their networks and volumes in one file, and docker compose up creates them on a project network where service names resolve as hostnames.

services:
  api:
    build: .
    image: orders-api:dev
    ports: ["127.0.0.1:8000:8000"]
    environment:
      DATABASE_URL: postgresql://app@db:5432/orders
    depends_on:
      db:
        condition: service_healthy
  db:
    image: postgres:16
    environment:
      POSTGRES_USER: app
      POSTGRES_DB: orders
      POSTGRES_HOST_AUTH_METHOD: trust   # local development only
    volumes: ["pgdata:/var/lib/postgresql/data"]
    healthcheck:
      test: ["CMD", "pg_isready", "-U", "app"]
      interval: 5s
      retries: 10
volumes:
  pgdata:

Note the health check and condition: service_healthy: plain depends_on only orders container starts, and a database container is started long before it accepts connections. Even with the condition, write the application to retry its first connection, because the same race reappears whenever the database restarts. Compose is excellent for development, CI and single-host deployments; scheduling across machines, self-healing and rolling updates belong to an orchestrator.

Failure modes worth recognising

  • Exit 137 with no log line. The OOM killer, or SIGKILL after a stop timeout. Check OOMKilled and the memory limit, then the shutdown handling.
  • Slow builds after every edit. A broad COPY early in the file invalidates the cache. Move dependency installation above the source copy and add a .dockerignore.
  • Works on my laptop, fails in CI. A mutable tag resolved to a different image, or an ARM laptop pulled a different platform than the x86 runner. Pin digests and build with an explicit platform when it matters.
  • Permission denied on a volume. The container's numeric user id does not own the mounted host path. Ids, not names, are what the kernel compares.
  • Disk full on the host. Old images, stopped containers, build cache and unbounded JSON log files accumulate. Configure log rotation and prune on a schedule; docker system df shows where the space went.
  • Accidental exposure. A published port on all interfaces, a mounted Docker socket, or --privileged turn a container compromise into a host compromise.

Trade-offs: containers versus the alternatives

ConcernContainersVirtual machinesBare processes
Isolation boundaryShared kernel; namespaces and cgroupsSeparate kernel per guestUnix users only
Start timeMilliseconds to secondsSeconds to minutesImmediate
PackagingImage carries userland and dependenciesFull disk imageHost must provide dependencies
DensityHighLower: each guest has its own OSHighest, with no isolation
When to preferMost services and CI jobsUntrusted or multi-tenant code, other kernelsSimple single-purpose hosts

When you need container ergonomics with a stronger boundary, sandboxed runtimes such as gVisor or lightweight VMs such as Kata Containers plug in as alternative OCI runtimes. For packaging specifics on the JVM, see Java containerization; for managed hosting, containers in the cloud; and for building and pushing images automatically, CI/CD pipelines.

What to do next

  1. Run a container, find its process on the host with ps, and compare its PID and hostname from inside and outside. Seeing the namespaces makes the model stick.
  2. Take one service you own and rewrite its Dockerfile as two stages: dependencies before source, a non-root numeric user, an exec-form entrypoint, and a .dockerignore.
  3. Read docker image history for the result and remove any layer that carries credentials, caches or build tools.
  4. Add a SIGTERM handler, then measure a docker stop: it should exit well inside the grace period with no dropped requests.
  5. Set memory, CPU and pids limits on every long-running container, bind published ports to localhost unless they must be public, and pin base images by digest with an automated update.
  6. Describe the service and its database in a Compose file with a health check, and use it as the local and CI environment before moving to an orchestrator.
Key takeaway: A container is a normal Linux process given a private view through namespaces, bounded by cgroups and started on a stack of read-only image layers; Docker builds those images, moves them through registries and drives containerd and runc to start them. Hold that model and the practical rules follow: order Dockerfiles for the cache and keep secrets out of layers, set resource limits, run as a non-root user, make your program PID 1 and handle SIGTERM, and treat the Docker socket as root.