Kubernetes is easiest to learn if you ignore its feature list and start with one idea: you tell it what should be true, and a set of programs works continuously to make it true. You write down that three copies of a container should exist, and when a machine dies at three in the morning, a program notices that only two exist and starts another. Almost everything else in Kubernetes is a consequence of that idea.
This article builds the mental model from first principles. It explains what each component does when you run kubectl apply, introduces the handful of objects you need on the first day, walks through a complete deployment of a small web service, and then gives a triage table for the failures every beginner meets. It is about Kubernetes itself rather than any one cloud: the differences between managed offerings are covered in managed Kubernetes across clouds, and Google's version in GKE architecture.
The problem it solves
Containers package an application with its dependencies, so the same image runs on a laptop and on a server. Running one container on one machine is easy. Running forty services, each with several copies, across twenty machines is not. Someone has to decide which machine runs which copy, restart copies that crash, move them when a machine fails, route traffic only to copies that are ready, roll out new versions without downtime and hand each copy its configuration and secrets.
Before Kubernetes, that someone was usually a set of scripts that ran once and assumed success, and the real state of the system drifted away from anyone's idea of it. Kubernetes replaces them with declared state plus control loops: the desired state is stored in a database, and programs called controllers compare it with what is running and act to close the gap, forever.
# The shape of every Kubernetes controller (pseudocode)
loop forever:
desired = read spec of my objects from the API server (via a watch)
observed = read status of the things I manage
diff = desired - observed
for each difference:
take one step towards desired # create a Pod, delete a Pod, update a route
write what I observed into status
# never assume the step worked: the next pass will see whether it didTwo properties follow. First, it is eventually consistent: kubectl apply returns as soon as the desired state is stored, not when the containers are running, which is why kubectl rollout status exists. Second, it is self-healing only for things that are declared. A container you started by hand on a node is invisible to the loops and will not be restarted by anything.
The architecture, component by component
A cluster has a control plane, which stores state and makes decisions, and worker nodes, which run your containers. On managed services the cloud provider runs the control plane for you, but the components are the same.
| Component | Where | What it does |
|---|---|---|
| API server | Control plane | The only entry point. Authenticates, authorizes (RBAC), runs admission checks, stores objects. Every other component is its client. |
| etcd | Control plane | Consistent key-value store holding every object. Back it up: losing it loses the cluster's definition. |
| Scheduler | Control plane | Picks a node for each unscheduled Pod: filters out nodes that cannot fit it, scores the rest, writes the choice. |
| Controller manager | Control plane | Runs the built-in control loops: Deployments create ReplicaSets, ReplicaSets create Pods, dead nodes get marked. |
| kubelet | Every node | Starts the Pods bound to its node through the runtime, runs probes, reports status. |
| Container runtime | Every node | containerd or CRI-O, behind the Container Runtime Interface. |
| kube-proxy and the CNI plugin | Every node | The CNI plugin gives each Pod a routable IP; kube-proxy (or a CNI replacing it) maps Service addresses to Pod addresses. |
The first five objects
Kubernetes has dozens of object types, but a first deployment needs five, and they all hang together through labels: key-value pairs attached to objects, and selectors that match them.
A Pod is the smallest unit that runs: one or more containers that share a network address and can share volumes. Pods are disposable, get a new IP when recreated, and are rarely created directly. A Deployment declares how many copies of a Pod template should run and manages a ReplicaSet per version of that template, which is how rolling updates and rollbacks work. A Service gives a stable name and virtual IP to whatever Pods match its selector, and only sends traffic to Pods that report ready. A ConfigMap holds non-secret configuration and a Secret holds credentials; both can be mounted as files or exposed as environment variables. Note that Secrets are only base64-encoded by default, so anyone who can read them through the API can read the values; protect them with RBAC and, where your platform supports it, encryption at rest. A Namespace groups objects so that names, permissions and quotas can be scoped per team or environment.
What happens on kubectl apply
Following one request end to end is the fastest way to make the architecture stick. Suppose you apply a Deployment with three replicas.
- kubectl sends the object to the API server, which authenticates you, checks that RBAC lets you create Deployments in that namespace, runs admission control (which may add defaults or reject the object, for example for missing resource requests), and writes it to etcd. kubectl now returns. Nothing is running yet.
- The Deployment controller, watching Deployments, sees a new one with no ReplicaSet and creates a ReplicaSet for the current Pod template.
- The ReplicaSet controller sees a ReplicaSet that wants three Pods and owns zero, and creates three Pod objects. They have no node yet, so their phase is Pending.
- The scheduler sees three unscheduled Pods, finds nodes with enough unreserved CPU and memory to satisfy each Pod's requests, and writes a node name into each Pod.
- On each chosen node, the kubelet sees a Pod bound to it, asks the runtime to pull the image and start the container, and begins running the probes.
- When the readiness probe passes, the kubelet marks the Pod ready. The EndpointSlice controller adds its IP to the Service's EndpointSlices, and kube-proxy or the CNI updates routing so traffic reaches it.
Every step is a separate loop talking to the API server. When something does not work, ask which loop stopped making progress; kubectl describe usually tells you, because each component records Events against the object it is acting on.
Worked example: deploying a small web service
Here is a complete, minimal deployment of an HTTP service called orders: three replicas, resource requests and a memory limit, readiness and liveness probes, and a Service in front.
apiVersion: apps/v1
kind: Deployment
metadata:
name: orders
labels: {app: orders}
spec:
replicas: 3
selector:
matchLabels: {app: orders}
template:
metadata:
labels: {app: orders}
spec:
containers:
- name: api
image: registry.example.com/orders:1.4.2 # pin a tag or digest, never :latest
ports:
- containerPort: 8080
resources:
requests: {cpu: 250m, memory: 256Mi} # what the scheduler reserves
limits: {memory: 512Mi} # exceed this and the kernel kills it
readinessProbe:
httpGet: {path: /ready, port: 8080}
periodSeconds: 5
livenessProbe:
httpGet: {path: /healthz, port: 8080}
periodSeconds: 10
failureThreshold: 3
---
apiVersion: v1
kind: Service
metadata:
name: orders
spec:
selector: {app: orders} # must match the Pod labels exactly
ports:
- port: 80
targetPort: 8080Apply it and watch it converge; these commands answer most beginner questions about a running workload.
kubectl apply -f orders.yaml
kubectl rollout status deployment/orders # wait for ready Pods
kubectl get pods -l app=orders -o wide # node, IP, restarts
kubectl describe pod <pod-name> # read the Events
kubectl logs <pod-name> --previous # last crash
kubectl get endpointslices -l kubernetes.io/service-name=orders
kubectl rollout undo deployment/orders # previous ReplicaSetNow change the image to orders:1.4.3 and apply again. The Deployment controller creates a second ReplicaSet and, under the default RollingUpdate strategy, scales it up while scaling the old one down. The Service routes only to ready Pods, so users see no errors if the readiness probe tells the truth. If the new version never becomes ready, the rollout stalls instead of taking the service down, and kubectl rollout undo returns to the previous ReplicaSet, which is still recorded in the Deployment's history. Application-specific concerns such as graceful drain and autoscaling signals for agent workloads are covered in running ADK Java agents on Kubernetes.
Requests, limits and what the kernel does with them
Resource settings are where most production surprises start, because the two numbers do different jobs. A request is a reservation used by the scheduler: a node is considered full when the sum of requests on it reaches its allocatable capacity, regardless of how much the containers actually use. A limit is enforced at run time by the Linux kernel through cgroups.
CPU and memory behave differently when a limit is hit. CPU is compressible: a container that exceeds its CPU limit is throttled, so it runs slower but keeps running, and the symptom is latency, not failure. Memory is not compressible: a container that exceeds its memory limit is killed by the kernel's out-of-memory killer, and Kubernetes reports the container as OOMKilled and restarts it. That is why the example sets a memory limit but no CPU limit, a common choice that avoids throttling while still protecting the node; teams that need strict isolation set both.
Requests also decide eviction order when a node runs short of memory. Pods whose requests equal their limits for every resource are Guaranteed, Pods with some requests are Burstable, and Pods with none are BestEffort and go first. Omitting requests is therefore not neutral: the scheduler packs blindly and those Pods are the first casualties.
Probes: telling Kubernetes the truth
Kubernetes knows nothing about your application except what probes tell it. A readiness probe answers "should this Pod receive traffic right now?"; failing it removes the Pod from Service endpoints but does not restart it. A liveness probe answers "is this process stuck beyond recovery?"; failing it repeatedly makes the kubelet kill and restart the container. A startup probe holds off the other two until a slow-starting application has finished initialising.
The classic mistake is a liveness probe that checks the database: a brief database outage fails every Pod's liveness at once, every Pod restarts, and a blip becomes a full outage. Check dependencies in readiness if anywhere, keep liveness to "this process can still answer", and give liveness generous thresholds so a long garbage-collection pause does not trigger a restart.
Triage: why is my Pod not running?
Almost every beginner failure falls into one of a few states. The status column of kubectl get pods names the state and the Events section of kubectl describe pod usually names the cause.
| Symptom | What it means | Usual causes and first check |
|---|---|---|
| Pending | No node has been chosen, or the chosen node cannot start it | Requests larger than any node's free capacity, a taint with no matching toleration, an unbound volume. Read the scheduler's FailedScheduling event. |
| ImagePullBackOff or ErrImagePull | The kubelet cannot fetch the image | Typo in the image or tag, a private registry without an image pull secret, or a registry rate limit. Try pulling the exact reference yourself. |
| CrashLoopBackOff | The container starts and exits, and restarts are being delayed with growing back-off | Missing configuration, a bad command, a failing migration, or a liveness probe killing it. Run kubectl logs with --previous to see the last crash. |
| OOMKilled in the last state | The kernel killed the container for exceeding its memory limit | The limit is below real peak usage, or a JVM heap setting leaves no room for native memory under the limit. Measure usage, then raise the limit or lower the heap. |
| Running but 0/1 ready | The readiness probe is failing | Wrong path or port in the probe, or a dependency the readiness check waits on. Call the endpoint from inside the Pod. |
| Service returns connection refused or times out | The Service has no endpoints | The selector does not match the Pod labels, or no Pod is ready. Check the EndpointSlices for the Service. |
Trade-offs and when not to use it
Kubernetes costs real attention: regular upgrades because each minor version is supported for a limited window, networking and storage plugins, RBAC design, and a large YAML surface where a selector typo fails silently. Managed control planes remove the hardest operations but not the learning curve.
For a handful of stateless HTTP services, a serverless container platform is usually cheaper to run and easier to operate; cloud container services below Kubernetes compares those options. Kubernetes earns its cost when you have many services and teams, need consistent deployment and policy across them, run mixed workloads such as batch jobs and data processing alongside services (see Spark on Kubernetes), or need portability across clouds and on-premises hardware. For databases, a managed service is the safer first choice.
What to do next
- Create a local cluster with kind or minikube and apply the orders example with any small HTTP image.
- Delete one Pod and watch the ReplicaSet replace it; then cordon and drain a node in a multi-node kind cluster and watch Pods move.
- Break things on purpose: a wrong image tag, a selector typo in the Service, a memory limit of 16Mi, and a liveness probe on a path that does not exist. Diagnose each with describe, logs --previous and EndpointSlices.
- Roll out a new image, then roll it back with rollout undo, and read the Deployment's revision history.
- Set requests and a memory limit on every container you deploy, and decide deliberately whether you want CPU limits.
- Write readiness and liveness endpoints into your application that report what each probe is supposed to mean, and keep dependency checks out of liveness.
- Then move to a managed cluster, after reading your provider's upgrade policy.