Kubernetes ships controllers for Deployments, Jobs and Services, but it knows nothing about how to run your database, cache or message broker: how to add a replica safely, which order to upgrade nodes in, what to clean up when the cluster is deleted. The operator pattern closes that gap. You add a custom resource that describes the thing you want, and a controller that keeps the real world matching it, using the same loop the built-in controllers use.

This article builds that up from first principles: the level-triggered control loop, how to design the custom resource's spec and status, a complete reconciler written with controller-runtime, deletion with owner references and finalizers, a traced worked example, the failure modes that bite in production, and how to run operators safely. It ends with a checklist.

The control loop from first principles

A controller compares desired state with observed state and acts to shrink the difference. The key word is level-triggered. The controller is not told "replicas changed from 3 to 5"; it is told "something about object default/orders-cache may have changed" and must re-read everything and work out what to do. If it misses an event, crashes halfway, or receives ten events at once, the next reconcile still converges, because it starts from the current state of the world rather than from a history of changes.

Three pieces make this efficient. An informer opens a watch on the API server and maintains a local, indexed cache of objects, so reads in the reconciler do not hit the API server. Event handlers turn each change into a key (namespace and name) and push it onto a workqueue, which deduplicates: a key that is already queued is not queued twice, and one key is never processed by two workers at the same time. Failed keys are re-added with exponential backoff by the queue's rate limiter. Finally, the reconcile function pops a key and does the work.

Changes to child objects are mapped back to the parent. When the StatefulSet that a Cache owns changes its ready count, the owner reference on the StatefulSet tells the controller which Cache key to enqueue. Periodic resyncs re-enqueue everything at an interval as a backstop against missed events. Taken together, this means the reconciler must be idempotent: running it twice in a row on unchanged state must change nothing the second time.

An operator is a control loop: watch, queue, reconcile, write, repeatAPI server + etcdCache CR, StatefulSetInformer cachewatch + local indexWorkqueuekeys, deduped, rate-limitedReconcile(key)read, diff, actwatchenqueuepopwrites: create/update children, patch status, add/remove finalizerWhat one Cache object ownsCache (custom resource)spec: replicas, versionStatefulSetownerReference to CacheServiceownerReference to Cacheownsownsstatus: observedGeneration, readyReplicas, conditions[Ready] finalizer: cache.example.com/deregister
The control loop. The informer cache feeds a deduplicating workqueue; the reconciler reads from the cache and writes to the API server, and its own writes come back as watch events.

Designing the custom resource

The custom resource definition (CRD) is your API, and it deserves the same care as any public API. Split it in two. The spec is written by users and expresses intent: replica count, version, memory size. The status is written only by the controller and reports what it observed. Enable the status subresource so the two have separate write paths: users cannot overwrite status, and the controller's status updates do not bump the object's generation.

The API server increments metadata.generation whenever the spec changes. The controller copies the generation it acted on into status.observedGeneration. A client that sees observedGeneration < generation knows the latest change has not been processed yet; without that field, a Ready condition is ambiguous because it might describe the previous spec. Report health as standard conditions, each with a type, status, reason, message and its own observed generation, so tools such as kubectl wait --for=condition=Ready work without knowing your schema.

Validate in the schema rather than in the reconciler. Bounds, enums, required fields and defaults declared through kubebuilder markers become OpenAPI validation in the CRD, so a bad object is rejected at admission instead of being accepted and then failing forever in a loop.

type CacheSpec struct {
	// +kubebuilder:validation:Minimum=1
	// +kubebuilder:validation:Maximum=9
	Replicas int32  `json:"replicas"`
	// +kubebuilder:validation:Pattern=`^\d+\.\d+\.\d+$`
	Version  string `json:"version"`
}

type CacheStatus struct {
	ObservedGeneration int64              `json:"observedGeneration,omitempty"`
	ReadyReplicas      int32              `json:"readyReplicas,omitempty"`
	Conditions         []metav1.Condition `json:"conditions,omitempty"`
}

// +kubebuilder:object:root=true
// +kubebuilder:subresource:status
type Cache struct {
	metav1.TypeMeta   `json:",inline"`
	metav1.ObjectMeta `json:"metadata,omitempty"`
	Spec   CacheSpec   `json:"spec,omitempty"`
	Status CacheStatus `json:"status,omitempty"`
}

Writing the reconciler

The reconciler below manages a StatefulSet for each Cache. It follows the shape every good reconciler has: fetch the object, handle deletion, ensure children match the spec, report status, and decide whether to come back later. Note what it does not do: it never edits the spec, never keeps state in memory between calls, and never assumes it knows why it was called.

const cacheFinalizer = "cache.example.com/deregister"

// +kubebuilder:rbac:groups=cache.example.com,resources=caches,verbs=get;list;watch;update;patch
// +kubebuilder:rbac:groups=cache.example.com,resources=caches/status,verbs=get;update;patch
// +kubebuilder:rbac:groups=cache.example.com,resources=caches/finalizers,verbs=update
// +kubebuilder:rbac:groups=apps,resources=statefulsets,verbs=get;list;watch;create;update;patch;delete

func (r *CacheReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error) {
	var cache cachev1.Cache
	if err := r.Get(ctx, req.NamespacedName, &cache); err != nil {
		return ctrl.Result{}, client.IgnoreNotFound(err) // already gone
	}

	if !cache.DeletionTimestamp.IsZero() {
		if controllerutil.ContainsFinalizer(&cache, cacheFinalizer) {
			if err := r.deregister(ctx, &cache); err != nil {
				return ctrl.Result{}, err // retried with backoff
			}
			controllerutil.RemoveFinalizer(&cache, cacheFinalizer)
			return ctrl.Result{}, r.Update(ctx, &cache)
		}
		return ctrl.Result{}, nil
	}
	if controllerutil.AddFinalizer(&cache, cacheFinalizer) {
		if err := r.Update(ctx, &cache); err != nil {
			return ctrl.Result{}, err
		}
	}

	sts := &appsv1.StatefulSet{ObjectMeta: metav1.ObjectMeta{
		Name: cache.Name, Namespace: cache.Namespace}}
	_, err := controllerutil.CreateOrUpdate(ctx, r.Client, sts, func() error {
		sts.Spec.Replicas = ptr.To(cache.Spec.Replicas)
		r.fillPodTemplate(sts, &cache) // image tag from cache.Spec.Version
		return controllerutil.SetControllerReference(&cache, sts, r.Scheme)
	})
	if err != nil {
		return ctrl.Result{}, err
	}

	before := cache.Status.DeepCopy()
	ready := sts.Status.ReadyReplicas == cache.Spec.Replicas
	cond := metav1.Condition{Type: "Ready", Status: metav1.ConditionFalse,
		Reason: "Progressing", ObservedGeneration: cache.Generation,
		Message: fmt.Sprintf("%d/%d replicas ready", sts.Status.ReadyReplicas, cache.Spec.Replicas)}
	if ready {
		cond.Status, cond.Reason = metav1.ConditionTrue, "AllReplicasReady"
	}
	meta.SetStatusCondition(&cache.Status.Conditions, cond)
	cache.Status.ObservedGeneration = cache.Generation
	cache.Status.ReadyReplicas = sts.Status.ReadyReplicas
	if equality.Semantic.DeepEqual(*before, cache.Status) {
		return ctrl.Result{}, nil // nothing changed: no write, no extra watch event
	}
	return ctrl.Result{}, r.Status().Update(ctx, &cache)
}

func (r *CacheReconciler) SetupWithManager(mgr ctrl.Manager) error {
	return ctrl.NewControllerManagedBy(mgr).
		For(&cachev1.Cache{}).
		Owns(&appsv1.StatefulSet{}).
		Complete(r)
}

Returning an error tells the workqueue to retry the key with backoff. Returning ctrl.Result{RequeueAfter: d} asks for a fresh look after a fixed delay, which is right for polling something outside Kubernetes that cannot send watch events. Status is written only when it differs from what was read, so an unchanged pass makes no write and triggers no further event. Here no requeue is needed while the StatefulSet is progressing, because Owns makes every change to the StatefulSet's status enqueue the Cache again.

CreateOrUpdate reads the child, runs the mutate function, and writes only if something changed, which is what keeps the reconciler idempotent. Set only the fields you own inside the mutate function; overwriting the whole spec fights with defaulting and with other controllers and produces an update on every pass.

Deletion: owner references and finalizers

Kubernetes offers two cleanup mechanisms, and operators need to know which is which. Owner references handle in-cluster children: when the Cache is deleted, the garbage collector deletes the StatefulSet and Service it owns. You write no deletion code. Owner references only work within a namespace (or from cluster-scoped owners), and they cannot clean up anything outside the cluster.

Finalizers handle everything else: a DNS record, a cloud load balancer, a registration in an external catalogue. A finalizer is a string in metadata.finalizers. When a user deletes an object that has one, the API server only sets deletionTimestamp; the object stays visible until every finalizer is removed. The reconciler sees the timestamp, performs external cleanup, and removes its finalizer, after which deletion completes.

The cleanup must be idempotent and must treat "already gone" as success, because the reconciler may crash after cleanup but before removing the finalizer and run the cleanup again. And it must eventually succeed or be bypassable: a finalizer whose controller has been uninstalled leaves objects stuck in Terminating, and the namespace containing them cannot be deleted either. Uninstall procedures should delete custom resources before removing the operator.

Worked example: scaling from 3 to 5

Trace a user scaling orders-cache from 3 to 5 replicas with kubectl patch.

  1. The API server stores the new spec and bumps generation from 4 to 5. The watch delivers the change; the key default/orders-cache is enqueued.
  2. Reconcile runs. The finalizer is present, so no update there. CreateOrUpdate sees replicas 3 in the StatefulSet, sets 5 and updates. Status is written with observedGeneration 5, readyReplicas 3 and Ready=False, reason Progressing, message "3/5 replicas ready".
  3. The status write is itself a watch event on the Cache, so the key is enqueued again. That reconcile finds nothing to change: the StatefulSet already says 5, and the recomputed condition and counts match what is stored. This empty pass is normal and is why idempotence matters.
  4. The StatefulSet controller creates pods 3 and 4 in order. Each time its ready count changes, the owner reference enqueues the Cache, and the reconciler records 4/5 and then 5/5 with Ready=True.
  5. kubectl wait --for=condition=Ready cache/orders-cache returns. A script that also checks observedGeneration equals generation knows the Ready refers to the 5-replica spec, not the old one.

Now suppose the operator pod is killed between steps 2 and 4. The new leader starts, its informers list every Cache, and each key is enqueued. The reconciler re-reads the world, finds the StatefulSet at 5 and pods coming up, and carries on. Nothing was lost because nothing important lived in memory.

Failure modes

  • Conflicts on update. Every object has a resourceVersion; an update based on a stale copy fails with a conflict. Return the error and let the queue retry with a fresh read. Do not loop and retry inside Reconcile.
  • Stale cache reads. The informer cache lags the API server. Right after creating a child, a read may not find it, and a naive reconciler creates a second one. Use deterministic child names so a duplicate create fails with AlreadyExists instead of producing two objects.
  • Hot loops. Writing status on every pass, with a timestamp or a reordered list, generates a watch event that enqueues the object again, forever. Write status only when it changed; meta.SetStatusCondition leaves the transition time alone when the status does not change.
  • Fighting another controller. If an autoscaler also sets replicas on the StatefulSet, the two will flip it back and forth. Decide which controller owns each field, and leave the other fields alone.
  • Blocking work in Reconcile. Waiting minutes for a backup holds a worker. Start the work, record it in status, and return with RequeueAfter to check progress.
  • Stuck finalizers. External cleanup that always fails leaves objects Terminating forever. Alert on objects whose deletionTimestamp is older than a threshold.
  • Missing RBAC. A forbidden error on a watch is easy to miss because the controller simply never sees events. Generate RBAC from markers and test the operator under its real service account.

Running operators in production

Run two or more replicas of the operator with leader election enabled, so only one reconciles at a time and a standby takes over when the lease expires. Leader election is built on a Lease object and is covered conceptually in the leader election article; the important property is that a deposed leader must stop working when it loses the lease, which controller-runtime's manager does by exiting.

Controller-runtime exposes Prometheus metrics for reconcile counts, errors, durations and workqueue depth. Alert on rising queue depth and on reconcile error rate per controller, not on individual errors, which are routine conflicts. Add your own condition-based alerts, such as Ready=False for longer than an expected rollout.

Treat the CRD schema as a versioned API. Additive optional fields are safe. Renaming or retyping a field needs a new version and either a conversion webhook or a migration, and stored objects stay in the old storage version until they are rewritten. Test upgrades by installing the old operator, creating objects, upgrading, and checking that every object reconciles cleanly. All of this state lives in etcd, so the etcd article explains the consistency guarantees behind watches and resource versions.

Trade-offs

Operators are not free. Each one is a long-running program with cluster-wide permissions that your team must build, test, upgrade and be paged for. If an application only needs templated manifests installed once, Helm or Kustomize is simpler and has no runtime. The operator earns its keep when day-two operations need judgement encoded in code: ordered upgrades, failover, backup and restore, rebalancing after scale-out.

There is also a choice of scope. A narrow operator that manages one resource kind with a small, honest status is easier to reason about than one that tries to automate everything a human operator ever did. Prefer composing existing controllers (let the StatefulSet controller manage pods) over reimplementing them. And remember that the reconcile loop is a general pattern: the same level-triggered, idempotent design shows up in infrastructure drift reconciliation, and the same backoff discipline in retry and backoff patterns.

What to do next

  1. Write down the day-two operations your application needs; if there are none, use Helm instead.
  2. Design the CRD with spec for intent, status for observation, conditions and observedGeneration, and validation in the schema.
  3. Scaffold with kubebuilder and write a reconciler that is idempotent and keeps no state in memory.
  4. Use owner references for in-cluster children and finalizers only for external cleanup, with idempotent cleanup code.
  5. Return errors to retry, use RequeueAfter only for external polling, and write status only when it changed.
  6. Test with envtest, including a crash between steps and a conflict on update.
  7. Deploy two replicas with leader election and alert on queue depth, error rate and long Terminating objects.
  8. Rehearse a CRD version upgrade before your first breaking schema change.
Key takeaway: An operator is a custom resource plus a level-triggered controller that repeatedly reads the world and nudges it towards the spec. Keep the reconciler idempotent and stateless, separate spec from status and report observedGeneration and conditions, use owner references for in-cluster children and finalizers only for external cleanup, let the workqueue handle retries, and run the operator with leader election and alerts on queue depth and stuck deletions.