DigitalOcean Kubernetes (DOKS) is DigitalOcean's managed Kubernetes service. DigitalOcean runs the control plane, worker nodes are ordinary Droplets grouped into node pools, and controllers inside the cluster turn Kubernetes objects into DigitalOcean load balancers and block storage volumes. It aims at teams that want conformant Kubernetes with fewer moving parts than the hyperscalers' offerings.

This article explains what DOKS manages and what stays your job, how to create and size a cluster, how networking, load balancers, storage and autoscaling behave, how upgrades replace nodes, the failure modes that catch teams, and when DOKS is the right choice. Specific numbers and flags were checked against DigitalOcean's documentation on 2026-10-04; managed services change, so confirm limits in the current docs before relying on them. Pricing is deliberately not quoted. For a cross-provider view see managed Kubernetes compared.

What DOKS manages and what you own

DigitalOcean operates the API server, etcd, scheduler, controller manager and the DigitalOcean cloud controller manager (CCM). It also manages the operating system, packages and system components on worker nodes, including the Cilium networking agent, CoreDNS and the CSI driver. You own everything you deploy: workloads, their resource requests, PodDisruptionBudgets, network policies, ingress controllers, observability and backups of your data.

Control plane high availability is version-dependent. The limits page states that clusters created on DOKS version 1.36.0 and later have HA enabled by default, with earlier versions running single-replica control plane components; the doctl reference still lists a --ha flag defaulting to false. Check what your cluster actually has with doctl kubernetes cluster get. A non-HA control plane means API calls can fail during control plane maintenance; running pods keep running, but deployments, autoscaling and anything that talks to the API pause.

Published limits at the time of writing: up to 110 pods per node, 250 ports per Service, and no IPv6 on nodes or clusters. Cluster size is quoted as up to 512 worker nodes on the limits page and up to 1,000 worker nodes with VPC-native networking on the features page; if you are planning near either number, ask DigitalOcean which applies to your configuration.

DOKS: managed control plane, worker Droplets in your VPC, cloud resources via controllersControl plane (DigitalOcean)API server, etcd, scheduler, CCMoptional HA + control plane firewallYour VPCNode pool: generals-4vcpu-8gb x 3-6, autoscaleNode pool: memorytainted for cachesNode pool: GPUdevice plugin, scale to 0each node: kubelet, containerd, Cilium (eBPF, kube-proxy replacement), CSI node pluginLoad Balancerfrom Service type=LoadBalancerVolumes (block)do-block-storage, RWOContainer Registrypull secrets per namespaceEdit LBs and volumes through Kubernetes objects; changes made in the control panel can be reconciled away
Figure 1. DOKS layout. DigitalOcean runs the control plane; worker nodes are Droplets in your VPC grouped into node pools; the cloud controller manager and CSI driver create load balancers and volumes from Kubernetes objects.

Creating a cluster from code

Create clusters from code, not the control panel, so they can be rebuilt. The doctl CLI is enough to start; Terraform's DigitalOcean provider is the usual next step. List valid values first, then create:

doctl kubernetes options versions
doctl kubernetes options sizes

doctl kubernetes cluster create prod-fra1 \
  --region fra1 \
  --vpc-uuid <vpc-uuid> \
  --ha \
  --surge-upgrade \
  --auto-upgrade \
  --maintenance-window "sunday=03:00" \
  --enable-control-plane-firewall \
  --control-plane-firewall-allowed-addresses "203.0.113.10/32" \
  --node-pool "name=general;size=s-4vcpu-8gb;auto-scale=true;min-nodes=3;max-nodes=6" \
  --tag prod

doctl kubernetes cluster node-pool create prod-fra1 --name cache \
  --size m-2vcpu-16gb --count 2 \
  --taint "workload=cache:NoSchedule" --label "pool=cache"

doctl kubernetes cluster kubeconfig save prod-fra1

Defaults worth knowing: the pod network defaults to 10.244.0.0/16 and the service network to 10.245.0.0/16; surge upgrades default to on; automatic upgrades default to off; the default node size is a small s-1vcpu-2gb-intel Droplet, which is too small for most real workloads. Pick the VPC deliberately, and make sure its range and the pod and service ranges do not overlap networks you will peer with later, because subnets cannot be changed after creation. Run doctl kubernetes cluster create --help on your installed version to confirm the flags, since newer releases add options.

Sizing node pools

A node's size is not what your pods get. System components (kubelet, containerd, Cilium, CoreDNS and others) reserve memory, and the DOKS limits page gives the allocatable figure per size; a 2 GiB node leaves about 1 GiB for pods. Always read kubectl describe node for the Allocatable section before sizing.

A worked example. A service needs 24 replicas, each requesting 0.5 vCPU and 1.5 GiB, plus about 20 percent headroom for rolling updates and one node failure. Suppose the 8 GiB node size you picked shows roughly 6.5 GiB and 3.8 vCPU allocatable (measure your own; these are illustrative). Memory fits 4 pods per node and CPU fits 7, so memory binds: 24 / 4 = 6 nodes, plus headroom gives 8. A hypothetical 16 GiB, 4 vCPU size that fits 7 pods by both CPU and memory would need 4 nodes, 5 with headroom: fewer, larger nodes mean less per-node overhead, but each node failure removes a larger share of capacity and surge upgrades replace bigger units.

Use separate node pools for workloads with different shapes: a general pool, a memory-heavy pool with a taint so only tolerating pods land there, and GPU pools with the NVIDIA or AMD device plugin enabled (--enable-nvidia-gpu-device-plugin or --enable-amd-gpu-device-plugin). Node pools are made of Droplets, so Droplet size families and account limits apply.

Networking and load balancers

Cilium provides pod networking and supports Kubernetes NetworkPolicy, and on VPC-native clusters DOKS uses Cilium's eBPF-based kube-proxy replacement, so Service routing is handled in eBPF rather than iptables rules. Pods can reach other resources in the VPC, such as managed databases, over private addresses. Start every namespace with a default-deny ingress policy and open what is needed; this matters more on a flat VPC where many services share a network.

A Service of type LoadBalancer makes the CCM create a DigitalOcean load balancer. Behaviour is set by annotations; the CCM documentation lists, among others:

apiVersion: v1
kind: Service
metadata:
  name: web
  annotations:
    service.beta.kubernetes.io/do-loadbalancer-name: "prod-web"
    service.beta.kubernetes.io/do-loadbalancer-type: "REGIONAL"
    service.beta.kubernetes.io/do-loadbalancer-protocol: "http"
    service.beta.kubernetes.io/do-loadbalancer-tls-ports: "443"
    service.beta.kubernetes.io/do-loadbalancer-certificate-id: "<certificate-id>"
    service.beta.kubernetes.io/do-loadbalancer-redirect-http-to-https: "true"
    service.beta.kubernetes.io/do-loadbalancer-healthcheck-protocol: "http"
    service.beta.kubernetes.io/do-loadbalancer-healthcheck-path: "/healthz"
    service.beta.kubernetes.io/do-loadbalancer-size-unit: "2"
spec:
  type: LoadBalancer
  selector: { app: web }
  ports:
    - { name: http,  port: 80,  targetPort: 8080 }
    - { name: https, port: 443, targetPort: 8080 }

Per the CCM documentation the default type is REGIONAL_NETWORK with REGIONAL as the alternative, the default protocol is TCP, the default health check path is /, the default size unit is 1, and the HTTP idle timeout defaults to 60 seconds. Set the name explicitly; the default is derived from the Service UID, which is unreadable in billing and dashboards. Most teams run one load balancer in front of an ingress controller rather than one per Service. Annotation values are strings, so quote booleans and numbers. General load-balancer design is covered in cloud load balancers.

Persistent storage

The DigitalOcean CSI driver provisions block storage volumes from PersistentVolumeClaims using the do-block-storage StorageClass. Facts that shape designs: block volumes support only ReadWriteOnce, so a volume attaches to one node at a time; sizes range from 1 GB to 10,000 GB; volumes can be expanded but never shrunk; and the default reclaim policy is Delete, which means deleting the PVC deletes the volume and its data. For shared ReadWriteMany storage DigitalOcean offers NFS shares with a separate CSI plugin (--enable-nfs-csi-plugin).

apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: do-block-retain
provisioner: dobs.csi.digitalocean.com
reclaimPolicy: Retain
allowVolumeExpansion: true

Create a Retain class like this for databases and anything you cannot regenerate, and confirm the provisioner name against the default class on your cluster with kubectl get storageclass -o yaml. Because RWO volumes follow one node, a StatefulSet pod that moves must wait for detach and attach; keep replica counts and PodDisruptionBudgets in mind during upgrades.

Autoscaling

Each node pool can run the cluster autoscaler with --auto-scale --min-nodes N --max-nodes M. The minimum can be zero, which suits GPU or batch pools, but keep at least one small fixed pool so cluster components always have somewhere to run. The maximum is bounded by your account's Droplet limit (25 by default, minus Droplets you already have), and a scale-up that hits that limit fails quietly from the workload's point of view: pods stay Pending. Request a higher limit before you need it. The create command also exposes --scale-down-unneeded-time and --scale-down-utilization-threshold to tune how eagerly nodes are removed.

The node autoscaler acts on pending pods and their requests, not CPU usage, so requests must be realistic. Pair it with a HorizontalPodAutoscaler for replica counts; the general patterns are in cloud autoscaling.

Upgrades and maintenance

DOKS releases patch versions within a minor version and supports a limited window of minor versions. Automatic upgrades apply only patch releases and non-breaking component updates, inside a 4-hour maintenance window you set; minor version upgrades are always started by you, with doctl kubernetes cluster get-upgrades and doctl kubernetes cluster upgrade.

Upgrades replace nodes rather than patching them in place. With surge upgrades, DOKS creates up to 10 new nodes at once, drains old ones pool by pool, and deletes them. Pods are evicted honouring PodDisruptionBudgets, but the documented eviction timeout is 15 minutes and the drain timeout is 30 minutes, after which pods are deleted anyway. A PDB that can never be satisfied (for example minAvailable: 1 on a single-replica Deployment) does not block the upgrade; it only delays it and then loses the pod abruptly. Run at least two replicas of anything that matters, set PDBs that allow one disruption, handle SIGTERM, and avoid local state on nodes, since every node is eventually replaced.

Private images from the registry

For private images, doctl kubernetes cluster registry add <cluster> integrates DigitalOcean Container Registry: it creates a docker-registry secret in every namespace, which you reference as an imagePullSecrets entry or attach to each namespace's default service account. Namespaces created later need the integration to have created their secret too; check before the first deploy to a new namespace, since a missing secret shows up as ImagePullBackOff.

Failure modes

  • Control panel edits reverted. Renaming or reconfiguring a load balancer or volume in the control panel can break its link to the cluster or be overwritten by the controller. Change annotations and manifests instead.
  • Data deleted with the PVC. The default Delete reclaim policy removes the volume. Use a Retain class for stateful data and back up with snapshots.
  • Pending pods at peak. The autoscaler hit max-nodes or the account Droplet limit. Alert on Pending pods older than a few minutes.
  • Single-replica outage during upgrades. Every node is replaced; one replica means downtime. Two replicas and a PDB fix it.
  • Overlapping CIDRs. Pod or service ranges collide with a peered network or VPN. Plan ranges before creating the cluster.
  • API unavailable during maintenance. Without HA, deploys and autoscaling pause. Use HA for production clusters and schedule the window off-peak.
  • Undersized defaults. The default node size leaves little allocatable memory, so system pods and evictions crowd out workloads. Size nodes from measured allocatable.

Trade-offs

DOKS trades breadth for simplicity. Against EKS, GKE and AKS (see AKS in depth) you get a smaller set of knobs, integrated load balancers and volumes, and a short path from zero to a working cluster, but fewer managed add-ons, fewer instance types and regions, no IPv6, and a smaller ecosystem of managed services to integrate with. Compared with DigitalOcean's own App Platform, DOKS gives full Kubernetes control (custom controllers, sidecars, StatefulSets, GPU pools) and in exchange makes you own upgrades readiness, ingress, policies and observability. Choose DOKS when your team already runs Kubernetes and wants a simpler provider; choose App Platform when you only need to run stateless services from a repo.

What to do next

  1. Write the cluster definition as code (doctl script or Terraform) with region, VPC, CIDRs, HA, surge and auto-upgrade settings and a maintenance window.
  2. Measure allocatable CPU and memory per node size and size each node pool from your pods' requests plus failure headroom.
  3. Create a Retain StorageClass for stateful data and confirm snapshot backups restore.
  4. Put one load balancer with an explicit name in front of an ingress controller, with health checks on a real readiness path.
  5. Set two or more replicas and a PodDisruptionBudget for every production Deployment, then run a test minor upgrade on a staging cluster.
  6. Request a Droplet limit above your autoscaler maximum and alert on long-Pending pods and on LB health.
Key takeaway: DOKS runs the Kubernetes control plane for you while worker nodes are Droplets in your VPC, and controllers turn Services and PersistentVolumeClaims into DigitalOcean load balancers and volumes. Define clusters as code, size pools from measured allocatable resources, use a Retain storage class for data, manage load balancers only through annotations, and run two replicas with PodDisruptionBudgets because every upgrade replaces every node.