Oracle Kubernetes Engine, still called Container Engine for Kubernetes in some older documentation and CLI help and abbreviated OKE everywhere, is Oracle Cloud Infrastructure's managed Kubernetes service. Like every managed Kubernetes offering, it runs the control plane for you and leaves the nodes, network design, identity and upgrades as shared work. What makes OKE distinctive is the detail: two cluster tiers with different feature sets, three kinds of nodes, a choice between an overlay network and pods that get real VCN addresses, and an identity model built on OCI IAM policies.
This page is for an engineer who has to build or operate an OKE cluster and wants the irreversible decisions right the first time: cluster and node types, subnet and pod address planning (with a worked sizing example), CLI creation, load balancers, keyless pod permissions and upgrades. For a side-by-side view of EKS, GKE and AKS, see the managed Kubernetes comparison; this page stays on OKE.
What OKE manages, and what it leaves to you
OKE runs the Kubernetes control plane: the API server, etcd, the scheduler and the controller manager, replicated and patched by Oracle in an Oracle-owned tenancy. You never see those machines. You interact with the cluster through its API endpoint, which OKE places in a subnet of your Virtual Cloud Network (VCN) and which can be private (reachable only from inside the VCN or through a bastion) or public.
Everything else is yours: the VCN and subnets, node shapes and placement, the CNI plugin, security rules, Kubernetes version and upgrade timing, and how workloads authenticate to OCI. OKE also integrates the OCI cloud controller manager, so a Kubernetes Service of type LoadBalancer creates an OCI load balancer, and the OCI CSI drivers, so a PersistentVolumeClaim can create a block volume.
Basic versus enhanced clusters
Every OKE cluster is either basic or enhanced, and the choice gates features. According to Oracle's documentation on working with enhanced and basic clusters, enhanced clusters support all available features, while basic clusters do not support virtual node pools, self-managed nodes, workload identity, flexible cluster add-on management, or requests to raise the limits on clusters or nodes per cluster. Basic clusters carry no control-plane charge; enhanced clusters have an hourly per-cluster charge and a financially backed control-plane SLA. Prices change, so take current figures from Oracle's Kubernetes Engine pricing page rather than from this article or a third-party calculator.
In practice, use enhanced clusters for anything in production. Workload identity alone justifies it, because without it pods inherit the permissions of the node they run on. Basic clusters are reasonable for short-lived experiments and training environments. A basic cluster can be upgraded to enhanced in place, but the reverse is not a normal path, so if you are unsure, start enhanced.
Node types: managed, self-managed and virtual
Managed nodes are OCI compute instances that OKE creates and manages as members of a node pool. You choose the shape (for example a flexible VM shape with a set number of OCPUs and memory, a bare-metal shape, or a GPU shape), the image, the placement across availability domains and fault domains, and the pool size. OKE handles joining nodes to the cluster and can replace them during upgrades. This is the default and the right choice for most workloads.
Self-managed nodes are compute instances you create yourself and join to an enhanced cluster, usually because you need a configuration a node pool cannot express, such as a specific cluster-network setup for tightly coupled GPU training. You take on their life cycle, including upgrades.
Virtual nodes remove node management altogether: pods are scheduled onto Oracle-managed capacity and billed by the resources the pods use. They suit bursty stateless services and batch jobs. The trade-off is that features which depend on access to the host do not behave as on a normal node; workloads such as DaemonSets, privileged containers and host-path volumes need checking against Oracle's current list of virtual-node limitations before you move them. Virtual nodes use VCN-native pod networking.
Pod networking: flannel overlay or VCN-native
OKE offers two CNI plugins, and the choice is effectively permanent for a cluster, so make it deliberately.
With the flannel CNI, pods get addresses from a private overlay range (a pod CIDR such as 10.244.0.0/16 that you set at creation) and traffic between nodes is encapsulated. The VCN sees only node addresses. This uses very few VCN addresses and is simple, but OCI security rules cannot tell one pod from another, and anything outside the cluster sees pod traffic as coming from the node.
With the OCI VCN-native pod networking CNI, every pod gets a real private IP address from a dedicated pod subnet, attached through secondary VNICs on the worker node. Pods are first-class VCN citizens: network security groups can apply to pod traffic, peered networks can reach pods directly, and load balancers can send traffic straight to pod addresses. The cost is address consumption and a hard limit on pods per node set by the shape's VNIC count. Oracle's documentation for the commercial service gives the formula:
max pods per node = MIN( (number of VNICs on the shape - 1) * 31, 256 )The first VNIC serves the node itself; each additional VNIC serves up to 31 pods. Larger shapes support more VNICs, and the VNIC count of a flexible shape usually grows with its OCPU count, so check the shape's VNIC limit, not only its CPU and memory, when sizing node pools. You can also set a lower maximum pods per node in the node pool, which reserves fewer addresses per node. Oracle recommends network security groups rather than security lists when pod and node traffic need distinct rules.
Worked example: sizing the pod subnet
Suppose a production cluster will run up to 20 worker nodes, each on a shape that supports 4 VNICs. Each node can then hold MIN((4 - 1) * 31, 256) = 93 pods. If you let each node use its maximum, the pod subnet must hold 20 * 93 = 1,860 addresses.
Now add headroom for upgrades. Node pool upgrades and node cycling typically bring up replacement nodes before removing old ones, so plan for a surge of, say, 5 extra nodes: 25 * 93 = 2,325 addresses. Add room for growth, say to 30 nodes in a year: 2,790 plus surge. OCI reserves three addresses in every subnet. A /21 gives 2,048 addresses, which is too small even before growth; a /20 gives 4,096 and covers the plan with margin. The worker node subnet, by contrast, needs only one address per node, so a /24 is ample.
If the VCN cannot spare a /20, reduce the maximum pods per node to what the workloads actually need, for example 40, which brings 35 nodes to 1,400 addresses and fits a /21. Do this arithmetic before creation: changing a pod subnet afterwards means building new node pools and migrating workloads.
Creating a cluster from the CLI
The console's quick-create workflow builds a VCN and subnets for you, which is fine for a trial. For anything repeatable, create the network with Terraform or the CLI and then create the cluster into it. The shape of the CLI calls is shown below; the OCIDs are placeholders, and the exact flags for pod networking options and node sources should be taken from oci ce cluster create --help and oci ce node-pool create --help for your CLI version.
# 1. Create an enhanced cluster in an existing VCN, with a private API endpoint.
oci ce cluster create \
--compartment-id "$COMPARTMENT" \
--name prod-oke \
--vcn-id "$VCN" \
--kubernetes-version "$K8S_VERSION" \
--type ENHANCED_CLUSTER \
--endpoint-subnet-id "$API_SUBNET" \
--endpoint-public-ip-enabled false \
--service-lb-subnet-ids "[\"$LB_SUBNET\"]"
# 2. Write a kubeconfig that reaches the private endpoint (run from inside the VCN or via bastion).
oci ce cluster create-kubeconfig \
--cluster-id "$CLUSTER" \
--file "$HOME/.kube/config" \
--region "$REGION" \
--token-version 2.0.0 \
--kube-endpoint PRIVATE_ENDPOINT
kubectl get nodesThe kubeconfig does not hold a long-lived credential. It invokes the OCI CLI to mint a short-lived token on each call, which means kubectl access is governed by OCI IAM: the user needs IAM permission on the cluster, and Kubernetes RBAC then decides what that user can do inside it. Bind RBAC roles to OCI user or group OCIDs rather than handing out cluster-admin.
Load balancers from Services
A Service of type LoadBalancer makes the OCI cloud controller manager create either an OCI Load Balancer (layer 7 capable, with a flexible bandwidth shape) or a Network Load Balancer (layer 4, pass-through). Annotations choose which and configure it:
apiVersion: v1
kind: Service
metadata:
name: web
annotations:
oci.oraclecloud.com/load-balancer-type: "lb" # or "nlb"
service.beta.kubernetes.io/oci-load-balancer-shape: "flexible"
service.beta.kubernetes.io/oci-load-balancer-shape-flex-min: "10"
service.beta.kubernetes.io/oci-load-balancer-shape-flex-max: "100"
spec:
type: LoadBalancer
selector: {app: web}
ports:
- port: 443
targetPort: 8443Every such Service creates a billable OCI resource with its own public or private address, so a cluster with many small Services gets expensive and hits limits. Use one ingress controller behind one load balancer for HTTP workloads, and reserve dedicated load balancers for non-HTTP protocols. Remember that the load balancer subnet's security rules and the node or pod NSGs must both allow the health-check and traffic paths, or the load balancer reports every backend as unhealthy.
Identity: workload identity instead of node permissions
Pods that call OCI APIs, for example to read Object Storage or fetch secrets from Vault, need credentials. The old pattern was instance principals: put the nodes in a dynamic group and grant that group permissions. Every pod on those nodes then has the same permissions, which breaks least privilege.
Workload identity, available only on enhanced clusters, grants permissions to a Kubernetes service account in a namespace of a specific cluster. The IAM policy uses conditions on the request principal. This example follows the form in Oracle's documentation:
Allow any-user to read objects in compartment app-data where all {
request.principal.type = 'workload',
request.principal.namespace = 'payments',
request.principal.service_account = 'invoice-reader',
request.principal.cluster_id = 'ocid1.cluster.oc1..<placeholder>'}Inside the pod, the OCI SDK authenticates with the workload's projected service-account token. In Java the provider class is OkeWorkloadIdentityAuthenticationDetailsProvider; Oracle's documentation lists minimum SDK versions for each language (for Python, 2.111.0 or later), and the signer to use in other languages should be taken from that SDK's reference. Once this works, remove the node dynamic-group policies so that pods cannot fall back to node permissions.
Storage, upgrades and day-two operations
For persistent volumes, the OCI block volume CSI driver provisions a block volume per claim; block volumes are attached to one node at a time and live in one availability domain, so a pod using one can only reschedule to nodes in that domain. For shared read-write storage, the File Storage service CSI driver provides NFS-style volumes. Spread node pools across availability or fault domains, and use topology-aware scheduling so stateful pods land where their volumes are.
Upgrades happen in two steps. First you upgrade the control plane to a newer supported Kubernetes version; Oracle does this without downtime to the API in normal conditions. Then you bring node pools up to the same version, either by creating a new pool and draining the old one or by using node cycling, which replaces nodes in an existing pool with controlled surge and unavailability. Before either step, check for removed Kubernetes APIs in your manifests and make sure every important workload has a PodDisruptionBudget and more than one replica, or the drain will either stall or cause an outage. OKE supports a limited window of Kubernetes versions, so plan an upgrade cadence rather than waiting to be forced.
Failure modes seen in practice
- Pods stuck in ContainerCreating with VCN-native networking. The pod subnet is out of addresses or the node hit its VNIC-derived pod limit. Fix: size the subnet as in the worked example and set realistic maximum pods per node.
- Load balancer with all backends unhealthy. Security rules allow client traffic but not health checks to nodes or pods. Fix: add NSG rules for the health-check path from the load balancer subnet.
- kubectl cannot reach a private endpoint. The workstation is outside the VCN. Fix: use the Bastion service or a VPN, and generate kubeconfig with the private endpoint option.
- Pods with too much power. A node dynamic group has broad policies and every pod inherits them. Fix: migrate to workload identity on an enhanced cluster and delete the node policies.
- Upgrade drains that never finish. A PodDisruptionBudget allows zero disruptions or a single-replica Deployment blocks eviction. Fix: audit budgets and replicas before upgrade day.
- Stateful pods that will not reschedule. The block volume lives in a different availability domain from the remaining nodes. Fix: keep node capacity in every domain that holds volumes.
What to do next
- Choose an enhanced cluster for production and note which features you depend on (workload identity, virtual nodes, add-ons).
- Pick the CNI: VCN-native if pods need VCN-level addressing and security rules, flannel if addresses are scarce and pods only talk inside the cluster.
- Size the pod and node subnets with the formula and a surge allowance, and define NSGs for endpoint, nodes, pods and load balancers.
- Create the cluster with a private API endpoint from code (Terraform or CLI), and give engineers access through IAM plus RBAC bindings.
- Move every pod that calls OCI to workload identity and remove node-level policies.
- Write the upgrade runbook now: version check, PodDisruptionBudgets, node cycling settings, and a rehearsal on a staging cluster.