A normal Compute Engine VM shares its physical server with VMs from other customers, and Google moves it between servers during maintenance without asking. For most workloads that is exactly right. For some it is a problem: a compliance rule says workloads must run on hardware dedicated to the organisation, a software licence is priced per physical core or socket, or a team wants to pack many small VMs onto a server it pays for as a unit. Sole-tenant nodes solve these by giving one project, or a set of projects you choose, exclusive use of whole physical servers.
This article explains the resource model, how VMs are placed with affinity labels, the three host maintenance policies and why they matter for licensing, CPU overcommit, and how billing works, then works through a SQL Server estate with per-core licences. Facts and flags come from the Compute Engine documentation and pricing page as read on 2026-10-03.
Sole tenancy in one picture
What sole tenancy gives you, and what it does not
A sole-tenant node is a physical Compute Engine server dedicated to hosting VMs from your project, or from the projects or organisation you share its node group with. It gives you three things. Physical separation: no VM from an unrelated project runs on the same host, which some compliance regimes require and which removes a class of cross-tenant side-channel concerns. Hardware affinity: you can keep certain VMs together on one server, or deliberately apart. Licence control: you can bound the number of physical cores and servers your VMs ever touch, which matters for licences counted per physical core or processor.
It does not give you encryption of data in use, a different hypervisor, or access to the host. If the requirement is protection from the cloud operator rather than from other tenants, look at Confidential VMs; the two answer different threat models and can be combined only where a machine series supports both.
Node types, templates and groups
Three resources describe sole tenancy. A node type names the server shape, for example n2-node-80-640 (80 vCPUs, 640 GB), n2d-node-224-896, c3-node-176-352, m3-node-128-1952 or GPU types such as a3-highgpu-node-208-1872-lssd. VMs on a node must belong to the same machine series as the node type: an n2 node hosts only n2 VMs.
A node template is a regional resource that fixes the node type and optional custom affinity labels, GPUs, Local SSDs and CPU overcommit. Its properties are copied immutably into every node, so changing a template means creating a new one and a new group. A node group is a zonal set of identical nodes made from one template, with a target size, a host maintenance policy, an optional maintenance window and an autoscaler mode of off, on or only-scale-out.
Some things do not work on sole-tenant nodes. The docs list T2D, T2A, E2, C2D, A3 Ultra, A3 Edge, A4, A4X and bare metal as unsupported; sole-tenant VMs cannot specify a minimum CPU platform; and preemptible VMs are not supported. Node groups are zonal, so high availability means groups in at least two zones with the application replicated across them.
Placing VMs with node affinity
A VM asks for sole tenancy through node affinity. Every node carries three default labels: compute.googleapis.com/node-group-name, compute.googleapis.com/node-name and compute.googleapis.com/projects. Custom labels such as workload=sql or environment=prod can only be set on the template, and you cannot add custom labels to nodes after the group exists. A VM's affinity is a list of rules, each with a key, an operator IN or NOT_IN, and values; all rules must match.
# node-affinity-sql-prod.json: production SQL VMs on prod SQL nodes only
[
{"key": "workload", "operator": "IN", "values": ["sql"]},
{"key": "environment", "operator": "IN", "values": ["prod"]}
]
# node-affinity-dev.json: development VMs anywhere except prod nodes
[
{"key": "workload", "operator": "IN", "values": ["sql"]},
{"key": "environment", "operator": "NOT_IN", "values": ["prod"]}
]Prefer group-level or custom labels to node-name. Pinning a VM to a named node is brittle: if that node fails and is replaced, or the group is resized, the VM has nowhere to go. Affinity to a label lets the scheduler choose any matching node with capacity.
Host maintenance policies
Google still has to patch and repair hosts. Host events happen roughly every four to six weeks, and the node group's host maintenance policy decides what your VMs experience.
| Policy (gcloud value) | What happens at a host event | Use it when |
|---|---|---|
Default (default) | VMs live migrate, as a group, to a different sole-tenant node; VMs that cannot live migrate are restarted on a fresh node | You need isolation but not a fixed set of physical servers |
Restart in place (restart-in-place) | VMs stop and restart on the same physical server after the event; VMs must use onHostMaintenance=TERMINATE | You need physical-server affinity and can tolerate around an hour of downtime per event |
Migrate within node group (migrate-within-node-group) | VMs live migrate only among a fixed set of servers, including reserved holdback nodes | Per-core licences plus high availability |
The migrate-within-node-group policy reserves holdback capacity: a group needs at least 2 nodes; 2 to 20 nodes reserve 1 holdback, 21 to 40 reserve 2, and so on up to 5 for 81 to 100. That holdback is the price of live migration that never leaves your licensed servers.
You can also give a group a maintenance window, a 4-hour block starting at 00:00, 04:00, 08:00, 12:00, 16:00 or 20:00 GMT, which controls when maintenance may begin. It cannot be changed after the group is created, maintenance is not guaranteed to finish inside it, and it is not supported with migrate-within-node-group. You can simulate a host maintenance event to test the policy before it happens for real. Hardware failures are separate: Compute Engine retires the failed server, replaces it, and restarts affected VMs if they are set to restart automatically.
CPU overcommit
Many VMs are idle most of the time. CPU overcommit lets VMs on a node share spare CPU cycles, so you can place more vCPUs on a node than it physically has. The template must enable it (--cpu-overcommit-type=enabled); each VM then declares a guaranteed minimum with --min-node-cpu, and its machine type's vCPU count becomes the ceiling it can burst to. The minimum can be as low as half the VM's vCPUs, so the maximum overcommit ratio is 2.0. Nodes with overcommit enabled are charged an extra 25 percent, and CPU quota is still counted on the node type's vCPUs.
Overcommit pays when bursts are uncorrelated. If every VM runs a batch job at 02:00, they all hit their minimums at once and run degraded. It can also lower per-VM licence costs, because the licence follows the physical cores while more VMs share them.
How billing works
You pay for the node, not the VMs. Billing covers all the vCPU and memory of every node in the group from the moment it exists, plus a sole-tenancy premium of 10 percent of that vCPU and memory cost; VMs placed on the node cost nothing extra. GPUs and Local SSDs in the template are billed for every one on the node, and the premium excludes them. Sustained use discounts apply to the vCPU, memory and premium. Resource-based committed use discounts cover the vCPU and memory but not the premium; compute flexible CUDs do discount the premium.
The pricing page works an example for an n1-node-96-624 in us-west1, as read on 2026-10-03: 96 vCPU at $0.031611 per hour is $3.034656, plus 624 GB at $0.004237 is $2.643888, so the base is $5.678544 per hour; the 10 percent premium adds $0.5678544, giving $6.2463984 per hour or about $4,560 for a 730-hour month before discounts. The lesson for planning is that an empty node costs the same as a full one, so utilisation is the whole game.
Worked example: SQL Server with per-core licences
A team must move a SQL Server estate whose licences are counted per physical core: 10 production instances and 12 development instances. Production must run on servers dedicated to the company, keep licensed core counts bounded, and survive a zone failure through an Always On availability group spanning two zones.
Sizing. Production instances are n2-highmem-16 (16 vCPUs, 128 GB). An n2-node-80-640 holds five of them by memory (640 GB) and by vCPU (80). Put five in zone a and five in zone b: one node per zone carries the load. With migrate-within-node-group, each group needs at least two nodes, and a 2-node group reserves one as holdback, so each zone's group has two nodes and the licensed core set stays fixed to those two servers. Development uses a separate group with the default policy and CPU overcommit at 1.5x (12 vCPU VMs guaranteed 8), because development load is bursty and licence-sensitive but downtime-tolerant.
# Template with custom labels for production SQL
gcloud compute sole-tenancy node-templates create sql-prod-tmpl \
--region=europe-west1 --node-type=n2-node-80-640 \
--node-affinity-labels=workload=sql,environment=prod
# One group per zone, pinned to a fixed set of servers during maintenance
gcloud compute sole-tenancy node-groups create sql-prod-b \
--zone=europe-west1-b --node-template=sql-prod-tmpl --target-size=2 \
--maintenance-policy=migrate-within-node-group --autoscaler-mode=off
# Production VM placed by label affinity
gcloud compute instances create sqlprod-b-01 --zone=europe-west1-b \
--machine-type=n2-highmem-16 --node-affinity-file=node-affinity-sql-prod.json \
--image=YOUR_IMPORTED_BYOL_SQL_IMAGE
# Development template and group with overcommit; VMs guarantee 8 of 12 vCPUs
gcloud compute sole-tenancy node-templates create sql-dev-tmpl \
--region=europe-west1 --node-type=n2-node-80-640 \
--node-affinity-labels=workload=sql,environment=dev --cpu-overcommit-type=enabled
gcloud compute sole-tenancy node-groups create sql-dev-b \
--zone=europe-west1-b --node-template=sql-dev-tmpl --target-size=1 \
--maintenance-policy=default --maintenance-window-start-time=04:00
gcloud compute instances create sqldev-01 --zone=europe-west1-b \
--machine-type=n2-standard-12 --min-node-cpu=8 --node-group=sql-dev-b \
--image=YOUR_IMPORTED_BYOL_SQL_IMAGEThe image is a placeholder: a BYOL estate imports its own SQL Server images rather than using pay-as-you-go images, which bill the licence again. Then test: simulate a host maintenance event in each production group, confirm the availability group fails over cleanly, and confirm the VMs stayed on the group's servers.
Failure modes
- No capacity on create. Node types are available only in some zones, and a large node type can be temporarily unavailable. Create groups before you need them and use reservations or a minimum autoscaler size for critical capacity.
- Fragmentation. A node with 30 spare vCPUs but only 20 GB of free memory cannot take a highmem VM. Track free vCPU and memory per node and pick VM shapes that tile the node type.
- Wrong policy for the licence. The default policy can live migrate VMs to any sole-tenant server, so the set of physical cores your VMs have touched grows over time. If the licence counts those cores, use migrate-within-node-group or restart-in-place.
- Restart-in-place surprise. Teams choose it for licence reasons and then discover about an hour of downtime every few weeks. Pair it with replication across groups.
- Template drift. Templates are immutable, so a label or node type change means new templates, new groups and VM moves. Keep templates in infrastructure as code and version their names.
- Paying for empty servers. A group sized for a peak runs all month. Use the node group autoscaler where licences allow, or the
only-scale-outmode where removing nodes would break licence counting.
Trade-offs
Sole tenancy trades flexibility and cost for control. You give up preemptible VMs, some machine series and the scheduler's freedom to place VMs anywhere; you pay a 10 percent premium on capacity you may not fill; and you take on capacity planning that the cloud normally hides. In return you get physical isolation that auditors accept, a bounded core count for licensing, and the ability to bin-pack and overcommit your own servers. For pure isolation from other tenants with no licence angle, check first whether the requirement really needs dedicated hardware or would be met by Confidential VMs, VPC Service Controls and organisation policy.
Related reading on this site: Compute Engine in depth, Compute Engine instance scheduling, managed instance group autoscaling, Spot and preemptible VMs and Google Cloud IAM.
What to do next
- Write down why you need sole tenancy: compliance isolation, per-core licensing, affinity or bin-packing; each points to different settings.
- Pick a node type whose vCPU and memory ratio tiles your VM shapes, and confirm it is offered in your zones.
- Choose the maintenance policy from the licence terms, and set a maintenance window at creation if you use default or restart-in-place.
- Define custom affinity labels on templates and place VMs by label, not by node name.
- Create node groups in at least two zones and replicate the application across them.
- Enable overcommit only for bursty, uncorrelated workloads, and measure CPU steal time inside the guests after you do.
- Model cost per node including the 10 percent premium, and decide between resource-based and flexible CUDs.
- Simulate a host maintenance event in each group before going live, and keep templates in infrastructure as code.