Multi-Instance GPU, MIG, lets a data-centre NVIDIA GPU from the A100 onward be split in hardware into as many as seven instances, each with its own slice of streaming multiprocessors, its own L2 cache and memory-controller path, and its own fixed share of memory. How that partitioning works inside the chip is covered in the MIG architecture article. This article is about the other half of the job: deploying MIG in a fleet. That means choosing geometries from your workload mix, applying them reproducibly, exposing them to a scheduler, keeping them observable, and changing them without taking a cluster down.

The reason to bother is money. A small embedding model, a reranker or a 7B model at modest traffic often needs ten or twenty gigabytes and a fraction of the compute of an 80 GB GPU. Giving each of those services a whole GPU strands most of the device. MIG lets several of them share one GPU with hard memory limits and stable latency, which time slicing and MPS do not give you. The cost is rigidity: geometries are coarse, changing them evicts work, and some tooling behaves differently on a sliced GPU. Deployment is mostly the art of managing that rigidity.

Instances, profiles and the real profile table

Three terms carry the whole operational model. A GPU instance (GI) owns a set of memory slices and compute slices; memory isolation happens at this level. A compute instance (CI) lives inside a GI and owns some of its compute slices. Most deployments create exactly one CI per GI using all of its compute, so the pair behaves as one device. The MIG device is what CUDA and Kubernetes see: a GI plus CI pair with its own UUID.

Profiles are named by compute slices and memory: 1g.10gb is one-seventh of the compute with 10 GB. Profile tables differ by GPU model and memory size, and drivers occasionally add profiles, so the only authoritative table is the one your own driver prints. As an orientation, an 80 GB A100 or H100 offers profiles along these lines:

ProfileCompute slicesMemoryMax per GPU
1g.10gb1 of 710 GB7
1g.20gb1 of 720 GB4
2g.20gb2 of 720 GB3
3g.40gb3 of 740 GB2
4g.40gb4 of 740 GB1
7g.80gb7 of 780 GB1

The 40 GB A100 has the same shapes with halved memory, 1g.5gb up to 7g.40gb, and larger memory parts such as the H100 NVL and H200 have their own tables. Numeric profile IDs are also GPU-specific, so scripts should use profile names, never IDs copied from another machine. Note that the usable memory reported inside an instance is a little less than the name suggests.

Doing it once by hand

Before automating anything, do it once by hand on a single node. It teaches the failure messages you will later see from operators.

# 1. Enable MIG mode on GPU 0. It persists across reboots; the GPU must be idle,
#    and some platforms need a GPU reset or reboot before it takes effect.
sudo nvidia-smi -i 0 -mig 1

# 2. Ask the driver what this GPU supports, and where each profile can be placed.
nvidia-smi mig -i 0 -lgip
nvidia-smi mig -i 0 -lgipp

# 3. Create GPU instances by profile name, with -C creating a default
#    compute instance inside each one.
sudo nvidia-smi mig -i 0 -cgi 3g.40gb,3g.40gb -C

# 4. List the MIG devices and their UUIDs.
nvidia-smi -L

# 5. Pin a process to one instance.
CUDA_VISIBLE_DEVICES=MIG-<uuid> python serve.py

# 6. Tear down: compute instances first, then GPU instances.
sudo nvidia-smi mig -i 0 -dci
sudo nvidia-smi mig -i 0 -dgi

Two facts from this exercise drive everything later. First, MIG mode persists across reboots but the instances do not; something must recreate the geometry every boot. Second, creating or destroying instances fails while processes hold the GPU, so every change is a drain. Also note that on current CUDA releases a process uses a single MIG device; asking for two instances gives you two separate devices with no peer-to-peer path, so MIG is not a way to build a small multi-GPU job.

Declarative geometry with mig-parted

Hand-typed nvidia-smi commands do not survive a fleet. NVIDIA's mig-parted tool, which the GPU Operator's MIG manager wraps, turns the geometry into a declarative file. Each named configuration lists devices and the instances to create on them, and applying a name is idempotent: if the GPU already matches, nothing happens.

version: v1
mig-configs:
  all-disabled:
    - devices: all
      mig-enabled: false

  serving-small:
    - devices: all
      mig-enabled: true
      mig-devices:
        "1g.10gb": 7

  serving-mixed:
    - devices: [0, 1, 2, 3]
      mig-enabled: true
      mig-devices:
        "3g.40gb": 2
    - devices: [4, 5, 6, 7]
      mig-enabled: true
      mig-devices:
        "1g.10gb": 7

Keep this file in version control next to your cluster configuration and review changes to it like code, because a geometry change is a capacity change. Keep the names few and meaningful. A fleet with twenty bespoke geometries cannot be reasoned about, and the scheduler will strand capacity across them.

Kubernetes: the operator, labels and strategies

In Kubernetes the GPU Operator ties the pieces together. A platform engineer sets the label nvidia.com/mig.config on a node to the name of a configuration, such as all-1g.10gb from the defaults or serving-mixed from your own ConfigMap. The MIG manager DaemonSet on that node notices, sets nvidia.com/mig.config.state to pending, stops the GPU pods and operator components using the GPUs, applies the geometry, restarts them, and sets the state to success or failed. The device plugin then advertises the new resources and the scheduler can place pods.

MIG in Kubernetes: who changes the geometry, who advertises it, who schedules onto itplatform engineerkubectl label nodemig.config=MIG manager (DaemonSet)applies mig-parted configdriver + GPUGIs and CIs createddevice pluginadvertises MIG resourcesenumerates via NVMLkube-schedulermatches pod limitspod specnvidia.com/mig-1g.10gb: 1pod on nodesees one MIG devicedcgm-exporterper-instance metricsLabel state moves pending, then success or failed; GPU pods on the node are evicted during the change.
The control loop: a label change drives the MIG manager, the device plugin re-advertises, and pods schedule against the new resource names.

How resources are named depends on the operator's MIG strategy, recorded on the node as nvidia.com/mig.strategy. With the single strategy every GPU on a node must have the same geometry, and instances are advertised as ordinary nvidia.com/gpu resources, with the profile visible in node labels so you can select it with node affinity. Existing pod specs work unchanged. With the mixed strategy each profile becomes its own resource name, so a pod asks for exactly what it needs:

apiVersion: v1
kind: Pod
metadata:
  name: reranker
spec:
  containers:
    - name: server
      image: registry.example.com/reranker:1.4
      resources:
        limits:
          nvidia.com/mig-1g.10gb: 1

Pick single when nodes are uniform and you want old manifests to keep working; pick mixed when one node carries several shapes. Do not mix the two conventions across node pools without documenting it, because a pod written for one stays Pending forever on the other with no error beyond insufficient resources.

Capacity planning and a bin-packing example

Geometry should come from measurement, not from round numbers. For each service record peak memory including the KV cache or activation workspace at the batch size you will actually run, the throughput one instance size achieves at your latency target, and the replica count that traffic requires. Then bin-pack services onto profiles.

A worked example. Ten small services each peak below 9 GB and meet their p99 on one compute slice. Four medium services need up to 35 GB and three slices each. On H100 80 GB GPUs, the ten small services fit on two GPUs configured 7 x 1g.10gb, with four slices spare. The four medium services fit on two GPUs configured 2 x 3g.40gb. Four GPUs carry fourteen services that would otherwise hold fourteen whole GPUs.

Worked example: 10 small and 4 medium services on H100 80GB GPUsGPU 0 all-1g.10gb1g1g1g1g1g1g1gGPU 1 all-1g.10gb1g1g1gfreefreefreefreeGPU 2 all-3g.40gb3g.40gb | 3g.40gbGPU 3 all-3g.40gb3g.40gb | 3g.40gbEach box is one of seven compute slices. Two 3g instances use six; the seventh idles by design.Four GPUs carry fourteen services that would otherwise hold fourteen whole GPUs.The four free 1g slots on GPU 1 are your headroom for the next small tenant.
Bin-packing the worked example. The idle seventh slice on the 3g GPUs is the price of two equal halves.

Two lessons hide in the picture. The 2 x 3g.40gb geometry leaves one compute slice idle by construction; a 4g.40gb plus 3g.40gb split uses all seven if one service can use the extra compute, and the driver's placement listing tells you whether a combination is legal. And spare slots are only useful if they match future demand: four free 1g slices cannot host a fifth medium service. Keep a small pool of whole GPUs, or a node you can reconfigure, for shapes the plan did not foresee.

Monitoring a sliced GPU

Observability changes on a sliced GPU, and teams usually discover it from a broken dashboard. The classic utilisation counter, DCGM_FI_DEV_GPU_UTIL, is not reported per MIG instance. Use the profiling metrics instead: DCGM_FI_PROF_GR_ENGINE_ACTIVE for how busy an instance's compute is, the tensor-pipe and DRAM-activity profiling fields for what it is doing, and framebuffer used and free per instance. The DCGM exporter labels series with the GPU instance ID and profile, so dashboards and alerts can key on the tenant's slice rather than the physical GPU. The DCGM article covers the exporter itself.

Point autoscalers at request latency and queue depth, not GPU utilisation. Alert on per-instance memory headroom, since an instance cannot borrow memory from its neighbours and an out-of-memory error is the first symptom. And keep watching device-level health: Xid errors, ECC events and temperature are still properties of the physical GPU, and a fault that needs a full GPU reset takes every instance on it down together.

Changing geometry safely

Reconfiguration is the operation that hurts, because every GPU pod on the node is evicted. Treat it like a kernel upgrade:

  1. Cordon the node and drain GPU workloads gracefully, honouring PodDisruptionBudgets, so replicas move before the MIG manager kills them.
  2. Change the nvidia.com/mig.config label.
  3. Wait for nvidia.com/mig.config.state to read success, and alert if it reads failed or stays pending; a failed state usually means a process still held the GPU or the geometry is not legal for that model.
  4. Check the node's allocatable resources match the plan before uncordoning.
  5. Roll through nodes one at a time, never a whole pool.

Plan geometry changes weekly or monthly, not per deployment. If demand shapes swing within a day, MIG is the wrong tool for that pool; time slicing or MPS, discussed in GPU sharing strategies, adapt in seconds at the cost of isolation. Virtualised hosts add another layer: MIG-backed vGPU profiles, covered in the vGPU article, put the geometry decision on the hypervisor rather than the Kubernetes node.

Failure modes

  • Geometry lost on reboot. MIG mode survived, the instances did not, and nothing recreated them; pods sit Pending. Make the MIG manager or a boot unit own the geometry.
  • Wrong resource name. A pod asks for a mixed-strategy name on a single-strategy node, or for a profile name from a different memory size, and never schedules.
  • Stranded slices. Free 1g slots everywhere and no 3g slot anywhere. Track free capacity per profile, not per GPU.
  • Out-of-memory inside a slice. The service fit at batch 8 and a config change raised it to 16. Instances cannot borrow memory, so cap batch and context at deploy time.
  • Dashboards showing zero. GPU_UTIL is not reported per MIG instance; switch to profiling metrics before someone downsizes a busy pool.
  • Assuming multi-GPU works. NCCL jobs across MIG devices do not get NVLink or peer-to-peer; training jobs belong on whole GPUs.

Trade-offs

ChoiceGainCost
MIG with single strategyOld manifests work, simple poolsOne geometry per node
MIG with mixed strategyExact sizing per podMore resource names, more stranding
Few large instancesHeadroom for bursts and long contextsFewer tenants per GPU
Many 1g instancesHighest density for small modelsLow per-tenant throughput
Whole GPUs or MPS insteadElastic, no drainsWeaker isolation, noisy neighbours

What to do next

  1. Run nvidia-smi mig -lgip on each GPU model you own and save the real profile tables.
  2. Measure peak memory and the latency-bound throughput per service on one candidate profile.
  3. Bin-pack the services, choose two or three named geometries, and commit them as a mig-parted config.
  4. Decide single versus mixed strategy per node pool and write it in the platform docs.
  5. Rebuild GPU dashboards on per-instance profiling and memory metrics.
  6. Write a reconfiguration runbook with cordon, drain, label, wait for success and verify, and rehearse it on one node.
Key takeaway: Deploying MIG well is mostly operations: read the profile table from your own driver, derive a few geometries from measured service memory and latency, apply them declaratively through mig-parted and the GPU Operator label, pick single or mixed naming per pool, monitor per instance, and treat every geometry change as a drain.