Multi-Instance GPU (MIG) splits one data-centre GPU into hardware-isolated instances, each with its own streaming multiprocessors, its own share of L2 cache and memory bandwidth, and its own memory capacity. The site already covers how MIG partitions the hardware and what a CUDA program sees inside one slice. This page is about the other half of the job: running MIG as a shared service for many teams. That means deciding which tenants get which profiles, keeping a fleet of GPUs in layouts that match demand, enforcing quotas, changing layouts without surprising anyone, billing fairly, and knowing exactly which failures MIG does not contain.
Specific numbers here are for the H100 80GB and A100 40GB as documented in NVIDIA's MIG user guide. Other MIG-capable parts use different profile names and counts, so treat the tables as examples and run nvidia-smi mig -lgip on your own hardware before planning.
Tenants first: classes and profiles
Start from the tenant, not the GPU. Every tenant workload has a memory floor (weights, activations, KV cache or batch), a compute need, a latency target and a tolerance for interference. MIG profiles are named for both resources: in 2g.20gb, the 2g is two of seven compute slices and the 20gb is the memory. On an H100 80GB the documented profiles are:
| Profile | Compute slices | Memory slices | Max per GPU | Typical tenant |
|---|---|---|---|---|
| 1g.10gb | 1/7 | 1/8 | 7 | Small model inference, notebooks, CI tests |
| 1g.20gb | 1/7 | 2/8 | 4 | Memory-heavy, compute-light: embedding lookups, long KV cache |
| 2g.20gb | 2/7 | 2/8 | 3 | Mid-size inference with steady traffic |
| 3g.40gb | 3/7 | 4/8 | 2 | Larger inference, fine-tuning small models |
| 4g.40gb | 4/7 | 4/8 | 1 | Compute-heavy jobs that fit in 40 GB |
| 7g.80gb | 7/7 | 8/8 | 1 | Whole GPU, still in MIG mode |
Three tenant classes cover most fleets. Latency-sensitive inference gets dedicated instances sized to its memory floor with headroom, because MIG's value to it is predictable tail latency. Interactive and development users get 1g instances: enough to load a model and test, cheap enough to hand out freely. Distributed training generally should not be on MIG at all: NVIDIA's deployment notes state that NCCL is not supported with MIG, and P2P between instances on different GPUs is not supported, so multi-GPU jobs belong on whole GPUs in a non-MIG pool. Put that rule in your admission policy rather than in a wiki page.
Kubernetes: strategies, layouts and quotas
On Kubernetes, the NVIDIA GPU Operator manages MIG through two components: MIG Manager, which applies a layout to a node, and the device plugin, which advertises the resulting instances to the scheduler. The device plugin's MIG strategy decides how instances are named:
- single: every GPU on the node has the same layout, and instances are advertised as plain
nvidia.com/gpu. Pods need no changes, which is why it suits homogeneous pools such as an all-1g.10gbdevelopment pool. - mixed: layouts can differ, and each profile is its own extended resource of the form
nvidia.com/mig-<compute>g.<memory>gb, for examplenvidia.com/mig-3g.40gb. Pods ask for the profile by name.
You select a layout by labelling the node with nvidia.com/mig.config, and MIG Manager reports progress in nvidia.com/mig.config.state with values such as pending, rebooting, success and failed. Layouts live in a mig-parted config file; the operator ships named ones such as all-1g.10gb and all-balanced (on an H100 80GB, all-balanced yields two 1g.10gb, one 2g.20gb and one 3g.40gb), and you add your own:
version: v1
mig-configs:
tenant-inference-a: # four small tenants plus one 3g for a larger model
- devices: all
mig-enabled: true
mig-devices:
"1g.10gb": 4
"3g.40gb": 1
tenant-heavy:
- devices: all
mig-enabled: true
mig-devices:
"4g.40gb": 1
"3g.40gb": 1Quotas then work the way they do for any extended resource. Give each tenant namespace a ResourceQuota per profile, so one team cannot quietly take every 3g instance in the cluster:
apiVersion: v1
kind: ResourceQuota
metadata:
name: mig-quota
namespace: team-search
spec:
hard:
requests.nvidia.com/mig-1g.10gb: "6"
requests.nvidia.com/mig-3g.40gb: "1"
---
# inside the tenant's pod spec
resources:
limits:
nvidia.com/mig-1g.10gb: 1Planning the layout mix from demand
Layouts are a packing problem with two dimensions. Each profile consumes some of the seven compute slices and some of the eight memory slices, and the hardware also has placement rules about which positions an instance may start at, which you can list with nvidia-smi mig -lgipp. The practical approach is not to solve the general problem but to keep a short menu of layouts you have verified on real hardware, and choose how many GPUs get each layout so demand is covered with the fewest GPUs.
# Plan how many GPUs get each verified layout so that demand is met.
# Greedy is not optimal in general, but with a short menu it is close, fast and explainable.
LAYOUTS = { # verified with mig-parted on H100 80GB before being added here
"all-balanced": {"1g.10gb": 2, "2g.20gb": 1, "3g.40gb": 1},
"tenant-inference-a": {"1g.10gb": 4, "3g.40gb": 1},
"all-1g.10gb": {"1g.10gb": 7},
"tenant-heavy": {"4g.40gb": 1, "3g.40gb": 1},
"whole": {"7g.80gb": 1},
}
COMPUTE = {"1g.10gb": 1, "1g.20gb": 1, "2g.20gb": 2, "3g.40gb": 3, "4g.40gb": 4, "7g.80gb": 7}
def plan(demand, gpus_available):
need = dict(demand)
chosen = []
while any(v > 0 for v in need.values()):
def useful(layout): # compute slices that serve unmet demand
return sum(min(n, need.get(p, 0)) * COMPUTE[p] for p, n in LAYOUTS[layout].items())
def served(layout): # tie-break: distinct unmet profiles it serves
return sum(1 for p in LAYOUTS[layout] if need.get(p, 0) > 0)
best = max(sorted(LAYOUTS), key=lambda L: (useful(L), served(L)))
if useful(best) == 0:
raise ValueError(f"no layout serves remaining demand {need}")
chosen.append(best)
for prof, n in LAYOUTS[best].items():
need[prof] = max(0, need.get(prof, 0) - n)
if len(chosen) > gpus_available:
raise ValueError(f"demand needs more than {gpus_available} GPUs")
return chosen
print(plan({"1g.10gb": 4, "2g.20gb": 2, "3g.40gb": 2, "7g.80gb": 1}, gpus_available=4))
# ['all-balanced', 'all-balanced', 'whole']Run the planner on a schedule against a demand signal, not on every request. The demand signal is the count of pending pods per profile plus a forecast from the last few days of usage. Hysteresis matters: a layout change costs a drain, so only change a GPU when the plan has asked for the same new layout for several consecutive runs.
Worked example: four tenants, four GPUs
A platform team has four tenants on a node with four H100 80GB GPUs. Search wants four 1g.10gb instances for small rerankers. Support wants two 2g.20gb instances for a summariser. Ranking wants two 3g.40gb for a larger model. Research wants one whole GPU for a week of experiments. Demand in compute slices is 4 + 4 + 6 + 7 = 21, out of 28 available, and memory slices are 4 + 4 + 8 + 8 = 24 of 32.
The planner picks all-balanced twice and whole once, leaving GPU 3 free. Two all-balanced GPUs provide exactly four 1g.10gb, two 2g.20gb and two 3g.40gb. GPU 3 can stay out of MIG mode for the next training job, or take all-1g.10gb as a development pool. If the planner had instead given search its own tenant-inference-a GPU, the remaining two 2g.20gb and one 3g.40gb would have needed two more GPUs (all-balanced covers only one 2g), plus research's whole GPU: all four GPUs, with four 1g.10gb and one 3g.40gb instance idle. Matching layouts to the mix, not to one tenant, is the whole game.
Billing follows from the layout. Weight each instance by the larger of its compute share and its memory share, because whichever is larger is what it denies to others: a 1g.10gb weighs max(1/7, 1/8) = 0.143, a 2g.20gb 0.286, a 3g.40gb max(3/7, 4/8) = 0.5. On all-balanced those weights sum to 1.071, so charging them raw would bill 107% of the GPU. Normalise per GPU instead: divide each weight by its layout's sum. On all-balanced a 1g.10gb then costs 0.133 of a GPU, the 2g.20gb 0.267 and the 3g.40gb 0.467, which sums to exactly one. Over a 730-hour month, ranking's two 3g instances bill about 681 GPU-hours against research's 730 for a whole GPU. Publish the formula, including the normalisation, before the first invoice.
Re-layout is a drain
Changing a layout is a maintenance event. When the nvidia.com/mig.config label changes, MIG Manager stops all GPU pods on the node, including the device plugin, GPU feature discovery and the DCGM exporter, applies the new geometry and restarts them. Every tenant on every GPU of that node loses its instance, not only the tenants on the GPU being changed. On A100 and A30, enabling MIG mode also requires a GPU reset; from Hopper onward, NVIDIA documents that enabling MIG mode no longer needs a reset. Instances themselves are not persistent across reboots, which is why a declarative tool such as mig-parted, rather than a one-off script, has to own the layout.
# A safe re-layout of one node, as a runbook
kubectl cordon gpu-node-07
kubectl drain gpu-node-07 --ignore-daemonsets --delete-emptydir-data \
--pod-selector='gpu-tenant' # tenant pods only; operator pods stay
kubectl label node gpu-node-07 nvidia.com/mig.config=tenant-heavy --overwrite
kubectl get node gpu-node-07 -o jsonpath='{.metadata.labels.nvidia\.com/mig\.config\.state}'
# wait for: success (alert on: failed, or pending for more than 10 minutes)
kubectl uncordon gpu-node-07
# Manual equivalent on a node outside Kubernetes
sudo nvidia-smi -i 0 -mig 1 # enable MIG mode on GPU 0
sudo nvidia-smi mig -lgip # list profiles this GPU offers
sudo nvidia-smi mig -cgi 4g.40gb,3g.40gb -C # create GIs and default CIs; larger first
sudo nvidia-smi mig -dci && sudo nvidia-smi mig -dgi # tear downKeep spare capacity in the right shape so a drain does not cause an outage: before re-laying out a node, make sure each displaced tenant has a free instance of its profile elsewhere, and roll nodes one at a time.
What MIG does not isolate
MIG isolates SMs, L2 cache slices, memory bandwidth and memory capacity, and it contains most errors to the instance that caused them. Tenants still share several things, and your security and reliability reviews should list them:
- The physical GPU and its node. A fault that needs a full GPU reset, a driver upgrade, or a node reboot takes every tenant down together.
- Power and thermals. Instances share the board power limit and cooling. One tenant running dense matrix work can raise board power and temperature, and any resulting clock reduction applies to the whole GPU. Watch board power alongside per-instance metrics.
- The host. PCIe links, CPU, host memory and the NIC are shared unless you also partition them with CPU and NUMA pinning.
- Re-layouts. As above, a layout change on one GPU disrupts the node.
- Observability limits. NVIDIA notes that profiling of shared GPU resources is not supported in MIG mode, so some whole-GPU profiler views are unavailable to tenants.
For observability, the DCGM exporter attaches MIG instance identifiers to each series (GPU_I_ID and GPU_I_PROFILE labels), so you can attribute usage per tenant. Use DCGM_FI_PROF_GR_ENGINE_ACTIVE for activity rather than DCGM_FI_DEV_GPU_UTIL, which is not meaningful per instance. See the DCGM guide for setting up the exporter.
Failure modes
- Memory-stranded layouts. Two 3g.40gb on an H100 use all memory and six of seven compute slices. Fine if intended; costly if it is the default.
- Fragmentation. Free capacity spread as single 1g instances across many GPUs cannot satisfy a 3g request. Track free capacity per profile, not in GPUs.
- Layout thrash. A planner without hysteresis drains nodes every hour. Require several consecutive identical plans before changing a GPU.
- Stuck state.
nvidia.com/mig.config.state=failedor a longpendingleaves a node with no GPU resources advertised. Alert on it. - Wrong strategy. Pods requesting
nvidia.com/gpuon a mixed node will not get a MIG instance; pods requesting a profile on a single node will stay pending. - Training jobs on slices. Multi-GPU jobs fail because NCCL is not supported with MIG. Route them to a non-MIG pool at admission.
Trade-offs
| Option | Isolation | Utilisation | Change cost | Best for |
|---|---|---|---|---|
| MIG, dedicated instances | Hardware, per instance | Medium: fixed shapes | Drain per node | SLO inference, untrusted tenants |
| MIG plus MPS inside an instance | Hardware between tenants, soft within | Higher | Drain per node | Many small trusted processes of one tenant |
| Time slicing | None for memory | High on paper | Config only | Bursty development work |
| Whole GPUs | Full | Low for small jobs | None | Training, large models |
Most fleets run several pools: whole GPUs for training, mixed-strategy MIG nodes for inference tenants, and a single-strategy 1g pool or a time-sliced pool for development. The sharing strategies overview compares the mechanisms in more depth.
What to do next
- Run
nvidia-smi mig -lgipand-lgippon each GPU model you own and record the real profiles and placements. - Write down tenant classes with a memory floor, latency target and whether multi-GPU communication is needed; route multi-GPU work away from MIG.
- Build a menu of three to five verified mig-parted layouts and test each on a real node.
- Choose single or mixed strategy per node pool and add ResourceQuotas per profile.
- Run a planner against pending pods and usage with hysteresis before changing any layout.
- Publish a chargeback rule: the larger of compute and memory share, normalised so each GPU's layout sums to one.
- Alert on
nvidia.com/mig.config.state, board power, and per-instanceDCGM_FI_PROF_GR_ENGINE_ACTIVE.