Multi-Instance GPU (MIG) splits one data-centre GPU into hardware-isolated instances, each with its own streaming multiprocessors, its own share of L2 cache and memory bandwidth, and its own memory capacity. The site already covers how MIG partitions the hardware and what a CUDA program sees inside one slice. This page is about the other half of the job: running MIG as a shared service for many teams. That means deciding which tenants get which profiles, keeping a fleet of GPUs in layouts that match demand, enforcing quotas, changing layouts without surprising anyone, billing fairly, and knowing exactly which failures MIG does not contain.

Specific numbers here are for the H100 80GB and A100 40GB as documented in NVIDIA's MIG user guide. Other MIG-capable parts use different profile names and counts, so treat the tables as examples and run nvidia-smi mig -lgip on your own hardware before planning.

Tenants first: classes and profiles

Start from the tenant, not the GPU. Every tenant workload has a memory floor (weights, activations, KV cache or batch), a compute need, a latency target and a tolerance for interference. MIG profiles are named for both resources: in 2g.20gb, the 2g is two of seven compute slices and the 20gb is the memory. On an H100 80GB the documented profiles are:

ProfileCompute slicesMemory slicesMax per GPUTypical tenant
1g.10gb1/71/87Small model inference, notebooks, CI tests
1g.20gb1/72/84Memory-heavy, compute-light: embedding lookups, long KV cache
2g.20gb2/72/83Mid-size inference with steady traffic
3g.40gb3/74/82Larger inference, fine-tuning small models
4g.40gb4/74/81Compute-heavy jobs that fit in 40 GB
7g.80gb7/78/81Whole GPU, still in MIG mode

Three tenant classes cover most fleets. Latency-sensitive inference gets dedicated instances sized to its memory floor with headroom, because MIG's value to it is predictable tail latency. Interactive and development users get 1g instances: enough to load a model and test, cheap enough to hand out freely. Distributed training generally should not be on MIG at all: NVIDIA's deployment notes state that NCCL is not supported with MIG, and P2P between instances on different GPUs is not supported, so multi-GPU jobs belong on whole GPUs in a non-MIG pool. Put that rule in your admission policy rather than in a wiki page.

Kubernetes: strategies, layouts and quotas

On Kubernetes, the NVIDIA GPU Operator manages MIG through two components: MIG Manager, which applies a layout to a node, and the device plugin, which advertises the resulting instances to the scheduler. The device plugin's MIG strategy decides how instances are named:

  • single: every GPU on the node has the same layout, and instances are advertised as plain nvidia.com/gpu. Pods need no changes, which is why it suits homogeneous pools such as an all-1g.10gb development pool.
  • mixed: layouts can differ, and each profile is its own extended resource of the form nvidia.com/mig-<compute>g.<memory>gb, for example nvidia.com/mig-3g.40gb. Pods ask for the profile by name.

You select a layout by labelling the node with nvidia.com/mig.config, and MIG Manager reports progress in nvidia.com/mig.config.state with values such as pending, rebooting, success and failed. Layouts live in a mig-parted config file; the operator ships named ones such as all-1g.10gb and all-balanced (on an H100 80GB, all-balanced yields two 1g.10gb, one 2g.20gb and one 3g.40gb), and you add your own:

version: v1
mig-configs:
  tenant-inference-a:          # four small tenants plus one 3g for a larger model
    - devices: all
      mig-enabled: true
      mig-devices:
        "1g.10gb": 4
        "3g.40gb": 1
  tenant-heavy:
    - devices: all
      mig-enabled: true
      mig-devices:
        "4g.40gb": 1
        "3g.40gb": 1

Quotas then work the way they do for any extended resource. Give each tenant namespace a ResourceQuota per profile, so one team cannot quietly take every 3g instance in the cluster:

apiVersion: v1
kind: ResourceQuota
metadata:
  name: mig-quota
  namespace: team-search
spec:
  hard:
    requests.nvidia.com/mig-1g.10gb: "6"
    requests.nvidia.com/mig-3g.40gb: "1"
---
# inside the tenant's pod spec
resources:
  limits:
    nvidia.com/mig-1g.10gb: 1
Multi-tenant MIG on Kubernetes: who decides whatTenant demandprofile requestsLayout plannerdemand to mig.configNode labelnvidia.com/mig.configMIG Managerdrain, apply, restartDevice pluginnvidia.com/mig-*Scheduler + quotaper-namespace limitsTenant podsone slice eachDCGM exporterper-instance seriesMeteringslice-hours by tenantdemand signal
Control and data flow for a MIG tenancy service. Demand drives layout labels; MIG Manager applies them; the device plugin, scheduler and quotas hand slices to tenants; per-instance telemetry feeds metering and the next plan.

Planning the layout mix from demand

Layouts are a packing problem with two dimensions. Each profile consumes some of the seven compute slices and some of the eight memory slices, and the hardware also has placement rules about which positions an instance may start at, which you can list with nvidia-smi mig -lgipp. The practical approach is not to solve the general problem but to keep a short menu of layouts you have verified on real hardware, and choose how many GPUs get each layout so demand is covered with the fewest GPUs.

H100 80GB: compute slices (7) and memory slices (8) must both fitall-balanced1g.10gb1g.10gb2g.20gb3g.40gb7/74g + 3g4g.40gb3g.40gb7/73g + 3g3g.40gb3g.40gbidle6/7two 3g.40gb take all 8 memory slices but only 6 of 7 compute slices: one slice of SMs is strandedbars show compute slices; memory: 1g.10gb = 1/8, 2g.20gb = 2/8, 3g.40gb and 4g.40gb = 4/8
Three H100 80GB layouts. The third is legal and common, and it wastes a seventh of the GPU compute because memory, not compute, runs out first.
# Plan how many GPUs get each verified layout so that demand is met.
# Greedy is not optimal in general, but with a short menu it is close, fast and explainable.
LAYOUTS = {   # verified with mig-parted on H100 80GB before being added here
    "all-balanced":       {"1g.10gb": 2, "2g.20gb": 1, "3g.40gb": 1},
    "tenant-inference-a": {"1g.10gb": 4, "3g.40gb": 1},
    "all-1g.10gb":        {"1g.10gb": 7},
    "tenant-heavy":       {"4g.40gb": 1, "3g.40gb": 1},
    "whole":              {"7g.80gb": 1},
}
COMPUTE = {"1g.10gb": 1, "1g.20gb": 1, "2g.20gb": 2, "3g.40gb": 3, "4g.40gb": 4, "7g.80gb": 7}

def plan(demand, gpus_available):
    need = dict(demand)
    chosen = []
    while any(v > 0 for v in need.values()):
        def useful(layout):              # compute slices that serve unmet demand
            return sum(min(n, need.get(p, 0)) * COMPUTE[p] for p, n in LAYOUTS[layout].items())
        def served(layout):              # tie-break: distinct unmet profiles it serves
            return sum(1 for p in LAYOUTS[layout] if need.get(p, 0) > 0)
        best = max(sorted(LAYOUTS), key=lambda L: (useful(L), served(L)))
        if useful(best) == 0:
            raise ValueError(f"no layout serves remaining demand {need}")
        chosen.append(best)
        for prof, n in LAYOUTS[best].items():
            need[prof] = max(0, need.get(prof, 0) - n)
        if len(chosen) > gpus_available:
            raise ValueError(f"demand needs more than {gpus_available} GPUs")
    return chosen

print(plan({"1g.10gb": 4, "2g.20gb": 2, "3g.40gb": 2, "7g.80gb": 1}, gpus_available=4))
# ['all-balanced', 'all-balanced', 'whole']

Run the planner on a schedule against a demand signal, not on every request. The demand signal is the count of pending pods per profile plus a forecast from the last few days of usage. Hysteresis matters: a layout change costs a drain, so only change a GPU when the plan has asked for the same new layout for several consecutive runs.

Worked example: four tenants, four GPUs

A platform team has four tenants on a node with four H100 80GB GPUs. Search wants four 1g.10gb instances for small rerankers. Support wants two 2g.20gb instances for a summariser. Ranking wants two 3g.40gb for a larger model. Research wants one whole GPU for a week of experiments. Demand in compute slices is 4 + 4 + 6 + 7 = 21, out of 28 available, and memory slices are 4 + 4 + 8 + 8 = 24 of 32.

The planner picks all-balanced twice and whole once, leaving GPU 3 free. Two all-balanced GPUs provide exactly four 1g.10gb, two 2g.20gb and two 3g.40gb. GPU 3 can stay out of MIG mode for the next training job, or take all-1g.10gb as a development pool. If the planner had instead given search its own tenant-inference-a GPU, the remaining two 2g.20gb and one 3g.40gb would have needed two more GPUs (all-balanced covers only one 2g), plus research's whole GPU: all four GPUs, with four 1g.10gb and one 3g.40gb instance idle. Matching layouts to the mix, not to one tenant, is the whole game.

Billing follows from the layout. Weight each instance by the larger of its compute share and its memory share, because whichever is larger is what it denies to others: a 1g.10gb weighs max(1/7, 1/8) = 0.143, a 2g.20gb 0.286, a 3g.40gb max(3/7, 4/8) = 0.5. On all-balanced those weights sum to 1.071, so charging them raw would bill 107% of the GPU. Normalise per GPU instead: divide each weight by its layout's sum. On all-balanced a 1g.10gb then costs 0.133 of a GPU, the 2g.20gb 0.267 and the 3g.40gb 0.467, which sums to exactly one. Over a 730-hour month, ranking's two 3g instances bill about 681 GPU-hours against research's 730 for a whole GPU. Publish the formula, including the normalisation, before the first invoice.

Re-layout is a drain

Changing a layout is a maintenance event. When the nvidia.com/mig.config label changes, MIG Manager stops all GPU pods on the node, including the device plugin, GPU feature discovery and the DCGM exporter, applies the new geometry and restarts them. Every tenant on every GPU of that node loses its instance, not only the tenants on the GPU being changed. On A100 and A30, enabling MIG mode also requires a GPU reset; from Hopper onward, NVIDIA documents that enabling MIG mode no longer needs a reset. Instances themselves are not persistent across reboots, which is why a declarative tool such as mig-parted, rather than a one-off script, has to own the layout.

# A safe re-layout of one node, as a runbook
kubectl cordon gpu-node-07
kubectl drain gpu-node-07 --ignore-daemonsets --delete-emptydir-data \
    --pod-selector='gpu-tenant'            # tenant pods only; operator pods stay
kubectl label node gpu-node-07 nvidia.com/mig.config=tenant-heavy --overwrite
kubectl get node gpu-node-07 -o jsonpath='{.metadata.labels.nvidia\.com/mig\.config\.state}'
# wait for: success   (alert on: failed, or pending for more than 10 minutes)
kubectl uncordon gpu-node-07

# Manual equivalent on a node outside Kubernetes
sudo nvidia-smi -i 0 -mig 1               # enable MIG mode on GPU 0
sudo nvidia-smi mig -lgip                 # list profiles this GPU offers
sudo nvidia-smi mig -cgi 4g.40gb,3g.40gb -C   # create GIs and default CIs; larger first
sudo nvidia-smi mig -dci && sudo nvidia-smi mig -dgi   # tear down

Keep spare capacity in the right shape so a drain does not cause an outage: before re-laying out a node, make sure each displaced tenant has a free instance of its profile elsewhere, and roll nodes one at a time.

What MIG does not isolate

MIG isolates SMs, L2 cache slices, memory bandwidth and memory capacity, and it contains most errors to the instance that caused them. Tenants still share several things, and your security and reliability reviews should list them:

  • The physical GPU and its node. A fault that needs a full GPU reset, a driver upgrade, or a node reboot takes every tenant down together.
  • Power and thermals. Instances share the board power limit and cooling. One tenant running dense matrix work can raise board power and temperature, and any resulting clock reduction applies to the whole GPU. Watch board power alongside per-instance metrics.
  • The host. PCIe links, CPU, host memory and the NIC are shared unless you also partition them with CPU and NUMA pinning.
  • Re-layouts. As above, a layout change on one GPU disrupts the node.
  • Observability limits. NVIDIA notes that profiling of shared GPU resources is not supported in MIG mode, so some whole-GPU profiler views are unavailable to tenants.

For observability, the DCGM exporter attaches MIG instance identifiers to each series (GPU_I_ID and GPU_I_PROFILE labels), so you can attribute usage per tenant. Use DCGM_FI_PROF_GR_ENGINE_ACTIVE for activity rather than DCGM_FI_DEV_GPU_UTIL, which is not meaningful per instance. See the DCGM guide for setting up the exporter.

Failure modes

  • Memory-stranded layouts. Two 3g.40gb on an H100 use all memory and six of seven compute slices. Fine if intended; costly if it is the default.
  • Fragmentation. Free capacity spread as single 1g instances across many GPUs cannot satisfy a 3g request. Track free capacity per profile, not in GPUs.
  • Layout thrash. A planner without hysteresis drains nodes every hour. Require several consecutive identical plans before changing a GPU.
  • Stuck state. nvidia.com/mig.config.state=failed or a long pending leaves a node with no GPU resources advertised. Alert on it.
  • Wrong strategy. Pods requesting nvidia.com/gpu on a mixed node will not get a MIG instance; pods requesting a profile on a single node will stay pending.
  • Training jobs on slices. Multi-GPU jobs fail because NCCL is not supported with MIG. Route them to a non-MIG pool at admission.

Trade-offs

OptionIsolationUtilisationChange costBest for
MIG, dedicated instancesHardware, per instanceMedium: fixed shapesDrain per nodeSLO inference, untrusted tenants
MIG plus MPS inside an instanceHardware between tenants, soft withinHigherDrain per nodeMany small trusted processes of one tenant
Time slicingNone for memoryHigh on paperConfig onlyBursty development work
Whole GPUsFullLow for small jobsNoneTraining, large models

Most fleets run several pools: whole GPUs for training, mixed-strategy MIG nodes for inference tenants, and a single-strategy 1g pool or a time-sliced pool for development. The sharing strategies overview compares the mechanisms in more depth.

What to do next

  1. Run nvidia-smi mig -lgip and -lgipp on each GPU model you own and record the real profiles and placements.
  2. Write down tenant classes with a memory floor, latency target and whether multi-GPU communication is needed; route multi-GPU work away from MIG.
  3. Build a menu of three to five verified mig-parted layouts and test each on a real node.
  4. Choose single or mixed strategy per node pool and add ResourceQuotas per profile.
  5. Run a planner against pending pods and usage with hysteresis before changing any layout.
  6. Publish a chargeback rule: the larger of compute and memory share, normalised so each GPU's layout sums to one.
  7. Alert on nvidia.com/mig.config.state, board power, and per-instance DCGM_FI_PROF_GR_ENGINE_ACTIVE.
Key takeaway: Multi-tenant MIG is a fleet problem more than a hardware one: classify tenants by memory floor and latency target, keep a short menu of verified layouts, plan the mix from demand with hysteresis, enforce per-profile quotas, treat re-layouts as node drains, bill by the larger of compute and memory share normalised per GPU, and remember that power, the host and full-GPU faults are still shared.