Multi-Instance GPU (MIG) lets a data-centre GPU such as an A100 or H100 be split into hardware-isolated instances, each with its own memory, cache share and streaming multiprocessors. Most writing about MIG is from the operator's side: profiles, geometries, Kubernetes. This article is from the other side, the program running inside a slice. What does your CUDA process actually see? Will your model and its KV cache fit? Can you train on slices, share one with several processes, or measure whether the isolation you are paying for is real?
The hardware model, profile shapes and the MIG-versus-MPS argument are covered in NVIDIA MIG architecture, and the fleet side, from mig-parted to Kubernetes strategies, in Multi-Instance GPU Deployment, in depth. Here we assume an administrator has already created instances and you have been handed one.
The program's view of a slice
A MIG-enabled GPU is carved into GPU instances (GIs) and, inside each, one or more compute instances (CIs). A GI owns memory slices, and with them memory capacity, bandwidth and a share of L2 cache; it is the unit of capacity and fault isolation. A CI owns a subset of its GI's SMs. CUDA treats a CI and its parent GI as one device, named MIG-<uuid> rather than by index.
To an unmodified program, that device looks like a smaller GPU: fewer SMs, less memory, the same instruction set and libraries. That is MIG's strength. The differences appear at the edges, in device enumeration, inter-process sharing, collectives, profiling and graphics, and those edges decide whether your workload belongs on a slice.
What CUDA enumerates
Enumeration rules changed with driver R570 and CUDA 12. NVIDIA's guide now states that "a single CUDA process can enumerate across multiple GPU instances, but only one CI per GI", that CUDA_VISIBLE_DEVICES accepts compute instance UUIDs, at most one per GPU instance, and that CUDA supports at most 64 MIG instances across all GPUs. With older drivers a process could enumerate only a single MIG instance, and very old drivers used a MIG-<GPU-UUID>/<GI>/<CI> form. If your fleet mixes drivers, assume the older, stricter rule.
Never rely on what you think you were given. Check at start-up:
# list MIG devices and their UUIDs (run on the node, or inside the container)
nvidia-smi -L
# GPU 0: NVIDIA A100-SXM4-80GB (UUID: GPU-...)
# MIG 3g.40gb Device 0: (UUID: MIG-...)
CUDA_VISIBLE_DEVICES=MIG-<uuid> python serve.pyimport torch
def describe_device(min_free_gb: float) -> None:
assert torch.cuda.device_count() >= 1, "no CUDA device visible"
props = torch.cuda.get_device_properties(0)
free, total = torch.cuda.mem_get_info(0)
print(f"{props.name}: {props.multi_processor_count} SMs, "
f"{total/1e9:.1f} GB total, {free/1e9:.1f} GB free")
if free / 1e9 < min_free_gb:
raise SystemExit(f"slice too small: need {min_free_gb} GB free")
describe_device(min_free_gb=18.0)Log the SM count and memory at start-up. They are the two facts that explain most performance surprises later, and the usable memory is always a little below the profile's nominal figure.
Boundaries: IPC, P2P, graphics, profiling
The guide's sharing rules follow the isolation hierarchy. "CUDA IPC across GPU instances is not supported. CUDA IPC across Compute instances is supported." Peer-to-peer between MIG instances on different GPUs, or between a MIG instance and a non-MIG GPU, is not supported. No graphics APIs such as OpenGL or Vulkan are available. Debugging with cuda-gdb and checking with compute-sanitizer work, but "profiling of shared GPU resources is not supported", so counters that describe the whole chip are not something a slice tenant can read.
The practical design rule: if two processes must exchange tensors through device memory, such as a tokenizer-and-scheduler process feeding an inference engine, put them on compute instances of one GPU instance, or in one process. If they are separate tenants, put them in separate GPU instances and accept that anything they exchange goes through host memory or the network.
Worked example: sizing a model to a slice
Memory, not compute, is usually the binding constraint for serving on a slice. The budget is: usable instance memory, minus weights, minus runtime overhead (CUDA context, allocator fragmentation, activations), with the remainder available for KV cache. The KV cache per token is 2 x layers x kv_heads x head_dim x bytes. The math is derived in KV Cache Sizing for Deployments.
GB = 1e9
def kv_tokens(usable_gb, params_b, weight_bytes, layers, kv_heads, head_dim,
kv_bytes=2, overhead_gb=1.5):
weights_gb = params_b * weight_bytes # params in billions
budget = usable_gb - weights_gb - overhead_gb
per_token = 2 * layers * kv_heads * head_dim * kv_bytes
return max(0, int(budget * GB / per_token)), round(budget, 2)
# 8B model, 32 layers, 8 KV heads of dim 128 (grouped-query attention)
for slice_, usable in [("1g.10gb", 9.75), ("2g.20gb", 19.5), ("3g.40gb", 39.5)]:
for name, wb in [("fp16", 2), ("fp8", 1)]:
print(slice_, name, kv_tokens(usable, 8.03, wb, 32, 8, 128))Take an 8-billion-parameter model with grouped-query attention: 32 layers, 8 KV heads of dimension 128, FP16 cache. Each token costs 131,072 bytes, 128 KiB. The usable memory figures (9.75, 19.5 and 39.5 GB) and the 1.5 GB overhead are assumptions for the example; measure yours with the start-up check above.
| Slice | Weights | KV budget | Tokens of KV cache |
|---|---|---|---|
| 1g.10gb | FP16, 16.1 GB | negative | does not fit |
| 1g.10gb | FP8, 8.0 GB | 0.22 GB | about 1,700 |
| 2g.20gb | FP16, 16.1 GB | 1.94 GB | about 14,800 |
| 2g.20gb | FP8, 8.0 GB | 9.97 GB | about 76,000 |
| 3g.40gb | FP16, 16.1 GB | 21.9 GB | about 167,000 |
Read the table as a concurrency limit. With 4,096-token contexts, FP16 weights on a 2g.20gb slice hold only three concurrent sequences, which wastes the slice's compute; FP8 weights on the same slice hold about eighteen. The 1g.10gb slice technically loads the FP8 model but serves almost nothing. Quantizing weights is often what makes a small slice worthwhile: FP8 on a 1g.20gb slice, where that profile exists, holds the same 76,000 tokens on a single compute slice, and FP8 on a 2g.20gb gets most of the way without moving to a 3g slice. Also check compute: a 1g slice has roughly one seventh of the SMs, so time-to-first-token on long prompts rises by a similar factor.
Configure the inference engine to match. Engines that pre-allocate a fraction of device memory for the KV cache compute that fraction from the total they see, which on a slice is the instance's memory, so the defaults are usually right; but a hard-coded byte budget copied from a whole-GPU deployment will fail at start-up or, worse, leave the slice half used. Derive every memory setting from the start-up check rather than from a config file written for different hardware.
Sharing a slice: compute instances and MPS
Sometimes one slice should serve several small processes, such as a few low-traffic models. There are two ways to do it. You can split the GPU instance into several compute instances; each gets dedicated SMs, but they share the GI's memory and bandwidth, and within the R570 rules a single process still sees only one CI per GI. Or you can run CUDA MPS on top of a MIG device. The guide says MPS is supported on MIG, "the only limitation is that the maximum number of clients (48) is lowered proportionally to the Compute Instance size".
A sensible pattern is one MPS control daemon per MIG device, each started with CUDA_VISIBLE_DEVICES set to that device's UUID and its own pipe and log directories, so tenants on different slices never share a daemon. MPS gives work-conserving sharing within the slice, while the slice boundary still protects neighbours. Remember that MPS clients share fault fate: a fatal error in one client can take down the others on that daemon. Use it for cooperating workloads you own, never between tenants.
Training on slices
Training on MIG is a single-device affair. NVIDIA's guide states plainly that "NCCL is currently not supported with MIG", so data-parallel or tensor-parallel jobs across slices, even slices on the same card, are out. For distributed training use whole GPUs and the collectives described in NCCL collectives.
Within that limit slices are useful for training-adjacent work. Hyperparameter sweeps over small models run seven independent trials on one card that do not contend for SMs or GPU memory, though they still share the host's CPUs and PCIe. Continuous-integration jobs that run a few training steps to catch regressions need a GPU but not a whole one. LoRA fine-tuning of small models fits on 3g and 4g slices when the frozen base is quantized. Notebooks and evaluation jobs get a guaranteed slice instead of contending on a time-sliced GPU. In each case write the job as a single-device job and let the scheduler place it.
Measuring isolation yourself
MIG's value is that your latency does not depend on your neighbours. Verify it on your own hardware and workload before promising it to anyone. Run a fixed probe on slice A, first with the rest of the GPU idle, then with every other slice running a saturating load, and compare the latency distributions.
import torch, statistics
def probe(n=2000, size=4096):
a = torch.randn(size, size, device="cuda", dtype=torch.float16)
b = torch.randn(size, size, device="cuda", dtype=torch.float16)
times = []
for _ in range(n):
start, end = torch.cuda.Event(True), torch.cuda.Event(True)
start.record(); a @ b; end.record()
end.synchronize()
times.append(start.elapsed_time(end))
times.sort()
return {"p50": times[n // 2], "p99": times[int(n * 0.99)],
"mean": statistics.fmean(times)}On a correctly configured GPU the two runs should be close; a large p99 gap means something else is shared, such as a CPU, PCIe link or host memory bandwidth outside the GPU. Repeat with your real model and request mix, because a GEMM probe exercises compute and a decode loop exercises memory bandwidth. Watch per-instance metrics with DCGM rather than chip-wide counters.
Failure modes
- Addressing slices by number. MIG devices have no indices, so
CUDA_VISIBLE_DEVICES=1does not select a MIG instance; name them by UUID. On pre-R570 drivers only one is visible, socuda:1fails. - Expecting two slices in one process. On pre-R570 drivers only one is visible; on newer drivers only one CI per GI. Check the device count at start-up.
- Launching NCCL jobs on slices. Unsupported; fail fast in the launcher.
- Sizing by nominal memory. Usable memory is lower than the profile name; a model sized to the nominal figure runs out of KV cache under load.
- Treating CIs as isolation. Compute instances in one GI share memory capacity and bandwidth; separate tenants need separate GIs.
- Reconfiguring under load. Changing geometry needs idle instances, and on A100 and A30 enabling MIG mode needs a GPU reset; plan it as a drain.
Trade-offs
MIG trades flexibility for predictability. A slice cannot borrow idle compute from its neighbours, so an unevenly loaded card wastes capacity that time-slicing or MPS would use; the comparison is in GPU Time Slicing. Profiles are fixed shapes, so models that need 25 GB land on a 40 GB slice and strand the rest. And the lack of collectives keeps slices out of distributed training. In return you get bounded memory, bandwidth and SM shares per tenant, separate fault domains and latency you can write an SLO against. Choose MIG when tenants or jobs are independent, small relative to the card, and latency-sensitive.
What to do next
- Add the start-up check to every service: assert a visible device, log SMs and free memory, and fail below your minimum.
- Size each model with the KV budget script using measured usable memory, and pick the slice by concurrency, not by whether the weights load.
- Try FP8 or INT8 weights before asking for a larger slice.
- Run the isolation probe with neighbours idle and saturated, then repeat with your real decode workload.
- Keep training jobs that need NCCL on whole GPUs; move sweeps, CI steps and evaluation onto slices.
- Record driver versions per node, since enumeration rules differ before and after R570.