On one machine, nvidia-smi answers most questions: is the GPU there, how hot is it, how much memory is used. At a thousand GPUs the questions change. Which nodes should not receive the next job? Is this training run slow because of the model, the input pipeline or one bad GPU? Which job was running when that GPU threw an uncorrectable memory error? nvidia-smi was not built for those questions. NVIDIA Data Center GPU Manager (DCGM) was.

DCGM is a daemon and library that samples GPU telemetry continuously, runs health checks and diagnostics, and records statistics per job. This article covers how it is built, which metrics describe a training step, how to keep bad nodes out of jobs, dcgm-exporter, and what goes wrong in each layer.

Advertisement

Where DCGM sits: the host engine and its clients

Underneath everything is NVML, the driver library nvidia-smi also uses. DCGM adds a layer above it: a core that knows which values to sample, how often and for how long to keep them, plus health, policy, diagnostic and statistics modules. The core runs in one of two modes. In standalone mode it lives inside nv-hostengine, a daemon normally run as the nvidia-dcgm systemd service, and many clients connect to it. In embedded mode an agent loads DCGM as a shared library and owns it privately.

Standalone mode is the right default for a fleet: the engine samples each GPU once and every consumer reads the same cache, instead of three tools polling NVML separately and getting three slightly different answers.

GPUs + NVSwitchNVML / driverDCGM corewatches, cache, healthnv-hostenginestandalone daemondcgmioperator CLIdcgm-exporter :9400Prometheus textScheduler hooksprolog / epilog / statssamplePrometheusscrape, alert rulesNode controllercordon, drain, RMAscrapealertdiag -r 2One host engine per node owns sampling; every consumer reads its cache instead of polling NVML itself.
DCGM on one node: the host engine samples the GPUs, and operator tools, the exporter and scheduler hooks all read its cache.

Entities, groups, fields and watches

DCGM's data model has four ideas. An entity is something that can be measured: a GPU, an NVSwitch, a MIG instance, or (on supported systems) a NIC. A group is a named set of entities you act on together, such as the eight GPUs of one training node. A field is one measurable value with a numeric ID and a name like DCGM_FI_DEV_GPU_TEMP. A watch tells the engine to sample a set of fields for a group at a given interval and keep the samples for a given age.

Watches trade cost against resolution. Cheap device fields can be sampled every second; profiling fields are better at one to ten seconds, because very fast sampling is not free. Consumers should read the cache rather than ask for fresh samples.

# What does DCGM see on this node?
dcgmi discovery -l

# A group for the 8 training GPUs, then a health watch on every subsystem
dcgmi group -c train_gpus
dcgmi group -g 2 -a 0,1,2,3,4,5,6,7      # use the group id printed by -c
dcgmi health -g 2 -s a
dcgmi health -g 2 -c                      # read accumulated health results

# Live profiling view: engine active, SM active, tensor active, DRAM active
dcgmi dmon -e 1001,1002,1004,1005 -i 0
Advertisement

GPU-Util is not the number you want: the profiling metrics

The classic utilisation figure, DCGM_FI_DEV_GPU_UTIL, is the same duty cycle nvidia-smi prints: the share of time any kernel was running. A kernel using one SM out of 132 counts as 100 percent. For a training job that is nearly useless. DCGM's datacenter profiling (DCP) fields read hardware counters instead, on Volta and newer datacenter GPUs, and report ratios between 0 and 1:

Field (ID)What it measuresWhat it tells you in training
PROF_GR_ENGINE_ACTIVE (1001)Time the graphics/compute engine is busyRoughly the old GPU-Util; low means the GPU is waiting on the host, data or communication
PROF_SM_ACTIVE (1002)Cycles an SM has at least one warp, averaged over SMsHow much of the chip the kernels spread across; low with high engine-active means small kernels
PROF_SM_OCCUPANCY (1003)Resident warps relative to the maximumLatency hiding; low occupancy is not bad by itself for tensor-core GEMMs
PROF_PIPE_TENSOR_ACTIVE (1004)Cycles the tensor-core pipe is activeThe closest single number to useful matrix math for transformer training
PROF_DRAM_ACTIVE (1005)Cycles the HBM interface is moving dataNear 1 with low tensor activity means memory-bound work: attention, norms, optimiser

Read them together. High engine-active with low SM-active means kernels too small to fill the GPU, often a small micro-batch. High SM-active with low tensor-active means the math is not on tensor cores, perhaps because of FP32 layers or awkward shapes. Low engine-active on every GPU means the GPUs are waiting on something outside them, which is where interconnect metrics such as NVLink bandwidth and the collective timing described in NCCL collectives come in. The shipped default counters file enables engine-active, tensor-active and DRAM-active but leaves SM-active and SM-occupancy commented out, so add them if you want the full picture.

Worked example: one slow node in a 64-GPU job

A 64-GPU pretraining job on eight nodes has dropped from 1.9 to 1.4 steps a second. The dashboard shows, averaged over ten minutes:

NodesGR_ENGINE_ACTIVESM_ACTIVETENSOR_ACTIVEDRAM_ACTIVEPower
Nodes 1-6, 80.720.660.410.38~640 W
Node 7, GPU 30.970.950.550.47~700 W, clocks lower
Node 7, other GPUs0.700.640.400.37~630 W

In a synchronous data-parallel job every rank waits for the slowest at each all-reduce. The healthy GPUs are idle about 28 percent of the time: that is waiting. One GPU is busy almost constantly, at the power limit with lower clocks, so it is throttling, and all 63 others wait for it.

The fix is operational: drain node 7, check its airflow, run a longer diagnostic and restart from the last checkpoint on a spare. The straggler rule later in this article, which compares each GPU's SM clock with its training job's average, would have flagged it within minutes.

Health watches and policies

A health watch monitors subsystems such as PCIe, memory, InfoROM, thermal and power, NVLink and the driver in the background. dcgmi health -g 2 -s a turns on all of them, and -c returns pass, warning or failure per GPU with a reason. It is cheap enough to leave on permanently.

Policies register conditions such as double-bit ECC errors, PCIe errors, retired pages, thermal or power limits, NVLink errors or Xids, and notify a client when one fires, optionally with an action such as a GPU reset. A reset under a running job kills the job, so most fleets use notify-only and let a node controller decide. Policy flags vary by release; read dcgmi policy --help on your version.

Diagnostics as a node gate

dcgmi diag runs active tests, not passive sampling. The plugins include software (driver and deployment checks), pcie, memory, memory_bandwidth, diagnostic (large matrix work), targeted_stress, targeted_power, memtest, pulse_test, nvbandwidth and nccl_tests. Run levels bundle them: -r 1 is a quick check that takes seconds, -r 2 adds short stress, and -r 3 and -r 4 add long stress, power, bandwidth and NCCL tests. NVIDIA documents -r 3 at under 35 minutes and -r 4 at under 2.25 hours on an 8-GPU system. Which plugin sits in which level has changed between releases, so check the table for your version before building a policy on it.

Map levels onto a lifecycle: level 1 in the scheduler prolog before every job, level 2 after a failed job or in a nightly sweep of idle nodes, level 3 or 4 for burn-in after a repair. -j prints JSON and the exit code reports errors, so this gate fails closed on either without depending on one version's JSON schema:

#!/usr/bin/env python3
# Node gate for a scheduler prolog: run a short DCGM diagnostic, fail closed.
import json, subprocess, sys

def failed_results(node):
    # Walk the JSON without assuming its schema: any string value "fail" counts.
    if isinstance(node, dict):
        return sum(failed_results(v) for v in node.values())
    if isinstance(node, list):
        return sum(failed_results(v) for v in node)
    return 1 if isinstance(node, str) and node.strip().lower() == "fail" else 0

proc = subprocess.run(["dcgmi", "diag", "-r", "1", "-j"],
                      capture_output=True, text=True, timeout=300)
try:
    fails = failed_results(json.loads(proc.stdout))
except json.JSONDecodeError:
    fails = -1                      # unparseable output is a failure, not a pass
if proc.returncode != 0 or fails != 0:
    print(f"dcgm gate: rc={proc.returncode} fails={fails}", file=sys.stderr)
    sys.exit(1)                     # scheduler drains the node, job goes elsewhere

Per-job statistics

When someone asks why last Tuesday's run was slow, the node has run twenty jobs since. Job statistics fix that: the prolog starts a recording tagged with the job id, and the epilog stops it and prints energy, utilisation, clocks, ECC, PCIe replays, Xids and throttling for exactly that job's GPUs and window. Keep the report with the job logs.

# Prolog: attribute everything that happens on these GPUs to the job
dcgmi stats -g 2 --enable
dcgmi stats -g 2 -s "$SLURM_JOB_ID"

# Epilog: stop recording and keep the per-job report with the job's logs
dcgmi stats -x "$SLURM_JOB_ID"
dcgmi stats -j "$SLURM_JOB_ID" -v > "/var/log/gpu-jobs/$SLURM_JOB_ID.txt"

dcgm-exporter: from engine to Prometheus

dcgm-exporter reads DCGM fields and serves them in Prometheus text format, by default on port 9400 at /metrics. Which fields it exports is set by a three-column CSV (field name, Prometheus type, help text) whose default lives at /etc/dcgm-exporter/default-counters.csv. The default collection interval is 30 seconds. With --kubernetes it queries the kubelet's pod-resources API and labels each GPU's series with the pod that holds it, which is what lets you group metrics by job. --remote-hostengine-info points it at a standalone host engine instead of an embedded one.

The shipped CSV is a reasonable start, but for training fleets add the SM fields and the volatile ECC counter, which it leaves commented out:

# /etc/dcgm-exporter/training-counters.csv   (field, prometheus type, help)
DCGM_FI_DEV_GPU_TEMP,                    gauge,   GPU temperature (C).
DCGM_FI_DEV_POWER_USAGE,                 gauge,   Power draw (W).
DCGM_FI_DEV_FB_USED,                     gauge,   Framebuffer used (MiB).
DCGM_FI_DEV_XID_ERRORS,                  gauge,   Last Xid seen.
DCGM_FI_DEV_ECC_DBE_VOL_TOTAL,           counter, Volatile double-bit ECC errors.
DCGM_FI_DEV_UNCORRECTABLE_REMAPPED_ROWS, counter, Rows remapped for uncorrectable errors.
DCGM_FI_DEV_ROW_REMAP_FAILURE,           gauge,   Row remapping failed.
DCGM_FI_DEV_PCIE_REPLAY_COUNTER,         counter, PCIe retries.
DCGM_FI_PROF_GR_ENGINE_ACTIVE,           gauge,   Graphics engine active ratio.
DCGM_FI_PROF_SM_ACTIVE,                  gauge,   SM active ratio.
DCGM_FI_PROF_PIPE_TENSOR_ACTIVE,         gauge,   Tensor pipe active ratio.
DCGM_FI_PROF_DRAM_ACTIVE,                gauge,   DRAM interface active ratio.
DCGM_FI_DEV_NVLINK_BANDWIDTH_TOTAL,      gauge,   NVLink throughput (MB/s).

Every field multiplies by GPU count, and high-churn labels such as per-process ids are what break Prometheus. Keep job identity to one label.

Alerting and Xid triage

Xids are the driver's error codes; some point at hardware, some at the application. A rough triage, to confirm against NVIDIA's Xid catalogue for your driver:

XidUsual meaningAction
13, 31, 43Graphics engine exception, MMU fault, GPU stopped processing; very often a bug or bad pointer in user codeLook at the job first; act on the node only if it recurs across different jobs
48, 94, 95Double-bit ECC error; contained and uncontained ECC errors on newer GPUsUncontained: drain the node and reset; repeated contained: plan replacement
63, 64Row remapping recorded or failedPending remap needs a GPU reset; a failed remap means replace the GPU
74NVLink errorCheck link counters and NVLink topology; reseat or replace if persistent
79GPU fell off the busDrain; usually power, PCIe or thermal hardware; needs a physical check

Note that DCGM_FI_DEV_XID_ERRORS is a gauge holding the last Xid seen, not a counter. Alerting on its value fires forever after a single event; alert on change instead:

# A new Xid appeared (the gauge holds the LAST Xid, so alert on change, not value)
changes(DCGM_FI_DEV_XID_ERRORS[10m]) > 0

# Uncorrectable memory trouble: page for a human, cordon automatically
increase(DCGM_FI_DEV_ECC_DBE_VOL_TOTAL[15m]) > 0 or DCGM_FI_DEV_ROW_REMAP_FAILURE == 1

# A GPU that is busy but not doing tensor math (training jobs only)
avg_over_time(DCGM_FI_PROF_GR_ENGINE_ACTIVE[15m]) > 0.9
  and avg_over_time(DCGM_FI_PROF_PIPE_TENSOR_ACTIVE[15m]) < 0.1

# Straggler: a GPU clocked well below the other GPUs of the same training job.
# train_job is YOUR label, e.g. relabelled from the pod labels --kubernetes adds;
# do not use Prometheus's own "job" label, which is the scrape job.
DCGM_FI_DEV_SM_CLOCK
  < on(train_job) group_left 0.9 * avg by (train_job) (DCGM_FI_DEV_SM_CLOCK)

MIG and the profiler conflict

On GPUs partitioned with Multi-Instance GPU, DCGM exposes GPU and compute instances as entities you can group and watch per slice. Temperature, power, ECC and NVLink stay physical-GPU fields; check which profiling fields your version supports per instance before building per-tenant dashboards.

The profiling fields use the same hardware counters as Nsight Systems and Nsight Compute, and only one consumer can own them at a time. Run dcgmi profile --pause before a profiling session and --resume afterwards, and expect a gap in that node's dashboards.

Failure modes and trade-offs

  • Two engines on one node. An embedded exporter plus the nvidia-dcgm service double the sampling and can conflict over profiling counters.
  • Averages hide spikes. A 30-second interval smooths out short stalls. DCGM shows that a GPU waits, not why; use Nsight or framework traces for that.
  • Version skew. After a driver upgrade, confirm dcgmi discovery -l sees every GPU and profiling fields return values.
  • Diagnostics on a busy GPU. Stress plugins give meaningless results on a GPU running work. Run them only on drained nodes.
  • Silent exporter death. Dead exporters silence error alerts too; alert on up == 0 for the target.
  • Automatic resets. They kill running jobs. Automate cordoning; reset only empty nodes.

What to do next

  1. Run DCGM standalone on every GPU node and point dcgm-exporter at it.
  2. Add SM-active, SM-occupancy and ECC counters to your exporter CSV and chart engine, SM, tensor and DRAM activity side by side per job.
  3. Turn on dcgmi health -s a for all GPUs and put a level-1 diagnostic in the scheduler prolog, failing closed.
  4. Wrap every job in dcgmi stats start and stop and keep the report with the job logs.
  5. Replace value-based Xid alerts with change-based ones, and route Xid 48, 63, 64, 74, 79, 94 and 95 to automatic cordoning.
  6. Add a per-job straggler rule that compares each GPU's SM clock with its training job's average.
  7. Schedule level 3 diagnostics after every hardware repair before the node rejoins the pool.
Key takeaway: DCGM is the fleet layer above NVML: one host engine per node samples the GPUs and every tool reads its cache. Use its profiling fields, not GPU-Util, to see what a training step is doing; read engine, SM, tensor and DRAM activity together, and compare each GPU with its job's peers to catch stragglers. Keep health watches on, gate every job with a seconds-long diagnostic that fails closed, save deeper diagnostics for drained nodes, and record per-job statistics. Export through dcgm-exporter with a curated field list, alert on Xid changes rather than values, and automate cordoning, not resets.