A GPU job has a carbon footprint for two reasons: the electricity it uses while running, and its share of the emissions from manufacturing the hardware it runs on. The first depends on how much energy the job draws and how clean the grid is at that place and hour. The second depends on how long the hardware lives and how much of that life your job occupies. Both can be measured well enough to make decisions, and most of the levers that reduce them are in software: utilisation, precision, batching, power limits, scheduling and not wasting runs.
This article builds the calculation from first principles, starting from the energy counter inside the GPU and working outward to the facility and the grid. It shows code that measures a job's energy, a worked example for a fine-tuning run, and the levers ranked by how much they usually help. It ends with failure modes and a checklist. Power delivery to racks is covered in GPU datacenter power and telemetry collection in NVIDIA DCGM.
What is counted
Carbon accounting uses carbon dioxide equivalent (CO2e), which folds other greenhouse gases into one unit. For a compute job there are two parts. Operational emissions are energy times the carbon intensity of the electricity, in kilograms of CO2e per kilowatt-hour. Embodied emissions come from manufacturing, transporting and disposing of the hardware, spread over its useful life.
Two accounting conventions coexist for electricity. Location-based figures use the average intensity of the grid where the energy was consumed. Market-based figures reflect contracts such as renewable energy purchases. Corporate reports under the GHG Protocol usually show both. For engineering decisions, such as where and when to run a job, location-based figures and, where available, marginal intensity are more informative, because they reflect what the grid actually does when you add load.
The Green Software Foundation's Software Carbon Intensity specification, published as ISO/IEC 21031:2024, combines these into a rate: SCI = ((E x I) + M) per R, where E is energy, I is location-based carbon intensity, M is the embodied share and R is a functional unit you choose, such as one training run, one thousand generated tokens or one API request. Expressing results per unit is what makes them comparable across models and releases.
Measuring GPU energy
NVIDIA GPUs expose a cumulative energy counter through NVML, nvmlDeviceGetTotalEnergyConsumption, which returns millijoules consumed since the driver was last loaded. Reading it at the start and end of a job and subtracting gives the job's GPU energy directly, without integrating sampled power readings and missing the spikes between samples. Not every device or driver supports the counter, so check for a not-supported error and fall back to integrating nvmlDeviceGetPowerUsage samples.
import pynvml, time, json
class GpuEnergyMeter:
def __init__(self, indices):
pynvml.nvmlInit()
self.handles = [pynvml.nvmlDeviceGetHandleByIndex(i) for i in indices]
def read_mj(self):
return [pynvml.nvmlDeviceGetTotalEnergyConsumption(h) for h in self.handles]
def __enter__(self):
self.t0, self.e0 = time.time(), self.read_mj()
return self
def __exit__(self, *exc):
self.t1, self.e1 = time.time(), self.read_mj()
deltas = [b - a for a, b in zip(self.e0, self.e1)]
if any(d < 0 for d in deltas): # driver reloaded mid-job
raise RuntimeError("energy counter reset; use sampled power instead")
self.gpu_kwh = sum(deltas) / 1000 / 3.6e6 # mJ -> J -> kWh
self.hours = (self.t1 - self.t0) / 3600
with GpuEnergyMeter(range(8)) as m:
train() # your training loop
print(json.dumps({"gpu_kwh": round(m.gpu_kwh, 3), "hours": round(m.hours, 2),
"avg_w_per_gpu": round(m.gpu_kwh * 1000 / m.hours / 8, 1)}))Run the meter on every node of a distributed job and sum the results. Two caveats: the counter updates at intervals of tens of milliseconds, so it is unsuitable for timing individual kernels, and on shared GPUs the device total cannot be split exactly between tenants. For fleet-wide collection, DCGM exports equivalent energy and power fields; the DCGM article covers its exporter.
From GPU to facility
The GPU counter sees only the GPU board. The host CPUs, memory, network cards, storage and fans also draw power, and in an eight-GPU training server they add a substantial fraction on top of GPU energy. The best source for that overhead is the server's own power measurement, from the baseboard management controller or the rack PDU, compared against the summed GPU energy over the same interval. Measure it once per server type under a realistic load and record the ratio as the node overhead factor.
The facility then adds cooling and power-conversion losses, captured by power usage effectiveness (PUE): total facility energy divided by IT energy. A PUE of 1.2 means 20 percent overhead. On premises, use the measured PUE of the site, ideally for the month the job ran. In the cloud, check what the provider's carbon reporting includes: if its figures already include facility overhead, applying your own PUE on top counts it twice.
Grid intensity depends on where and when
Grid carbon intensity varies by an order of magnitude between regions, from grids dominated by hydro, nuclear or wind to grids that still burn a lot of coal, and within one grid it varies by hour as solar and wind come and go. For a job measured in hours, use hourly intensity for the actual hours it ran rather than an annual average. Grid operators and commercial data services publish historical and forecast intensity; pick one source, record which, and use it consistently.
Average intensity answers the question "what was the mix when I ran". Marginal intensity answers "which power plant ramped up because I added load", which is the better guide when deciding whether to shift a job. Marginal data is harder to obtain and is modelled rather than measured, so treat it as an estimate and say so in reports.
Embodied emissions
Manufacturing a GPU server, especially its chips, memory and boards, emits a large one-off amount of CO2e. Published figures vary widely between vendors and methods, and few cover current accelerators in detail, so treat any single number with caution. The allocation, though, is simple: divide the server's embodied total by its expected service hours and charge each job for the hours it occupied, whether or not the GPUs were busy.
That allocation has a direct consequence: idle hardware still accrues embodied emissions. A cluster running at 40 percent utilisation charges more than twice the embodied carbon per useful GPU-hour of one running at 90 percent. Extending hardware life also reduces the per-hour share, which is a real argument for running older GPUs on work they do efficiently instead of retiring them early.
Worked example: one fine-tuning run
A team fine-tunes a model on 64 GPUs across 8 servers for 30 hours. The meters report 1,056 kWh of GPU energy, an average of 550 W per GPU. Their BMC measurements show node overhead of 1.3 for this server type, giving 1,373 kWh at the servers. The on-premises site PUE that month is 1.2, so the facility energy is about 1,647 kWh.
The grid figures below are illustrative, not sourced. At an average intensity of 0.35 kg CO2e per kWh for the hours the job ran, operational emissions are about 577 kg. Had the same job run in a region averaging 0.05 kg per kWh, they would be about 82 kg.
For embodied emissions, again with an illustrative figure, assume 3,000 kg CO2e per server and a five-year life of 43,800 hours. Each server then carries about 0.068 kg per hour, and 8 servers for 30 hours add about 16 kg. In the dirty-grid case embodied carbon is under 3 percent of the total; in the clean-grid case it is about 17 percent. That pattern is typical: on carbon-intensive grids energy dominates, and on clean grids embodied carbon becomes a meaningful share, so utilisation and hardware lifetime matter more there.
With R defined as one run, SCI is about 593 kg CO2e per run on the dirty grid. If the team runs four hyperparameter trials before the final run and keeps none of them, the true cost of the delivered model is five times that, which is why wasted runs appear near the top of the lever list.
The levers, roughly in order of impact
| Lever | How it reduces CO2e | What to watch |
|---|---|---|
| Avoid wasted runs | Fewer failed, duplicate and abandoned jobs | Smoke tests, small-scale sweeps, checkpoint resume |
| Choose region and hour | Lower grid intensity per kWh | Data residency, latency, queue delays |
| Raise utilisation | More work per powered and embodied hour | Input pipeline stalls, small batches, idle reservations |
| Use efficient precision and kernels | Fewer joules per step | Accuracy checks for lower precision |
| Batch inference | More tokens per joule | Latency targets |
| Cap GPU power | Lower energy, often small slowdown | Measure throughput per watt; job takes longer |
| Right-size models | Smaller model for the task | Quality on your evaluation set |
Power capping deserves a measurement rather than a rule of thumb. nvidia-smi -i 0 -pl 450 sets a board power limit in watts (it needs administrator rights and must be within the supported range shown by nvidia-smi -q -d POWER). Many training workloads lose less throughput than they save in power at moderate caps, because performance per watt peaks below maximum power, but the curve depends on the model and GPU. Run the same job for a few hundred steps at several caps and compare joules per step, as the cost article GPU cost optimisation suggests doing for price.
Carbon-aware scheduling
Flexible jobs, such as batch evaluation, embedding backfills, sweeps and many training runs, can wait for cleaner hours or move to a cleaner region. The scheduler needs an intensity forecast, the job's estimated energy and duration, and a deadline. A simple policy picks the start time that minimises forecast emissions within the deadline.
def best_start(forecast, job_hours, deadline_h, now_h=0):
# forecast: list of (hour_offset, kg_per_kwh); job assumed to draw evenly
best = None
for start in range(now_h, deadline_h - job_hours + 1):
window = [i for h, i in forecast if start <= h < start + job_hours]
if len(window) < job_hours:
break # forecast does not reach this far
avg = sum(window) / job_hours
if best is None or avg < best[1]:
best = (start, avg)
return best # (start_offset_hours, avg_intensity)
start, intensity = best_start(forecast, job_hours=6, deadline_h=24)
if intensity > 0.9 * current_intensity: # not worth waiting for <10% gain
start = 0Delaying a job has costs: idle reserved GPUs still accrue embodied emissions and money, and people wait. Carbon-aware scheduling therefore works best for jobs on shared or elastic capacity and pairs naturally with the queueing policies in GPU scheduling.
Reporting per job
Record a carbon line for every job alongside its cost, so trends are visible per team and per model version:
{"job": "ft-legal-v3", "gpus": 64, "hours": 30.0, "gpu_kwh": 1056,
"node_overhead": 1.3, "pue": 1.2, "facility_kwh": 1647,
"intensity_source": "grid-operator hourly", "kg_per_kwh_avg": 0.35,
"operational_kg": 577, "embodied_kg": 16, "embodied_basis": "illustrative 3000 kg/server, 5 y",
"functional_unit": "run", "sci_kg": 593}
Failure modes
| Failure | Effect | Defence |
|---|---|---|
| Counting only GPU energy | Underestimates by the node and facility overhead | Measure node overhead; apply PUE once |
| Annual average intensity | Hides the benefit of time shifting | Hourly data for the hours the job ran |
| Counter reset mid-job | Negative or tiny energy delta | Detect and fall back to sampled power |
| Ignoring failed runs | Reported footprint far below the real one | Charge sweeps and failures to the delivered model |
| Mixing market and location figures | Numbers not comparable over time | State the method on every report |
What to do next
- Wrap your training and batch inference entry points in an NVML energy meter and log kWh per job.
- Measure node overhead for each server type against BMC or PDU readings.
- Get the PUE for your site, or confirm what your cloud provider's figures already include.
- Choose one hourly grid-intensity source, record it, and compute operational CO2e per job.
- Pick a functional unit per workload and publish SCI alongside cost on your job dashboard.
- Run a power-cap sweep on one representative job and adopt the cap with the best joules per step.
- Move one flexible batch workload to carbon-aware scheduling with a deadline.