Where a GPU sits decides much more than its electricity bill. It decides whether a chat response arrives before the user loses patience, whether a training run can use the dataset without a transfer bill larger than the compute, whether a customer's data is allowed to be processed at all, and whether the cluster can grow next year. Placement is often treated as a real estate decision made once by a facilities team. For AI workloads it is an architecture decision, and the software team should be in the room.

This article explains placement from first principles: the different needs of training and inference, the hard constraints that remove sites from consideration, the factors worth scoring, and a small model you can run to make the trade-offs explicit. What happens inside the building is covered in GPU datacenter deployment and the electrical chain in GPU datacenter power; this page is about choosing the building and deciding what runs in it.

Advertisement

Placement is a workload question first

No site is best for everything, because AI workloads pull in opposite directions. Large pretraining runs need many accelerators in one place on a fast fabric, lots of power for months, and almost nothing from the end user; a few hundred milliseconds of distance to the user is irrelevant to them. Interactive inference is the reverse: each request uses little compute, but it must arrive close enough to users that network round trips fit inside the response budget. Between those sit fine-tuning, evaluation, batch inference and retrieval-augmented serving, each with its own mix of needs.

WorkloadNeeds mostToleratesUsually placed
PretrainingContiguous power and fabric, long-term capacityHigh user latency, remote locationWhere power is plentiful and cheap
Fine-tuning and evaluationAccess to proprietary data, moderate scaleSome latencyNear the data, often a regional hub
Batch inferenceCheap compute, throughputHours of delayWherever spare capacity is, follow the price
Interactive inferenceLow round-trip time to usersHigher cost per GPU hourMetro regions near user populations
Regulated inferenceProcessing in a specific jurisdictionHigher cost, smaller poolsInside the required country or region

So placement strategy starts with an inventory of workloads, each with a size, a latency budget, a data location and a residency rule, not with a list of attractive regions.

The constraints that remove sites

Placement is a two-step decision: filter sites by hard constraints, then score the rest per workloadCandidate sitesexisting and proposedHard constraintspower date, residency, fibreWeighted scorecost, carbon, latency, riskRemote, power-rich sitepretraining, batch inferenceRegional hubfine-tuning, RAG, async APIsMetro edgeinteractive inferenceCheckpoints, datasetsbulk replicationModel weightspushed to serving sitesUser trafficrouted by latency budgetTraining follows power; interactive inference follows users; data residency overrides both.
Hard constraints filter the candidate list first; a weighted score ranks what remains, separately for each workload class.

Some factors are not trade-offs; they disqualify. Treat them as filters before any scoring, because a weighted score lets a great price compensate for an impossible requirement, which is how bad decisions get approved.

  • Power by the date you need it. The utility connection is the gating factor for large AI sites, and new connections can take years to arrive. A site whose power arrives after your capacity deadline is not a candidate, however cheap it is.
  • Data residency and sovereignty. If contracts or law require processing in a jurisdiction, sites outside it are excluded for that workload. Note that this is per workload: the same company may train globally and serve regulated customers locally.
  • Latency budget. For interactive workloads, a site whose round trip to the user population exceeds the network share of the budget is excluded.
  • Network reach. Diverse fibre routes into the site and enough capacity to move checkpoints and datasets.
  • Physical risk. Flood plains, seismic zones and grids with frequent curtailment may be acceptable with mitigation, but some are deal-breakers.
Advertisement

Power: the first and hardest constraint

A GPU cluster is sized in megawatts. The accelerators, hosts, network and cooling together set a facility load, and the grid has to deliver it continuously. Training adds a twist covered in the power article: synchronised steps make the load swing quickly between high and low, which some grids and on-site systems handle better than others.

Three questions decide whether a site's power is usable. How much can be delivered, and when? What does it cost per megawatt-hour, including network charges and whether the price is fixed or tracks a volatile market? And how clean is it, measured as grid carbon intensity over the hours you will run, not as an annual average? Sites with abundant hydro, nuclear or wind are attractive for training because those workloads can go where the power is. For flexible batch work, some operators also shift jobs in time towards cleaner or cheaper hours, which only works if the scheduler knows the site's price and carbon signals.

Climate, cooling and water

Cooling turns power into a second cost and sometimes a water bill. Cooler and drier climates allow more hours of economiser cooling, using outside air or water instead of mechanical chillers, which lowers facility overhead. That advantage is real but smaller than it used to be for dense GPU racks, because direct liquid cooling can run at higher coolant temperatures that can be rejected to warmer outside air. The thermal design is covered in GPU datacenter cooling and liquid cooling for GPU datacenters.

Climate cuts both ways. Hot regions can still host GPUs with suitable cooling designs, at a higher overhead, and some hot regions offer cheap power that outweighs the cooling penalty. Water is an increasingly binding constraint: evaporative cooling uses a lot of it, and local water stress can block permits or invite public opposition. Score water use explicitly rather than assuming a cool climate solves it.

The physics of latency

Light in optical fibre travels at roughly two thirds of its speed in vacuum, about 200 kilometres per millisecond, or 5 microseconds per kilometre one way. A round trip therefore costs about 1 millisecond per 100 kilometres of fibre path, before any routing, queueing or processing. Real fibre paths are longer than straight lines, often by a substantial factor, so measure routes rather than drawing circles on a map.

For interactive LLM serving, the network round trip matters most for time to first token and for chatty agent loops that make many sequential calls. A token stream that takes several seconds to generate barely notices 30 milliseconds of distance; an agent that makes twenty sequential tool and model calls across regions pays that distance twenty times. The estimate below is simple enough to put in a planning spreadsheet.

FIBRE_US_PER_KM = 5.0           # one way, propagation only

def network_rtt_ms(route_km: float, overhead_ms: float = 4.0) -> float:
    # overhead covers routing, TLS reuse misses and load balancers; measure your own
    return 2 * route_km * FIBRE_US_PER_KM / 1000 + overhead_ms

def fits_budget(route_km, sequential_calls, ttft_budget_ms, model_ttft_ms):
    network = sequential_calls * network_rtt_ms(route_km)
    return network + model_ttft_ms <= ttft_budget_ms, round(network, 1)

# Single chat turn vs a 20-step agent loop, 2,500 km of fibre route
print(fits_budget(2500, 1, 800, 450))    # (True, 29.0)
print(fits_budget(2500, 20, 800, 450))   # (False, 580.0)

The lesson is that latency sensitivity is a property of the application's call pattern, not just of the model. Co-locate the orchestration loop with the model it calls, and the distance is paid once per user turn instead of once per step.

Data gravity, residency and egress

Large datasets and checkpoints are expensive to move. Copying a multi-hundred-terabyte corpus between sites takes days over a dedicated link and can cost more in transfer fees than the compute that reads it. That pulls fine-tuning and evaluation towards where the data already lives, and it makes the pretraining site's storage and connectivity as important as its power.

Residency adds legal edges to data gravity. Processing personal data may be restricted to a jurisdiction, and prompts and retrieved documents are data too, as are KV caches and logs. A serving region for regulated customers must keep the whole request path in scope, including the vector store and the logging pipeline, which is why multi-region serving pins conversations to a region. Model weights usually move in the other direction: trained centrally, pushed out to serving sites, so weight distribution bandwidth and cold-start time are part of every serving site's cost.

Training across more than one site

When no single site can host a whole run, the question becomes whether training can span sites. Within a site, data-parallel gradient synchronisation uses a fabric with very high bandwidth and microsecond latency. Between sites, bandwidth is far lower and latency is milliseconds, so naive synchronous data parallelism across sites spends most of each step waiting.

The practical patterns keep the heavy, frequent communication inside a site. Tensor and pipeline parallelism stay within a site's fabric. Data parallelism across sites can be made hierarchical, reducing within each site first and exchanging only once between sites, and research approaches synchronise across sites less often, at some cost in convergence that has to be measured for each model. For placement, the implication is that sites meant to train together need dedicated, high-capacity links between them, and that a cluster split across two buildings is not equivalent to one cluster of the same size.

A placement model in code

The model below makes the two-step decision explicit: hard filters per workload, then a weighted score over normalised factors. The weights are policy, not physics; writing them down is the point, because it lets finance, security and engineering argue about numbers instead of adjectives.

from dataclasses import dataclass

@dataclass
class Site:
    name: str
    mw_available: float       # by the required date
    usd_per_mwh: float
    g_co2_per_kwh: float
    route_km_to_users: float
    jurisdictions: set
    water_stress: float       # 0 (none) .. 1 (severe)
    risk: float               # 0 .. 1, physical and grid risk

@dataclass
class Workload:
    name: str
    mw_needed: float
    max_rtt_ms: float | None      # None: latency-insensitive
    must_be_in: set | None        # residency rule
    weights: dict                 # factor -> weight

def feasible(s: Site, w: Workload) -> bool:
    if s.mw_available < w.mw_needed:
        return False
    if w.must_be_in and not (s.jurisdictions & w.must_be_in):
        return False
    if w.max_rtt_ms is not None and network_rtt_ms(s.route_km_to_users) > w.max_rtt_ms:
        return False
    return True

def norm(values):
    lo, hi = min(values), max(values)
    return [0.0 if hi == lo else (v - lo) / (hi - lo) for v in values]   # 0 best, 1 worst

def rank(sites, w: Workload):
    ok = [s for s in sites if feasible(s, w)]
    if not ok:
        return []
    factors = {
        "cost": norm([s.usd_per_mwh for s in ok]),
        "carbon": norm([s.g_co2_per_kwh for s in ok]),
        "latency": norm([s.route_km_to_users for s in ok]),
        "water": [s.water_stress for s in ok],
        "risk": [s.risk for s in ok],
    }
    scored = [(sum(w.weights.get(f, 0) * factors[f][i] for f in factors), s.name)
              for i, s in enumerate(ok)]
    return sorted(scored)          # lowest penalty first

Worked example

Take three illustrative candidate sites. Site A is remote, with 120 MW available next year, low-cost hydro power, low carbon, and 3,000 km of fibre to the main user base. Site B is a regional hub with 40 MW, mid-priced mixed-grid power, and 600 km to users, inside the customer's required jurisdiction. Site C is a metro site with 8 MW, expensive power and 50 km to users, also inside the jurisdiction. All figures are invented for the example.

A 100 MW pretraining workload with no latency or residency rule is feasible only at A, which settles it before any scoring. A 30 MW fine-tuning workload on regulated customer data must stay in the jurisdiction, which removes A, and B wins on cost and power over C, which lacks capacity anyway. An interactive serving workload of 6 MW with a 15 millisecond network budget excludes A at about 34 milliseconds round trip by the estimate above; B at about 10 milliseconds and C at about 4.5 both pass. With cost weighted heavily, B wins; with latency weighted heavily, C wins. That last choice is a genuine trade-off, and the model's value is that it shows exactly which weight decides it.

The resulting strategy is typical: train where power is, fine-tune where the data is allowed to be, and serve from a mix of a regional hub for most traffic and a small metro footprint for the most latency-sensitive paths.

Failure modes and trade-offs

  • Choosing on price alone. A cheap site whose power arrives late, or which lacks fibre diversity, delays the whole roadmap.
  • One site for every workload. Either inference latency suffers or training pays metro prices for power it does not need near users.
  • Ignoring the call pattern. Agents that loop across regions multiply distance; co-locate orchestration with the model.
  • Residency scoped too narrowly. The GPU is in-region but the logs, vector store or KV cache spill elsewhere.
  • Treating two buildings as one cluster. Cross-site links cannot carry intra-cluster traffic patterns.
  • Annual averages. Carbon and price vary by hour; flexible workloads should be scheduled on hourly signals.
  • Water as an afterthought. Local opposition and permits can stop expansion at an otherwise ideal site.

What to do next

  1. Inventory workloads with size in megawatts, latency budget, call pattern, data location and residency rule.
  2. List candidate sites with power-by-date, price, hourly carbon, measured fibre routes, water stress and risk.
  3. Apply hard filters per workload before scoring anything, and write down why each site was excluded.
  4. Agree weights with finance, security and engineering, and run the scoring model; review which weight decides close calls.
  5. Check the full data path of regulated workloads for residency, not just the GPUs.
  6. Plan checkpoint, dataset and weight replication bandwidth between the chosen sites.
  7. Revisit placement each planning cycle, because power availability and workloads change faster than buildings.
Key takeaway: GPU placement is a workload-matching problem: pretraining follows abundant power, interactive inference follows users, fine-tuning follows data, and residency overrides all of them. Filter sites by hard constraints such as power-by-date, jurisdiction and latency budget before scoring, then rank the rest with explicit weights for cost, carbon, latency, water and risk. Measure latency against the application's call pattern, plan data movement deliberately, and keep frequent training communication inside one site.