Crusoe started in 2018 by putting computers where gas would otherwise be burned. Oil wells produce natural gas that often cannot be piped to market, so operators flare it. Crusoe's Digital Flare Mitigation business captured that gas, generated electricity on site and ran modular data centres next to the well; by early 2025 it had deployed more than 425 of them in seven US states and Argentina. In March 2025 Crusoe agreed to sell that business, bitcoin mining included, to NYDIG. The company that remains calls itself a vertically integrated AI infrastructure provider: it builds gigawatt campuses such as the 1.2 GW Abilene, Texas site that hosts capacity for Oracle and OpenAI, and runs Crusoe Cloud, a GPU cloud on top.

So the old name describes a business Crusoe chose to leave, but the idea survived: start from where power is available, then build compute there. This page treats that idea as an engineering input. It explains why power-first siting exists, what it does to latency, data movement, capacity timing and carbon accounting for the jobs you run, and how to drive Crusoe Cloud from its CLI. For neocloud diligence in general see CoreWeave in depth.

Why build compute where the power is

A modern training cluster is a power problem before it is a chip problem. Rack density for current GPU systems runs from tens to well over a hundred kilowatts, so a hall of a few thousand GPUs needs tens of megawatts and a frontier campus needs gigawatts. Grid connections at that scale take years: utilities must study the load, build substations and sometimes transmission. GPUs that sit in a warehouse waiting for power earn nothing, so the scarce input is energised floor space, not silicon. See Power Density in AI Datacenters for the rack-level arithmetic.

An energy-first developer inverts the usual order. Instead of choosing a metro near users and queuing for power, it chooses sites where power can be delivered soon: places with spare grid capacity, cheap wind, or gas that can feed on-site generation behind the meter. Crusoe's public projects show the pattern. The Abilene campus sits on the Lancium Clean Campus in West Texas, a region with large wind resources; in March 2026 Crusoe announced a second, 900 MW Abilene campus for Microsoft next door; a 1.8 GW campus with its own power plant has been approved in Cheyenne, Wyoming; and Spark, its modular product, packages power, cooling and racks into factory-built units that can be placed beside an energy source, including second-life EV batteries from Redwood Materials. Crusoe says its contracted capacity is approaching 5 GW.

What power-first siting leaks into software

Energy-first AI cloud: the layers between a power source and your jobPower sourcegrid, on-site gas, wind,solar, batteriesCampus or SparkGW halls or modular unitsGPU clustersInfiniBand per clusterYour jobVMs, IB partitionWhat siting near power changes for the software above itLatencyusers may be far awayData gravitydatasets must move inCapacity timingpower arrives in phasesCarbondepends on the site mixDesign rule: put latency-tolerant, data-light, long work there firsttraining, fine-tuning, batch inference, evaluation
The stack from power source to job, and the four properties of a power-first site that leak through to software.

Four properties of a power-first site leak upward into your engineering. Latency: the site is chosen for power, not proximity, so it may be far from your users and your other infrastructure. Data gravity: training data, checkpoints and evaluation sets must move to the site and back. Capacity timing: large campuses energise in phases, so the capacity you are promised arrives building by building. Carbon: the emissions of a GPU-hour depend on that site's actual generation mix, which for a campus with on-site gas turbines is not the same as the flare-gas story the company started with.

None of these is a reason to avoid such a provider; each is a reason to place the right work there. Training, fine-tuning, large batch inference and evaluation sweeps tolerate a hundred milliseconds of network distance and run for hours or weeks on data that can be staged once. Interactive inference with a tight time-to-first-token budget is the workload that cares most where the site is.

Latency and data gravity, in numbers

Light in fibre travels about 200 km per millisecond, and real routes are longer than the straight line, so a useful planning figure is roughly 1 ms of round trip per 70 to 100 km of distance. A site 1,500 km from your users adds something like 15 to 25 ms of round trip before any queueing. For a chat product that streams tokens, that is added once to time-to-first-token and is negligible against decode time. For an agent that makes twenty sequential tool calls back to services in another region, it is paid twenty times. Measure your own path: the planning figure is no substitute for a ping from the client network you care about.

Data movement is usually the larger cost. The script below estimates how long it takes to stage a dataset and pull checkpoints back, and how that compares with the run itself. Use your measured sustained throughput, not the port speed.

# staging.py: is data movement a rounding error or the critical path?
def hours_to_move(terabytes, gbit_per_s, efficiency=0.7):
    bits = terabytes * 1e12 * 8
    return bits / (gbit_per_s * 1e9 * efficiency) / 3600

dataset_tb, ckpt_tb, ckpts_kept, link_gbps = 40, 1.1, 3, 25
stage_in = hours_to_move(dataset_tb, link_gbps)
pull_out = hours_to_move(ckpt_tb * ckpts_kept, link_gbps)
run_hours = 96
print(f"stage in {stage_in:.1f} h, pull out {pull_out:.1f} h, run {run_hours} h")
print(f"movement is {100 * (stage_in + pull_out) / run_hours:.0f}% of the run")

With these inputs, 40 TB at 25 Gb/s and 70 percent efficiency takes about 5 hours to stage and three 1.1 TB checkpoints take under half an hour to bring home: about 6 percent of a four-day run. Halve the run and double the dataset and movement grows to over a fifth of the wall clock, which is when you stage data before the reservation begins.

Driving Crusoe Cloud from the CLI

Crusoe Cloud exposes VMs, disks, VPC networking and InfiniBand through a console, an API and a CLI. The CLI is a good model of the resource graph: a VM has a type and a location, multi-node GPU VMs join an InfiniBand partition that belongs to an IB network in that location, and persistent disks attach separately. Do not hard-code VM type strings from a blog post, including this one; list them.

# discover what exists before creating anything
crusoe locations list
crusoe compute vms types
crusoe compute images list
crusoe networking ib-networks list          # IB networks per location
crusoe reservations --help                  # view reserved capacity, if you have any

# one partition per job limits which nodes can reach its nodes over InfiniBand
crusoe networking ib-partitions create --name ft-run-42 --ib-network-id "$IB_NET"

# create the four nodes atomically (named ft-run-42-<number>)
crusoe compute vms create \
  --name ft-run-42 --count 4 \
  --type "$GPU_TYPE" --location "$LOCATION" \
  --image "$IMAGE" --keyfile ~/.ssh/id_ed25519.pub \
  --ib-partition-id "$PARTITION_ID" \
  --startup-script ./node_up.sh \
  --shutdown-script ./node_down.sh \
  --json

crusoe compute vms list --json             # wait until every node is running
crusoe diagnostics vm --help               # collect node diagnostics for support cases

Two flags carry most of the operational weight. --count creates the nodes atomically, so you do not end up with three of four nodes and a half-formed job. --shutdown-script attaches a bash script under 64 KB as a shutdown hook; use it to flush the last checkpoint shard and logs to durable storage rather than trusting local NVMe, and remember a crash will not run it. The startup script, under the same limit, should run a GPU and InfiniBand health check before the node joins the job; FluidStack in depth walks through an acceptance burn-in that works on any provider.

Treat teardown as part of the same script. When the job ends, stop and delete the VMs, delete the IB partition, and list disks to catch any left attached to nothing; an orphaned set of multi-node GPU VMs is the most expensive bug a training team can ship. Wrap the whole flow in one driver that records every resource ID it creates and deletes them in reverse order on exit, including on failure, and use the --json output so the driver parses IDs rather than scraping tables.

Worked example: a 96-hour fine-tune far from home

Suppose you plan a 96-hour fine-tune of a 70-billion-parameter model on 64 GPUs at a Crusoe site 1,500 km away. With mixed-precision Adam the full training state is about 16 bytes per parameter, so a checkpoint is about 1.1 TB. Assume writing one takes 120 seconds and that, from your history, a 64-GPU job is interrupted about once every 12 hours. Young's approximation for the checkpoint interval is the square root of 2 x 120 s x 43,200 s, about 3,220 seconds, so checkpoint every 50 to 55 minutes and expect to lose about half an interval of work per interruption.

Staging the 40 TB dataset takes about 5 hours, so it starts the day before the reservation. Latency is irrelevant to the run itself; it matters only for the evaluation service that scores checkpoints, which you either run at the same site or feed with checkpoints copied home once a day. The cost comparison against a hyperscaler region near your data must include egress for every byte that comes back; Bare Metal vs Cloud GPU Cost Analysis gives the utilisation arithmetic to finish the comparison.

Energy and carbon accounting

Do not assume a GPU-hour from an energy-first provider is low-carbon, and do not assume it is high-carbon either. The flare-gas business displaced emissions that would have happened anyway; a campus with on-site gas generation, grid supply, wind, solar and batteries has a mix that varies by site and by hour. If you report emissions, ask the provider for per-site, location-based figures and compute your share from energy actually used:

def job_emissions_kg(gpu_hours, gpu_kw=1.0, node_overhead=0.35, pue=1.2, kg_per_kwh=0.40):
    """gpu_kw: average measured draw per GPU; node_overhead: CPUs, NICs, fans as a fraction;
    kg_per_kwh: the site's location-based intensity from the provider, not a national average."""
    kwh = gpu_hours * gpu_kw * (1 + node_overhead) * pue
    return kwh, kwh * kg_per_kwh

kwh, kg = job_emissions_kg(64 * 96)
print(f"{kwh:,.0f} kWh, {kg/1000:.1f} t CO2e at the assumed intensity")

Every default in that function is a placeholder. Use measured GPU power from nvidia-smi or DCGM over the run, the site's reported PUE, and the provider's intensity figure; report the inputs alongside the number.

Failure modes

  • Promised capacity arrives late. Phased campuses energise building by building. Write start dates and remedies into the contract and keep a fallback provider warm.
  • Partial node sets. Creating nodes one at a time leaves half-formed jobs. Use --count and check every node before launching.
  • Unflushed checkpoints on stop. Local NVMe disappears with the VM. Flush in the shutdown script and checkpoint to durable storage on a schedule.
  • Over-broad InfiniBand membership. Nodes left in a shared partition can reach other jobs' nodes. Give each job its own partition and delete it afterwards; note that partitions control membership, not bandwidth, so jobs can still share links.
  • Egress surprises. Pulling every checkpoint home can rival compute cost. Keep checkpoints at the site and copy only the ones you will keep.
  • Latency-bound agents placed far away. Sequential calls multiply round trips. Co-locate the agent loop with the model or with its tools, not split between them.

Trade-offs

A power-first provider can offer capacity when metro regions cannot, often in large contiguous blocks with a modern fabric, because it solved the energy problem first. You pay in distance and in dependence on one company's build schedule, and you take on the job of moving data and proving carbon claims. Hyperscaler regions are closer to your data and services and carry more managed products; a neocloud like Lambda or Crusoe is usually simpler, more GPU-focused and quicker to get large reservations from. Place work by its sensitivity to latency and data movement, not by brand.

What to do next

  1. Classify each workload by latency sensitivity and data volume; send long, latency-tolerant, data-light work to power-first sites first.
  2. Measure round trip and sustained throughput from your data source to the candidate location before signing anything.
  3. Run the staging estimate; if movement exceeds about a tenth of the run, stage early.
  4. Script the CLI flow: one IB partition per job, --count for atomic creation, health checks in the startup script, flush in the shutdown script.
  5. Set checkpoint cadence from Young's formula using your measured interruption rate.
  6. Ask for per-site energy and emissions data and compute job emissions from measured power.
  7. Write capacity delivery dates and remedies into the contract.
Key takeaway: Crusoe began by turning flared gas into compute, agreed to sell that business in 2025 and now builds gigawatt AI campuses where power can be delivered fast. For your jobs, that means capacity in large blocks at sites chosen for energy, not proximity: place latency-tolerant, data-light work there, stage data early, isolate jobs in their own InfiniBand partitions, and verify carbon claims per site.