Voltage Park launched in October 2023 with an unusual origin story. The Navigation Fund, a philanthropic foundation started by Stellar co-founder Jed McCaleb, put about $500 million into roughly 24,000 NVIDIA H100 GPUs, to rent them to AI developers at prices aimed below the big clouds. Press coverage often called it a non-profit cloud. The more precise description is a for-profit company, Voltage Park, Inc., owned by a 501(c)(3) foundation that receives its profits.
The company kept growing. It acquired the TensorDock GPU marketplace in March 2025 and deployed VAST Data storage across most of its sites. In January 2026 it merged with Lightning AI, the company behind PyTorch Lightning, in a deal reported at over $2.5 billion with about 35,000 GPUs across six US data centres. This article explains what you actually rent there, how the hardware maps onto a training job, and how to check a node in its first hour. It also shows how to choose between Ethernet and InfiniBand capacity, and how to estimate the bill before you start. The commercial details move quickly, so every fact below carries its date. Check current terms with the merged company before you sign anything.
What Voltage Park is
Three things set Voltage Park apart from other neoclouds.
It owns its hardware and was funded by a grant, not debt. Most GPU clouds borrow against their GPUs or their contracts. Reports at the time of the merger described Voltage Park as backed by about $900 million in grants from the Navigation Fund. For a buyer, that shows up mostly as pricing posture and stability, not as a feature you can call.
It sells bare metal. The fleet is HGX H100 servers, reported as Dell PowerEdge XE9680 with H100 SXM5 80 GB GPUs. They are connected with NVIDIA Quantum-2 InfiniBand and sit in Tier 3+ facilities in Washington, Texas, Virginia and Utah. You get the whole machine, with no hypervisor between your process and the GPUs.
It is now part of a software company. Since the Lightning AI merger, the capacity is sold alongside Lightning's development platform, and the old documentation domain now redirects to a Lightning-branded help centre. If you are evaluating it, ask which console, API and contract you will be on. Do not assume older product pages still describe them.
For comparison with providers built on the same hardware but with different commercial models, see CoreWeave and Lambda Labs, in depth.
The node a job actually sees
An HGX H100 node has eight H100 SXM GPUs, each with 80 GB of HBM3. Every pair of GPUs talks through NVSwitch at up to 900 GB/s per GPU over NVLink 4. That is where tensor parallelism belongs. Leaving the node goes through eight ConnectX-7 adapters at 400 Gb/s each, one per GPU, for 3200 Gb/s per node. That is the 3200 Gbps figure on Voltage Park's InfiniBand tiers. If the fabric is rail-aligned, as NVIDIA's reference designs are, GPU k on every node connects to the same leaf switch and most NCCL traffic stays one switch hop away. Ask the vendor how its fabric is cabled. The layout is explained in DGX H100 architecture.
Storage is a separate path. VAST's shared file, object and block namespace holds datasets and checkpoints, so any node can resume a job that another node started. Local NVMe is still the fastest scratch space for a shard you read every epoch.
Ways to buy
The product page, checked in October 2026, lists the offers below.
| Offer | Shape (product page, October 2026) | Fits |
|---|---|---|
| On-demand, Ethernet | 1 to 1016 H100 GPUs, no minimum term, about 15 minutes to deploy | single-node fine-tuning, inference, experiments |
| On-demand, InfiniBand | 8 to 1016 H100 GPUs on the 3200 Gbps fabric | multi-node training that is short or bursty |
| Reserved | 32 to 8000+ GPUs, contracts of 6+ or 12+ months | pre-training and steady production load |
| TensorDock marketplace | third-party hosts, including consumer GPUs | cheap inference and hobby work |
The pages also listed managed Kubernetes, virtual machines, storage and observability, and mentioned Blackwell systems. Later reporting describes the fleet as Hopper, Blackwell and Grace Blackwell. Treat any per-hour price you find as a point-in-time quote. Third-party listings during 2025 showed Ethernet H100 on-demand around $1.99 per GPU-hour and disagreed about the InfiniBand premium. Get a written quote for your own shape and term.
Ethernet or InfiniBand: do the arithmetic
The Ethernet versus InfiniBand choice is really a question about one number: how many bytes each node must exchange per optimizer step, compared with how fast it can move them. For data-parallel training with a hierarchical all-reduce, each node sends and receives about 2S(k - 1)/k bytes per step across its NICs, where S is the gradient size and k is the number of nodes.
Take an 8-billion-parameter model with bf16 gradients, so S = 16 GB, on k = 8 nodes. Each node moves about 28 GB per step. On the InfiniBand tier, eight 400 Gb/s NICs give 400 GB/s per node, so the exchange takes about 0.07 s at line rate. Suppose an Ethernet node had 100 Gb/s, which is 12.5 GB/s. Voltage Park does not publish that number, so measure it. The same exchange would then take about 2.2 s, which could easily exceed the compute time of the step. Gradient accumulation and sharded optimizers change the constant, but not the conclusion. Single-node jobs and inference do not care about the fabric. Multi-node training needs InfiniBand, and paying for it is cheaper than leaving GPUs idle. For link-level detail see InfiniBand NDR and XDR in depth.
The first hour on a new node
Bare metal means you inherit the hardware's faults as well as its speed. Spend the first hour of any new allocation proving that each node is healthy, before a 64-GPU job finds a bad one for you. The commands below are standard NVIDIA and Mellanox tools, not vendor-specific ones.
#!/usr/bin/env bash
# Run on every node before admitting it to a job.
set -euo pipefail
nvidia-smi --query-gpu=index,name,memory.total,ecc.errors.uncorrected.volatile.total \
--format=csv # 8 x H100 80GB, zero uncorrected ECC
nvidia-smi topo -m # expect NV18 between every GPU pair
ibstat | grep -E "State|Rate" # IB tiers: 8 ports Active, Rate 400
dcgmi diag -r 3 # medium hardware diagnostic, all PASS
# Single-node NCCL bandwidth (github.com/NVIDIA/nccl-tests)
./build/all_reduce_perf -b 1G -e 8G -f 2 -g 8
# Across two nodes: one rank per GPU, launched with MPI
mpirun -np 16 -N 8 --hostfile hosts ./build/all_reduce_perf -b 1G -e 8G -f 2 -g 1Record the bus bandwidth that nccl-tests reports, per node and per node pair. On the InfiniBand tier, inter-node bus bandwidth should come close to the 50 GB/s per GPU that a 400 Gb/s NIC can carry. A pair that comes in far below that points to a cable, a port or a misconfigured NIC, and a ticket with numbers in it gets fixed faster. Keep these baselines. They are your reference for every incident later.
Launching a multi-node job
Most problems on a new cluster come from NCCL picking the wrong interface. Set the choice explicitly and turn logging on for the first runs.
# InfiniBand nodes: use the RDMA adapters for collectives,
# and a named Ethernet interface only for bootstrap traffic.
export NCCL_DEBUG=INFO
export NCCL_IB_HCA=mlx5
export NCCL_SOCKET_IFNAME=<bootstrap-interface> # find it with: ip -br link
# Ethernet-only nodes without RDMA: tell NCCL not to look for InfiniBand.
# export NCCL_IB_DISABLE=1
torchrun --nnodes=8 --nproc-per-node=8 \
--rdzv-backend=c10d --rdzv-endpoint=$HEAD_NODE:29500 --rdzv-id=job42 \
train.py --ckpt-dir /mnt/shared/ckpt/job42With NCCL_DEBUG=INFO, each rank logs which network transport it chose. On InfiniBand nodes the lines should start with NET/IB. Fail the job if any rank reports NET/Socket. Interface names differ between images, so read them from ip -br link rather than copying them from another cloud.
Worked example: pricing a fine-tuning run
Estimate the bill before you reserve anything. Training compute is about 6ND FLOPs for N parameters and D tokens. Divide that by the GPU's dense peak times the model FLOPs utilisation (MFU) you can really reach, and you get GPU-seconds.
def gpu_hours(params, tokens, mfu, peak_flops=989e12): # H100 SXM, dense BF16
return 6 * params * tokens / (peak_flops * mfu) / 3600
base = gpu_hours(8e9, 20e9, mfu=0.40) # 8B model, 20B tokens
total = base * 1.2 # restarts, evals, idle time
nodes = 8
print(f"{base:.0f} GPU-h, {total:.0f} with overhead, "
f"{total / (nodes * 8):.1f} h wall on {nodes} nodes")
# 674 GPU-h, 809 with overhead, 12.6 h wall on 8 nodes
def cost(gpu_h, price_per_gpu_hour): # use your quoted price
return gpu_h * price_per_gpu_hourUse the dense BF16 peak of about 989 TFLOPS, not the 1,979 TFLOPS figure that assumes structured sparsity. Well-tuned dense training on H100s is commonly reported in the 30 to 50% MFU range. Measure yours on a short run before you trust the estimate. At a quoted $2 per GPU-hour this job is about $1,600. At that size, on-demand InfiniBand is clearly the right purchase. A reservation only makes sense when the same arithmetic, summed over months of planned work, keeps a block of GPUs busy.
Checkpoints are the other cost. Mixed-precision Adam keeps about 14 bytes per parameter: bf16 weights, fp32 master weights and two fp32 moments. That is 112 GB for 8B parameters. Divide by the write throughput you measure to the shared filesystem to get the stall per checkpoint, and pick an interval that keeps the stall to a few percent of runtime.
Failure modes
- A silent slow node. One GPU with throttled clocks or one degraded NIC slows every synchronous step. Per-node nccl-tests baselines and per-step timing logs find it, and so does a DCGM health watch.
- Wrong transport. NCCL falls back to TCP sockets and the job runs, just several times slower. Check the INFO log at startup.
- Buying Ethernet for multi-node training. The fabric arithmetic above shows how communication can swamp compute. Test your actual model at your actual node count before you commit.
- Checkpoints on local disk. A node that disappears takes its NVMe with it. Write checkpoints to the shared namespace and test a restore on a different node.
- Assuming last year's terms. After the merger, the console, billing and support paths may differ from the 2025 documentation. Confirm them in writing.
- Single-provider dependence. Keep container images, data and launch scripts portable, as described in GPU provider failover, so capacity problems become an inconvenience rather than an outage.
Trade-offs
Against a hyperscaler, Voltage Park offers bare-metal H100s with a 3200 Gb/s InfiniBand option and simpler pricing. It gives up the managed-service catalogue, the global regions and the compliance paperwork that large enterprises often need. Against other neoclouds, its distinguishing features were ownership and funding, not architecture: the hardware recipe of HGX, Quantum-2 and a parallel filesystem is industry standard. Since the merger, the open question is how much of the platform becomes Lightning-shaped. That could suit teams already using Lightning and Studios, and suit less well teams that want a plain rack of nodes and a Slurm login. Nebius makes an instructive comparison of a different integrated model; see Nebius, in depth.
What to do next
- Write down your job's shape: model size, tokens, node count and checkpoint size. Then run the GPU-hour and all-reduce arithmetic above with your own numbers.
- Ask the merged company for a written quote that covers the fabric, minimum term, storage, egress and support for that shape.
- On the first allocation, run the validation script on every node and save the nccl-tests baselines.
- Set the NCCL interface variables explicitly and fail the job if any rank reports a socket transport.
- Checkpoint to shared storage, and test a restore on a different node before the long run.
- Keep images and data portable to a second provider, and rehearse moving a job once.