Fluidstack is one of the "neoclouds": companies that build and operate large, dedicated accelerator clusters for AI labs and enterprises, rather than selling general-purpose cloud services. It became widely known in November 2025, when Anthropic announced a roughly $50 billion US data-centre investment with Fluidstack as its partner for facilities in Texas and New York, due to come online through 2026. Public reporting describes a model in which Fluidstack leases land and power from site owners, often converted crypto-mining campuses, then deploys and operates the accelerators, network and storage inside them. It has also been reported as a partner for hosting Google TPUs, with Google providing financial backstops on some site leases.
This article is not a news summary. Company specifics move fast, and the figures in press coverage (valuations, GPU counts, megawatts) conflict between sources, so they are left out. Instead it explains what a dedicated cluster from a provider like Fluidstack actually is, how the hardware is laid out and why that matters to your training job, how to accept a cluster before you start paying for idle time, how to schedule and checkpoint on it, and which contract and operational questions decide whether the deal works. The same checklist applies to any neocloud, and the commands are the standard NVIDIA, InfiniBand and Slurm tools rather than any one provider's console.
The product: a dedicated cluster, not instances
It helps to separate three ways of buying accelerator time. On-demand instances from a hyperscaler give you a few nodes quickly, with shared networking and no promise that the next 64 nodes will be on the same fabric. Marketplaces give you cheap, heterogeneous capacity with variable quality. A dedicated cluster, which is the core neocloud product, is a single-tenant block of nodes on one non-blocking fabric, reserved for a term, with a managed scheduler on top. Fluidstack describes its large deployments this way: dedicated clusters, NVIDIA parts of recent generations, and managed Kubernetes or Slurm operated by its own engineers.
For a training team the difference is not the GPU model on the invoice. It is whether all-reduce across 512 GPUs runs near line rate, whether a failed node is replaced in hours or days, and whether the scheduler knows which nodes share a leaf switch. Those are properties of the cluster build and the operator, which is why acceptance testing matters more than the spec sheet.
Inside the cluster: nodes, NVLink and rails
A typical NVIDIA HGX-class node has eight GPUs connected to each other by NVLink through NVSwitch, and one high-speed network adapter per GPU for the back-end (east-west) fabric. A separate front-end network carries storage, SSH and control traffic. In a rail-optimised design, GPU k of every node connects to the same leaf switch (rail k), so the inter-node legs of a ring or tree collective, where GPU k talks to GPU k on other nodes, stay within one switch tier as far as possible.
Why this matters to software: NCCL builds its rings and trees from the topology it detects. Inside a node, traffic goes over NVLink at far higher bandwidth than any network link. Between nodes it goes over the back-end fabric, and if your job's nodes are scattered across spine switches, the collective's slowest hop is the oversubscribed one. Tensor parallelism should stay inside the NVLink domain; data and pipeline parallelism can cross the fabric. On rack-scale systems such as GB200 NVL72 the NVLink domain is the whole rack rather than one 8-GPU node, which changes how large a tensor-parallel group can be before it touches the network.
Rack-scale parts also change operations. Higher power density means liquid cooling, so a cooling fault can take out a rack, not a node. Ask how the provider handles a rack-level failure and whether spare capacity is held in the same NVLink domain size.
Acceptance: burn in before you pay for idle
Do not start the reservation clock on a promise. Agree in the contract that the cluster is accepted only after a burn-in that you can reproduce, and run it yourself on day one. The tools are standard:
# 1. Topology and drivers on every node (look for NV18/NVLink, correct NIC-GPU affinity)
srun -N $NODES --ntasks-per-node=1 nvidia-smi topo -m
srun -N $NODES --ntasks-per-node=1 nvidia-smi --query-gpu=name,driver_version,ecc.errors.uncorrected.volatile.total --format=csv
# 2. GPU health: DCGM diagnostics (level 3 runs the longer stress tests)
srun -N $NODES --ntasks-per-node=1 dcgmi diag -r 3
# 3. InfiniBand ports up at the expected rate
srun -N $NODES --ntasks-per-node=1 ibstat | grep -E "State|Rate"
# 4. Collective bandwidth across the whole allocation (nccl-tests)
srun -N $NODES --ntasks-per-node=8 --gpus-per-node=8 \
./build/all_reduce_perf -b 8M -e 8G -f 2 -g 1For a multi-node run nccl-tests must be built with MPI=1; without it each rank runs its own single-GPU test and the numbers look plausible but measure nothing. Read the output by its bus bandwidth column at large message sizes, not by algorithm bandwidth; bus bandwidth is normalised so that it can be compared with the per-GPU link rate regardless of GPU count. Then do the important part: run the same test on every pair and every small group of nodes, not just the full allocation. A single node with a degraded cable or a misconfigured adapter drags the whole ring down to its speed, and the full-cluster number hides which node it is. Bisecting with node subsets finds it.
Finish acceptance with a real workload: a short run of your own model at production scale for a few hours, logging step time per rank. A healthy cluster gives a tight step-time distribution; a long tail on specific ranks points to a thermally throttling GPU or a slow link that the synthetic tests missed. Record everything as a baseline, because you will compare against it every time a node is replaced.
Scheduling on managed Slurm
Managed Slurm is the common interface for large training jobs. Two settings matter more than the rest. Topology awareness, usually through Slurm's tree topology plugin and a topology.conf that mirrors the switch hierarchy, lets the scheduler pack a job onto nodes under as few switches as possible. Check that the provider has filled it in correctly, then use --switches to express a preference:
#!/bin/bash
#SBATCH --job-name=pretrain
#SBATCH --nodes=64
#SBATCH --ntasks-per-node=8
#SBATCH --gpus-per-node=8
#SBATCH --switches=4@00:30:00 # at most 4 leaf switches, wait up to 30 min for that
#SBATCH --exclusive
export NCCL_DEBUG=WARN
srun python train.py --resume-from latestThe second is how failed nodes leave the pool. A good operator runs health checks between jobs and drains bad nodes automatically, so your next job does not land on them. Ask what checks run, how often, and whether you can add your own. On Kubernetes the same questions apply, expressed as node labels for topology, a gang scheduler so that a 512-GPU job starts all-or-nothing, and node conditions that cordon unhealthy machines. For the Slurm side in depth, see GPU scheduling with Slurm.
Getting data in and checkpoints out
Compute is only useful if data arrives fast enough. A dedicated cluster usually offers a shared parallel file system, an object store, and local NVMe on each node. Each has a role. Training data streams best from the object store or parallel file system in large sequential shards (WebDataset-style tar files or similar), prefetched by several loader workers per GPU, so that random small-file reads never sit on the critical path. Local NVMe is the right first landing spot for checkpoints: write there, release the training loop, then copy to durable storage in the background.
Plan the initial data transfer as its own project. Moving hundreds of terabytes into a new site can take days over the provider's internet uplink, and the reserved cluster bills from day one. Start the transfer before acceptance testing finishes, verify checksums on arrival, and keep a manifest so a partial copy is obvious. Ask whether the site has direct connectivity to your cloud region and what egress costs on the way out, because final weights and evaluation artefacts have to leave the cluster too. Finally, measure loader throughput per node in isolation before the first large run: a data pipeline that delivers 80% of the samples the GPUs can consume caps utilisation at 80%, and no amount of fabric tuning will show up in step time.
Checkpoint cadence from failure rate
At thousands of GPUs, hardware failure is a when, not an if, and the job design must assume it. The standard approach is to pick a checkpoint interval from the observed failure rate. If a checkpoint takes C seconds to write and the cluster's mean time between job-killing failures is M, Young's approximation gives an interval near the square root of 2CM. With C = 60 s and M = 8 hours (28,800 s) that is about sqrt(3,456,000), roughly 1,860 seconds, so checkpoint every 30 minutes or so; halving M only shrinks the interval by a factor of about 1.4.
import math
def checkpoint_interval(write_seconds, mtbf_seconds):
"""Young's first-order optimum for periodic checkpointing."""
return math.sqrt(2 * write_seconds * mtbf_seconds)
def goodput(interval, write_seconds, mtbf_seconds, restart_seconds):
# fraction of wall time spent on useful steps (first-order model)
lost = write_seconds / interval + (interval / 2 + restart_seconds) / mtbf_seconds
return max(0.0, 1.0 - lost)
t = checkpoint_interval(60, 8 * 3600) # ~1859 s
print(round(t), round(goodput(t, 60, 8 * 3600, 600), 3))The model makes the levers explicit. Faster checkpoint writes (asynchronous saves to local NVMe, then upload) shrink C. Faster restart, meaning a warm spare node and a scheduler that requeues automatically, shrinks the restart term. Measure M from your own job logs over the first weeks rather than trusting a datasheet, and revisit the interval when the cluster grows.
Contract and SLA questions
| Question | Why it matters | Good answer looks like |
|---|---|---|
| What is the acceptance test and who runs it? | Billing for a broken cluster | Reproducible burn-in with nccl-tests and DCGM; billing starts after pass |
| Node replacement time? | Each failed node idles a whole job | Hot spares on site; hours, stated in the SLA |
| What counts as available? | SLA credits depend on the definition | Per-node availability with degraded performance counted as down |
| Topology and switch layout? | Collective performance | Documented fabric; topology.conf provided |
| Storage throughput per node? | Checkpoint write time C | Measured numbers for parallel FS and local NVMe |
| Data egress and location? | Cost and compliance | Stated egress pricing; known jurisdiction of the site |
| Access and isolation? | Security of weights and data | Single-tenant fabric, your own identity provider, audit logs |
Compare these answers across providers instead of comparing list prices per GPU-hour. A cheaper rate on a cluster that loses ten percent of its time to slow replacement and a bad link is the more expensive cluster. If you are spreading risk across several suppliers, read running across multiple GPU providers and provider failover; for how a larger neocloud packages the same product, see CoreWeave in depth.
Failure modes in production
- One slow node. Synchronous training runs at the speed of the slowest rank. Detect with per-rank step-time logging; evict and re-run acceptance on the replacement.
- Silent data corruption. A GPU that computes wrong values without erroring shows up as loss spikes or NaNs. Keep periodic checksums or duplicate-compute probes, and retain checkpoints far enough back to roll past a corrupted window.
- Fabric flaps. Link errors cause NCCL timeouts that look like hangs. Watch IB error counters and set NCCL timeouts so a hang becomes a fast restart, not a stalled day.
- Storage contention at checkpoint time. Thousands of ranks writing at once saturate the file system. Stagger or shard writes and checkpoint asynchronously.
- Capacity that slips. New sites depend on power and construction schedules. Contract dates are plans; keep a fallback for the first months of any new build.
What to do next
- Write your acceptance test as a script now, before you sign: topology, DCGM level 3, ibstat, nccl-tests full and pairwise, and a two-hour run of your own model.
- Ask every provider the seven questions in the table and score the answers.
- Measure your checkpoint write time today and compute Young's interval for a plausible MTBF.
- Add per-rank step-time logging and an automatic slow-rank alert to your training loop.
- Decide in advance which failures you will roll back past, and how many checkpoints to keep.