Vast.ai is a marketplace for GPU rental. Instead of one provider running uniform data centres, thousands of independent hosts, ranging from individuals with a few consumer cards to certified data-centre operators, list machines and set their own prices. You search the listings, rent a machine by the hour, and get a Docker container on it with the GPUs you asked for. Prices for the same GPU model are often well below hyperscaler on-demand rates, and the catalogue includes consumer cards like the RTX 4090 that large clouds rarely offer.
The trade is variability. Machines differ in network speed, disk, CPU, driver version and track record, and the cheapest capacity can be interrupted. This article explains how the marketplace and its billing work, how to choose an offer with the command-line query language, how to run a training job that survives interruption, how to keep data and secrets safe on hardware you do not control, and the failure modes to plan for. Specific commands and field names were checked against the Vast.ai documentation in October 2026; prices are deliberately not quoted because they change daily.
How the marketplace works
Four things are worth understanding before you rent anything.
- Offers. Each listing is an offer: a specific machine (or a slice of its GPUs) with a price per hour, plus attributes such as GPU model and count, GPU memory, CPU cores, RAM, disk size and speed, internet bandwidth, CUDA and driver version, location and a reliability score. An offer creates exactly one instance.
- Instances. Your instance is a container started from an image you choose, for example an official PyTorch image. You reach it over SSH or Jupyter, either through Vast's proxy or directly on host ports when the offer has them.
- Billing meters. There are three. GPU compute is charged while the instance is running. Storage is charged from creation until you destroy the instance, including while it is stopped. Bandwidth is charged per byte at a rate each host sets, for both uploads and downloads. Stopping an instance pauses the first meter only; destroying it is the only way to stop the storage charge.
- Trust tiers. Search defaults to verified machines. A separate Secure Cloud filter restricts results to vetted data-centre partners in facilities certified to ISO 27001 and/or Tier 3/4 standards. Everything else is community hosting.
On-demand, reserved and interruptible
Every offer can be rented in one of three ways, and the choice decides how you must engineer the job.
| Type | Price | Priority | Use for |
|---|---|---|---|
| On-demand | listed price | high: cannot be displaced by bids | interactive work, inference endpoints, short critical runs |
| Reserved | prepaid, discounts up to about 50% | high | steady long-running work on a machine you have vetted |
| Interruptible | your bid, often half of on-demand or less | low: highest bid runs, on-demand beats any bid | checkpointed training, batch inference, sweeps |
Interruptible rentals are an auction per machine. You set --bid_price in dollars per hour for the machine. While yours is the highest bid and nobody rents the machine on demand, your container runs. If someone bids higher or rents it on demand, your instance is stopped and its processes killed. Your disk is preserved, and it resumes when you again hold top priority, which may be minutes later or never. The docs describe no advance warning period, so design for an abrupt kill, not a graceful shutdown. The hyperscaler equivalent, with its notice signals, is covered in GPU spot instances in depth.
Choosing an offer
The CLI is a pip package, vastai. After vastai set api-key <key>, vastai search offers takes a query string of field comparisons. The default query is external=false rentable=true verified=true, and -o sorts by comma-separated fields, with a trailing - for descending.
# one RTX 4090, verified, good track record, fast network, direct SSH port
vastai search offers 'gpu_name=RTX_4090 num_gpus=1 reliability>0.98 inet_down>500 \
disk_space>=100 cuda_vers>=12.4 direct_port_count>=1' -o 'dlperf_usd-'
# eight 80 GB GPUs on one box reporting NVLink, data-centre hosts, interruptible pricing
vastai search offers 'num_gpus=8 gpu_ram>=80 bw_nvlink>0 datacenter=true' --type bid -o 'dph'The fields that matter most for ML work:
dlperfanddlperf_usd: Vast's own estimate of deep-learning throughput, and that estimate per dollar. Useful for ranking, but it is a model, not your workload; benchmark the top few.reliability: a score from the machine's history of uptime and health. For runs longer than a day, filter high, because a host that drops offline takes your disk with it.inet_down,inet_down_costanddisk_bw: a host with a slow uplink or disk can spend an hour pulling a dataset and image while the GPU meter runs.cuda_vers: the highest CUDA version the host driver supports. Your image's CUDA runtime must not exceed it.pcie_bwandbw_nvlink: host-to-GPU and GPU-to-GPU bandwidth. Consumer multi-GPU boxes usually lack NVLink, so data-parallel gradient exchange runs over PCIe.min_bid: the lowest bid currently accepted for interruptible use.
Launching a job that survives interruption
Rule one for any rented GPU you do not control: the instance is disposable. Code comes from a git repository or an image, data from object storage, and checkpoints go back to object storage on a timer. The launcher below searches, picks the cheapest qualifying offer, and creates an interruptible instance whose start-up command fetches the code and resumes.
import json, subprocess
def vast(*args):
out = subprocess.run(["vastai", *args, "--raw"], check=True, capture_output=True, text=True)
return json.loads(out.stdout)
QUERY = ("gpu_name=RTX_4090 num_gpus=1 reliability>0.98 inet_down>500 "
"disk_space>=100 cuda_vers>=12.4 direct_port_count>=1")
offers = vast("search", "offers", QUERY, "--type", "bid", "-o", "dlperf_usd-")
offer = offers[0]
print(json.dumps(offer, indent=1)[:800]) # inspect field names on your CLI version
ONSTART = " && ".join([
"cd /workspace",
"git clone --depth 1 https://github.com/acme/train.git",
"cd train && pip install -r requirements.txt",
"nohup python train.py --ckpt-uri s3://acme-ckpt/run-42 > /workspace/train.log 2>&1 &",
])
resp = vast("create", "instance", str(offer["id"]),
"--image", "pytorch/pytorch:2.4.0-cuda12.4-cudnn9-runtime",
"--disk", "100", "--ssh", "--direct",
"--bid_price", "0.30", # $/hour for the machine; your ceiling
"--label", "run-42",
"--env", "-e AWS_ACCESS_KEY_ID=... -e AWS_SECRET_ACCESS_KEY=...", # short-lived, prefix-scoped
"--onstart-cmd", ONSTART)
print(resp) # contains the new instance idThe training script owns resumption. It must find the newest complete checkpoint at start, and it must write checkpoints so that a kill mid-upload never leaves a half-written file looking like the latest one.
def save(state, step, uri):
local = f"/workspace/ckpt_{step}.pt"
torch.save(state, local)
upload(local, f"{uri}/ckpt_{step}.pt")
upload_text(f"{uri}/LATEST", str(step)) # pointer written last = commit
def resume(model, opt, uri):
step = read_text_or_none(f"{uri}/LATEST")
if step is None:
return 0
state = torch.load(download(f"{uri}/ckpt_{step}.pt"), map_location="cuda")
model.load_state_dict(state["model"]); opt.load_state_dict(state["opt"])
torch.set_rng_state(state["rng"]) # plus dataloader position
return state["step"]
# in the loop: save on a wall-clock timer, not every N steps
if time.time() - last_save > 20 * 60:
save({...}, step, args.ckpt_uri); last_save = time.time()Save on wall-clock time, because the work lost to a kill is bounded by the interval in minutes, whatever the step rate on this particular machine. Store optimizer state, RNG state and data-loader position, or the resumed run will not be the same run. Training checkpointing in depth covers sharded and asynchronous checkpoints for larger models.
Worked example: cost per useful hour
Compare offers on cost per useful hour, not price per hour. Suppose, purely as labelled assumptions, a fine-tuning job needs 40 GPU-hours, an on-demand offer costs $0.50 per hour and interruptible capacity on a similar machine clears at $0.25.
| Item | On-demand | Interruptible |
|---|---|---|
| Compute, 40 h | $20.00 | $10.00 |
| Interruptions expected | 0 | 4 |
| Lost work per interruption (half a 20-min interval + 15-min restart) | - | 25 min |
| Extra GPU hours | 0 | 1.7 h = $0.42 |
| Image and dataset pulls (2 GB each relaunch, host rate $0.01/GB, assumed) | $0.02 | $0.10 |
| Disk, 100 GB, for the wall-clock duration | small | larger: runs while paused |
| Total, approximately | $20 | $11 |
Interruptible wins comfortably here, as it usually does for checkpointed work. Three things can flip the result: a long restart because the dataset is large and the host's uplink slow, a short checkpoint interval being impossible because checkpoints are huge, and a deadline, since a paused instance may sit idle for hours. Wall-clock time is the hidden cost. For the longer-term buy-versus-rent question, see bare metal versus cloud GPU cost analysis.
Security on hardware you do not control
On a community host, the machine's owner has physical access to the hardware, the disk and, in principle, the memory of a running container. Treat it the way you would treat any third-party machine.
- Do not put regulated or personal data, unreleased proprietary weights or long-lived credentials on community hosts. Use the Secure Cloud filter, or a different provider, for sensitive work.
- Pass short-lived, narrowly scoped credentials, for example a key that can only write to one checkpoint prefix, and rotate it after the run. Environment variables are visible to anyone with root on the box.
- Encrypt checkpoints client-side if their contents matter, and verify downloaded artefacts with checksums.
- Pin images by digest so the code you run is the code you reviewed.
Multi-GPU and multi-node
Single-machine jobs are the sweet spot. Within one machine, data and tensor parallelism work as usual, limited by the interconnect: check bw_nvlink and pcie_bw before assuming that eight consumer GPUs scale like eight NVLinked data-centre GPUs. Gradient all-reduce over PCIe can dominate step time for large models, as NCCL all-reduce in depth explains.
Multi-node training across separate marketplace hosts is rarely worth it. Hosts sit in different buildings connected by ordinary internet links, so collective operations crawl. If a job needs more GPUs than one machine has, prefer a provider that sells clusters with a high-speed fabric, or restructure the work into independent jobs such as hyperparameter sweeps and data-parallel batch inference shards.
Failure modes
| Symptom | Cause | What to do |
|---|---|---|
| Restart hangs in SCHEDULING | another renter took the GPU while yours was stopped | relaunch elsewhere from the last uploaded checkpoint; do not wait on a single host |
| Bill keeps growing after stopping | storage is billed until destroy | destroy finished instances; script a sweep of idle ones |
| First hour spent downloading | slow host uplink or a huge image | filter on inet_down; use slim images; stage data in a nearby region |
| CUDA error at start-up | image CUDA runtime newer than host driver | filter on cuda_vers; match the image to it |
| Throughput far below the benchmark | thermal or power limits, weak CPU, slow disk feeding the GPU | run a 5-minute benchmark before committing; check nvidia-smi clocks |
| Training restarts from step 0 | checkpoint only on local disk, or resume code untested | upload checkpoints off-host; test kill-and-resume before the real run |
| Run silently corrupted after resume | partial checkpoint treated as complete | write the LATEST pointer only after the upload succeeds |
Trade-offs
Choose Vast.ai when cost dominates, the job fits on one machine, the data is not sensitive and the code tolerates interruption. Choose a hyperscaler or a specialised GPU cloud when you need uniform hardware, multi-node fabrics, compliance commitments or integration with managed services. Many teams use both: marketplace capacity for experiments and sweeps, and a stable provider for production inference and large runs.
What to do next
- Install the CLI, set an API key, and run a search with a reliability, bandwidth and CUDA filter for the GPU you need.
- Rent one on-demand instance for an hour and benchmark your actual training step against the offer's
dlperf. - Move code into git or an image and data and checkpoints into object storage you control.
- Implement time-based checkpointing with a commit pointer, then kill the instance mid-run and confirm the resume is exact.
- Switch to interruptible with a bid ceiling, and script relaunch on another offer when a restart sticks in SCHEDULING.
- Add a scheduled job that lists instances and destroys any that are stopped and finished.