Lambda is a GPU cloud built for machine learning teams. It sells three things: on-demand instances with one to eight GPUs, rented by the minute; 1-Click Clusters of 16 to 512 GPUs on an InfiniBand fabric, reserved by the week; and private cloud capacity on longer contracts. Its appeal is that the product surface is small. There is no catalogue of a hundred services to learn, a GPU instance boots with the drivers and frameworks already installed, and the whole control plane is a short REST API.

A small surface does not mean no sharp edges. Capacity comes and goes by region, storage is regional and can only be attached at launch, there is no way to pause an instance, and the one billing mistake everyone makes once is easy to make twice. This article explains how the pieces fit, shows a working controller that finds capacity and launches within the API's rate limits, prices a worked fine-tuning run, and covers the failure modes. Where Lambda sits among providers is covered in the CoreWeave and neocloud article.

Three products, one model

ProductShapeHow you get itBillingGood for
On-demand instanceone machine with 1 to 8 GPUs; type names from the APIconsole or API, minutes to bootper minute, from passing health checks to terminationfine-tuning, experiments, inference tests
1-Click Cluster16 to 512 H100 or B200 SXM GPUs plus 3 CPU head nodesreservation for a number of weeksper GPU-hour, invoiced weeklymulti-node pre-training and large fine-tunes
Private clouddedicated capacitycontractper contractsteady, long-running demand

The on-demand instance is the unit most people start with. Each type reports its GPU count, vCPUs, memory, local storage, CPU architecture and hourly price through the API. Some types report arm64 and need arm64 Python wheels and containers; check the architecture field before you launch a script built for x86.

The instance lifecycle and the missing off switch

An instance moves through a small set of states: booting, active, unhealthy, terminating, terminated and preempted. You can launch, restart and terminate. You cannot stop or suspend. Lambda's documentation is explicit that shutting the machine down from inside, for example with sudo shutdown -h now, does not end billing: the instance goes into an alert state and keeps charging until you terminate it through the console or API.

That gives one rule that shapes everything else: terminate is the only off switch, and terminate deletes the local disk. Anything you want to keep must be on a filesystem or copied off before termination. The default image is the latest Lambda Stack, which bundles the NVIDIA driver, CUDA and the common deep learning frameworks; you can pick another image by ID or family at launch, and pass a cloud-init user_data script of up to 1 MB to configure the machine on first boot.

Operating on Lambda: your controller owns the lifecycle, the region owns the dataYour controllerscript, CI job or laptopBearer keyCloud APIGET /instance-typesPOST launch / terminateGET /instances/{id}launchRegion, e.g. us-east-1capacity varies by hourOn-demand instance1 to 8 GPUs, Lambda StackFilesystem (regional)attach at launch onlyLocal NVMefast, lost on terminatessh, run job1-Click Cluster: 16 to 512 H100 or B200 GPUs, rail-optimized 400 Gb/s InfiniBand3 CPU head nodes with public IPs as jump hosts; compute nodes on a private network; reserved in weeks
The controller drives the API; instances and filesystems live in one region; local NVMe disappears at termination. A 1-Click Cluster is the multi-node shape of the same model.

The Cloud API and a controller that owns the lifecycle

The Cloud API lives at cloud.lambda.ai/api/v1, authenticates with an API key sent as a Bearer token, and wraps every successful response in a data field. Two rate limits matter: about one request per second in general, and one launch request every 12 seconds. A 429 with the code global/rate-limited means you exceeded one of them.

The pattern that works is a controller that owns the whole lifecycle. It discovers which instance types have capacity, launches one in the region where your data lives, waits for it to become active, runs the job and terminates the instance in a finally block, so a crashed job does not leave a machine billing overnight. Never hardcode a type name or price; read both from /instance-types, which also lists, per type, the regions with capacity right now.

# lambda_ctl.py -- find capacity, launch, wait, run, always terminate
import os, time, subprocess, requests

API = "https://cloud.lambda.ai/api/v1"
S = requests.Session()
S.headers["Authorization"] = "Bearer " + os.environ["LAMBDA_API_KEY"]

def get(path):
    r = S.get(API + path, timeout=30)
    r.raise_for_status()
    return r.json()["data"]

def candidates(gpus, regions_ok):
    """Instance types with the GPU count we need and capacity in an allowed region."""
    out = []
    for name, item in get("/instance-types").items():
        spec = item["instance_type"]["specs"]
        free = [r["name"] for r in item["regions_with_capacity_available"]]
        usable = [r for r in free if r in regions_ok]
        if spec["gpus"] == gpus and usable:
            out.append((item["instance_type"]["price_cents_per_hour"], name, usable[0]))
    return sorted(out)

def launch(type_name, region, fs_name):
    body = {"region_name": region, "instance_type_name": type_name,
            "ssh_key_names": ["ops"], "file_system_names": [fs_name],
            "name": "ft-run"}
    r = S.post(API + "/instance-operations/launch", json=body, timeout=60)
    r.raise_for_status()
    return r.json()["data"]["instance_ids"][0]

def wait_active(iid):
    while True:
        inst = get("/instances/" + iid)
        if inst["status"] == "active":
            return inst["ip"]
        if inst["status"] in ("unhealthy", "terminated", "preempted"):
            raise RuntimeError(inst["status"])
        time.sleep(15)

def terminate(iid):
    S.post(API + "/instance-operations/terminate",
           json={"instance_ids": [iid]}, timeout=60).raise_for_status()

REGION = "us-east-1"          # must be the filesystem's region
while not (c := candidates(8, {REGION})):
    time.sleep(60)            # well inside 1 request per second
iid = launch(c[0][1], REGION, "datasets-east")
try:
    ip = wait_active(iid)
    subprocess.run(["ssh", "ubuntu@" + ip, "bash /lambda/nfs/datasets-east/run.sh"], check=True)
finally:
    terminate(iid)            # the only off switch

The sketch omits retries and logging, and it assumes the filesystem is mounted under its own name; check the mount path on your first instance. In production, back off on 429, alert when the wait exceeds your patience, and record each instance ID so a second process can find and terminate orphans.

Two security points belong in the same controller. Inbound access is governed by firewall rules: the API manages firewall rulesets, including a global one, and a launch request can name the rulesets to apply. Limit SSH to your own address ranges, and remember that the head nodes of a 1-Click Cluster have public addresses too. And treat the API key like a root credential, because it can launch and terminate everything in the account: keep it in a secret store, use separate keys for separate automations, and rotate them.

Storage: local NVMe and regional filesystems

Lambda has two kinds of storage with opposite properties. Each instance has local NVMe, sized per type and reported as storage_gib. It is fast and it disappears at termination. Filesystems are regional network file stores, mounted under /lambda/nfs/ or at a mount point you choose at launch. They persist independently of instances, are billed per GiB used per month while they exist, and carry no ingress or egress charge.

Filesystems come with three rules that catch people. They must be in the same region as the instance. They can only be attached when the instance is launched, never to a running one. And they cannot be moved between regions. Together these mean that your data's region decides where you can run: if capacity appears in another region, you cannot simply use it with your existing filesystem. Capacity limits differ by region too; at the time of writing, filesystems in us-south-1 are limited to 10 TB. Accounts can have up to 24 filesystems.

A reliable layout follows from the properties. Keep datasets and checkpoints on the filesystem. At job start, copy the shards you will read repeatedly to local NVMe and train from there. Write checkpoints to the filesystem, and test that the job resumes from them, as described in LLM training checkpointing. Lambda does not publish filesystem throughput figures in the pages cited here, so measure read and write bandwidth with fio before you size a staging step or a checkpoint interval.

Worked example: pricing a fine-tuning run

Plan a fine-tuning run on one 8-GPU instance: 48 hours of training on a 2 TB dataset already on a filesystem in the region. Let P be the instance price per hour that the API reports. The run itself costs 48P. Staging the dataset to local NVMe at a measured, say, 1 GB/s takes about 35 minutes, which costs about 0.6P and is repaid if the job reads the data more than once. Checkpointing every hour bounds lost work to an hour if the instance becomes unhealthy.

Now the failure. The job finishes at 6 p.m. on a Friday, the engineer shuts the machine down from a shell and goes home, and nobody looks again until 9 a.m. on Monday. Because shutdown does not stop billing, those 63 hours cost 63P, more than the training run. The controller above prevents this by terminating in a finally block. As a second line of defence, run a scheduled job that lists /instances, compares each one with your record of running jobs, and terminates anything nobody owns. The broader cost comparison of renting against owning is in bare metal against cloud GPU.

1-Click Clusters

A 1-Click Cluster is the multi-node shape. Lambda's documentation describes 16 to 512 H100 or B200 SXM GPUs connected by a non-blocking NVIDIA Quantum-2 InfiniBand fabric at 400 Gb/s in a rail-optimized topology, which allows GPUDirect RDMA at up to 3,200 Gb/s per node, eight rails of 400 Gb/s. Every node also has two 100 Gb/s Ethernet ports for IP traffic. Each compute node has 24 TB of usable local NVMe, and filesystems are created and attached for you. Three CPU head nodes have public IP addresses; they serve as jump hosts and as the place to run cluster administration and job scheduling. Compute nodes sit on an isolated private network, reachable over SSH through a head node.

# ~/.ssh/config on your workstation
Host lc-head
    HostName <head-node-public-ip>
    User ubuntu
Host lc-node-*
    ProxyJump lc-head
    User ubuntu

Treat the first hours of a reservation as acceptance testing, because the reservation is billed whether or not the fabric is healthy. Check that every compute node reports all eight InfiniBand ports active at 400 Gb/s, run a GPU-memory-to-GPU-memory bandwidth test on each rail, then a multi-node NCCL all-reduce, and keep the results. Place ranks so that local rank matches GPU index, which keeps same-rank traffic on its own rail, as explained in rail-aligned topology. Report any slow node immediately rather than training around it.

Failure modes

  • No capacity. The type you want shows no regions. The controller waits and polls; the fix is flexibility in GPU type or region, which only helps if your data can follow.
  • Wrong-region filesystem. Capacity appears in a region without your data, and launch with that filesystem fails. Keep a copy of critical data in a second region if you need to follow capacity.
  • Billing after shutdown. Covered above: terminate, never power off.
  • Lost local data. Checkpoints written only to local NVMe vanish at termination or when an instance turns unhealthy. Write them to the filesystem.
  • Rate limiting. A tight loop of launch attempts returns 429. Respect the 12-second launch interval and back off.
  • Account limits. New accounts have a cap on how many instances they can run, raised as invoices are paid. Plan this before a deadline, not during it.

Trade-offs

Against a hyperscaler, Lambda trades breadth for simplicity: no managed databases or queues next to your GPUs, but fewer moving parts and drivers that already work. Against reserved-capacity neoclouds, its on-demand tier lets you rent by the minute with no commitment, at the cost of capacity that may not be there when you need it. The 1-Click Cluster sits between: a real InfiniBand cluster without a long contract, paid by the week. Teams that need guaranteed capacity at scale move to reservations or private cloud; teams that need to fail over between providers should read multi-provider GPU strategies.

What to do next

  1. Create an API key, store it in a secret manager, and list /instance-types to see current types, prices and regions with capacity.
  2. Choose a home region, create a filesystem there, and upload your dataset.
  3. Adapt the controller so every launch ends in terminate, and add a scheduled orphan sweep.
  4. Measure filesystem and local NVMe bandwidth with fio on your first instance.
  5. Add checkpoint-and-resume to your training script and test it by terminating mid-run.
  6. For multi-node work, plan a 1-Click Cluster reservation and an acceptance test for its first hours.
Key takeaway: Lambda keeps the GPU cloud small: instances by the minute, clusters by the week, a short API. Run it through a controller that discovers capacity, launches where your data lives and always terminates; keep data on regional filesystems and scratch on local NVMe; and acceptance-test a cluster before you train on it.