Modal is a serverless platform for Python workloads that need GPUs. Instead of provisioning nodes, writing Dockerfiles and running a scheduler, you decorate ordinary Python functions with the image they need, the GPU type and the scaling policy, and Modal builds the container, places it on a machine, scales it from zero to many copies with demand and back to zero when idle. Billing is per second of container time, so an idle endpoint can cost nothing.
That model is excellent for bursty inference, batch jobs and experiments, and it brings its own engineering problems: cold starts, preemption, and storage that behaves differently from a local disk. This article explains the programming model, what happens on a cold start and how to shrink it, how to run batch fan-out and training jobs safely, and when a reserved cluster is the better choice. API names are taken from Modal's documentation as of October 2026; check the docs for changes.
The programming model
Four objects carry most programs. An modal.App groups functions that deploy together. An modal.Image describes the container, built layer by layer from Python method calls and cached per layer, so a change to one layer rebuilds it and every layer after it. A function decorated with @app.function(...) becomes a remote Function with its own hardware and scaling settings. A class decorated with @app.cls(...) adds lifecycle hooks: @modal.enter() runs once per container before any input, @modal.method() handles inputs, and @modal.exit() runs on shutdown with a 30-second grace period.
import modal
app = modal.App("embedder")
image = (
modal.Image.debian_slim(python_version="3.12")
.uv_pip_install("torch==2.8.0", "transformers==4.56.0") # pin versions
)
weights = modal.Volume.from_name("embedder-weights", create_if_missing=True)
@app.cls(
image=image,
gpu=["L40S", "A100-40GB"], # preference order, falls back if L40S is scarce
volumes={"/weights": weights},
scaledown_window=300, # keep idle containers 5 minutes
enable_memory_snapshot=True,
)
class Embedder:
@modal.enter(snap=True)
def load_to_cpu(self): # captured in the CPU memory snapshot
import torch
from transformers import AutoModel, AutoTokenizer
self.tok = AutoTokenizer.from_pretrained("/weights/model")
self.model = AutoModel.from_pretrained("/weights/model", torch_dtype=torch.bfloat16)
@modal.enter(snap=False)
def to_gpu(self): # runs after every restore
self.model = self.model.to("cuda").eval()
@modal.method()
def embed(self, texts: list[str]) -> list[list[float]]:
import torch
batch = self.tok(texts, padding=True, truncation=True, return_tensors="pt").to("cuda")
with torch.no_grad():
out = self.model(**batch).last_hidden_state.mean(dim=1)
return out.float().cpu().tolist()
@app.local_entrypoint()
def main():
print(len(Embedder().embed.remote(["hello", "world"])[0]))modal run app.py runs the entrypoint against an ephemeral app; modal deploy app.py creates a persistent deployment that other code, or an HTTP endpoint built with @modal.fastapi_endpoint() or @modal.asgi_app(), can call. The model path assumes you have already copied weights into the Volume; downloading from a hub inside @modal.enter on every cold start is the most common self-inflicted latency.
Choosing GPUs
The gpu argument takes a type string: "T4", "L4", "A10", "L40S", "A100-40GB", "A100-80GB", "RTX-PRO-6000", "H100", "H200", "B200" and "B300" at the time of writing. Append :n for several GPUs in one container: up to 8 on most types, 4 on A10. A list is a preference order, which trades a little predictability for much better availability.
Two modifiers matter for benchmarking. Modal may run a "H100" request on an H200 at the same price; "H100!" pins it to H100 so measurements stay comparable. Conversely "B200+" opts in to running on either B200 or B300, billed as B200, for a larger capacity pool; use it only if your code runs on both, because B300 requires CUDA 13.1 or later. Pick the smallest GPU whose memory holds the weights, KV cache or activations with headroom; on a per-second bill the expensive GPU only wins if it finishes proportionally faster.
Cold starts and how to shrink them
A cold start is everything between an input arriving and a container being ready to process it: scheduling a machine, fetching the image, importing Python packages, loading weights from storage into host and then GPU memory, and CUDA initialisation including any compilation or graph capture. For a multi-gigabyte model, weight loading usually dominates. Modal gives you knobs on both sides of the trade between latency and idle cost:
| Knob | Effect | Cost |
|---|---|---|
min_containers=k | k containers stay warm even with no traffic | you pay for idle GPUs around the clock |
buffer_containers=k | keeps k extra idle containers while the Function is active | pays for headroom only during busy periods |
scaledown_window=s | how long an idle container lives, 2 s to 20 min (default 60 s) | longer windows absorb gaps between bursts |
enable_memory_snapshot=True | restores a snapshot of CPU memory taken after @modal.enter(snap=True) | imports and CPU-side setup are skipped |
experimental_options={"enable_gpu_snapshot": True} | also snapshots GPU memory (alpha) | see caveats below |
GPU memory snapshots are an alpha feature and Modal's documentation lists their limits plainly: they do not work with most multi-GPU code, can fail when torch.compile has run before the snapshot, and generally do not help when most of the start-up time is spent loading weights. The dependable pattern is the one in the example above: load weights to CPU in the snap=True hook, move them to the GPU in the snap=False hook, and keep the weights in a Volume in the same region rather than downloading them.
Throughput within a warm container is the other half. @modal.concurrent(max_inputs=...) lets one container accept several inputs at once, which suits async servers such as an inference engine with its own batching; target_inputs tells the autoscaler the level to aim for before adding containers. Without it each container handles one input at a time and a GPU spends most of its life waiting.
Batch fan-out and job queues
For offline work, .map() fans an iterable of inputs across as many containers as the autoscaler will start and yields results in order; .starmap() does the same for argument tuples. .spawn() submits one call and returns a FunctionCall handle whose .get(timeout=...) collects the result later, which is how you build a job queue without holding a connection open. max_containers caps parallelism, for example to protect a downstream database, and Modal enforces a hard limit of 4,000 concurrent containers per Function.
@app.function(gpu="L4", timeout=60 * 30, retries=3, max_containers=200)
def score_shard(shard_uri: str) -> dict:
... # download one shard, run the model, write results, return counts
@app.local_entrypoint()
def main():
shards = [f"s3://bucket/eval/part-{i:05d}.parquet" for i in range(5000)]
totals = {}
for counts in score_shard.map(shards, return_exceptions=True):
if isinstance(counts, Exception):
print("failed shard:", counts)
continue
for k, v in counts.items():
totals[k] = totals.get(k, 0) + v
print(totals)Functions time out after 300 seconds by default; timeout raises that to as much as 24 hours per attempt. Design each input to be idempotent: retries, preemption and duplicate delivery all mean a shard can run more than once, so write results under a deterministic key rather than appending. If you only need generic GPU time for many short jobs, compare with Replicate and Together AI.
Training under preemption, and multi-node clusters
Training on a serverless platform works if you assume the container can disappear. All Modal Functions are preemptible by default; when one is preempted, Modal restarts it on the same input. The nonpreemptible=True option costs three times as much and is not supported for GPU Functions, so GPU training must checkpoint. Write checkpoints to a Volume, call commit() so other containers can see them, and make the function resume from the latest checkpoint when it starts. The training checkpointing article covers what to save and how often.
ckpts = modal.Volume.from_name("run-42-ckpts", create_if_missing=True)
train_image = (
modal.Image.from_registry("nvidia/cuda:12.9.1-devel-ubuntu22.04", add_python="3.12")
.apt_install("libibverbs1") # RDMA user-space library
.uv_pip_install("torch==2.8.0")
)
@app.function(image=train_image, gpu="H100:8", volumes={"/ckpt": ckpts},
timeout=24 * 60 * 60,
retries=modal.Retries(initial_delay=0.0, max_retries=10))
@modal.clustered(size=4, rdma=True)
def train():
cluster = modal.Cluster.from_context()
rank, ips = cluster.container_rank(), cluster.container_ips()
# ips are ordered by rank; ips[0] is the rendezvous host for torch.distributed.
# NCCL environment variables for the RDMA fabric are set by Modal.
state = load_latest("/ckpt") # your code: None on the first attempt
for step in range(state.step if state else 0, TOTAL_STEPS):
train_step()
if step % 500 == 0 and rank == 0:
save_checkpoint("/ckpt", step) # write to a new file, then commit
ckpts.commit()Multi-node jobs use @modal.clustered(size=...), which gang-schedules that many containers and gives them private networking. Each container must use all GPUs of its node ("H100:8", not "H100:4"), and clusters go up to 32 nodes and 256 GPUs. With rdma=True the containers get Modal's RDMA scale-out network, which Modal quotes at 3,200 Gb/s for most GPU types and 6,400 Gb/s for B300 clusters. RDMA also needs the verbs libraries in the image, which is why the example builds on a CUDA base image with libibverbs1 installed; without them NCCL quietly falls back to TCP. Failures and preemptions apply to the whole cluster, which is torn down and retried as a unit, and only rank 0's return value reaches the caller. Keep tensor parallelism inside each 8-GPU container and send only data-parallel or pipeline traffic across the RDMA network.
Volumes: commit, reload and limits
Volumes are a distributed file system tuned for write-once, read-many data such as model weights. Changes become visible to other containers only after a commit() (also done automatically every few seconds and on shutdown), and a running container sees others' changes only after reload(). Concurrent writes to the same file are last-write-wins. The original Volumes have a hard limit of 500,000 files and work best below 50,000; the newer v2 Volumes remove the file-count limit but cap files at 1 TiB and a directory at 262,144 entries. Store checkpoints as a small number of large files, and never let two ranks write the same path.
Failure modes
- Downloading weights on every cold start. Seconds to minutes per container, multiplied by every scale-up. Put them in a Volume.
- Unpinned images. A floating
torchversion rebuilds into something different months later. Pin every package. - Benchmarks that drift. An
"H100"request may land on an H200; use"H100!"when measuring. - Non-idempotent inputs. Retries and preemption replay inputs. Appending to a shared file double-counts.
- Checkpoint invisible after restart. Writing without
commit()or reading withoutreload()resumes from an old step. - Snapshots and Volumes drifting apart. Updating weights in a Volume does not refresh existing memory snapshots, and deleting Volume files a snapshot used during restore makes restores fail. Treat those files as immutable: write new weights under a new path and point the code at it.
- Warm pools left on.
min_containersset for a demo and forgotten bills idle GPUs indefinitely. - One input per GPU container. Serving without
@modal.concurrentor batching leaves the GPU mostly idle.
When serverless GPUs are the right choice
Serverless wins when utilisation is low or spiky: a GPU you would otherwise rent by the hour sits idle most of the time, and per-second billing with scale-to-zero removes that waste. It also removes cluster operations: no node images, schedulers or drivers to maintain. It loses when utilisation is high and steady, because reserved capacity is cheaper per GPU-hour, and when a workload needs control Modal does not expose, such as custom kernel drivers, specific fabric topologies or very long uninterrupted runs. A common split is to serve and run batch evaluation on Modal while large training runs live on reserved clusters, with the same containers in both places.
Make the comparison with numbers from your own quotes rather than intuition. Let r be the reserved price per GPU-hour, s the serverless price per GPU-hour of container time, and u the fraction of reserved hours your workload would actually keep busy. Serverless is cheaper while u < r / s, after adding the container time spent on cold starts and warm pools to the serverless side. A worked example with illustrative ratios: if the serverless rate is twice the reserved rate, the break-even is 50% utilisation. An inference endpoint that is busy 6 hours a day (25%) with a 20-minute warm-pool tail after each burst still sits well below it; a training team that keeps eight GPUs busy 20 hours a day (83%) sits well above it. Re-run the sum whenever traffic, prices or the GPU type change, and include engineering time: operating your own cluster is never free.
What to do next
- Install the client, run the embedder example with
modal runand time a cold and a warm call. - Move model weights into a Volume and measure the cold start again.
- Enable memory snapshots with the CPU-then-GPU hook split and compare a third time.
- Set
scaledown_window,buffer_containersand@modal.concurrentfrom your measured traffic, not defaults. - For batch jobs, make every input idempotent and use
max_containersto protect downstream systems. - For training, implement resume-from-checkpoint and test it by killing a run mid-step before trusting it.