Replicate is a hosted platform for running machine-learning models behind an HTTP API. You send JSON inputs to a model, Replicate schedules the work onto a GPU or CPU instance running that model's container, and you get back outputs: text, numbers, or URLs to generated files. Its catalogue is large and skewed towards media models (image, video, audio, upscaling) as well as language models, and anyone can publish a model into it by packaging the model with Replicate's open-source tool, Cog. Cloudflare announced in November 2025 that it was acquiring Replicate; Replicate has said its API is unchanged, and everything below describes that API.
The interface looks trivially simple, which is why teams get surprised in production. Every call is really an asynchronous job with a lifecycle; whether the job waits seconds or minutes for a GPU depends on which of three billing models the model falls under; outputs vanish after an hour; and webhooks arrive more than once. This article covers what a prediction is, how Cog packages a model, where cold starts come from, how to receive results safely, and how to choose between public, official and deployed models.
Models, versions and predictions
Three nouns carry the whole system. A model is a named thing, owner/name, such as a public image generator or your company's private fine-tune. A version is an immutable build of that model: a specific container image with specific weights and a specific input schema, identified by a long hash. A prediction is one execution of one version against one set of inputs. Predictions are the unit you create, poll, cancel and pay for.
Pinning matters. Calling a community model by owner/name:version gives you the exact build you tested; anything less than a pinned version can change under you when the owner publishes, including its input names, defaults or output format overnight. Official models are the exception: they are maintained by Replicate with a stable interface and are called through the model-scoped endpoint without a version.
There are three creation endpoints, one per way of addressing compute:
| Endpoint | Addresses | Typical use |
|---|---|---|
POST /v1/predictions | a specific version (in the body) | community models, pinned |
POST /v1/models/{owner}/{name}/predictions | an official model | stable, maintained models |
POST /v1/deployments/{owner}/{name}/predictions | your deployment | production traffic on dedicated instances |
A prediction moves through starting, processing and one terminal state: succeeded, failed or canceled. starting covers both queueing and any cold boot; processing means the model code is running. Running predictions can be stopped with POST /v1/predictions/{id}/cancel, which you should do whenever the user who asked for the result has gone away, because time-billed models keep billing.
Packaging a model with Cog
Cog is the packaging layer. It reads a cog.yaml describing the environment and a Python class describing how to load and run the model, then builds a container that exposes an HTTP prediction server with an input schema derived from your type hints. Current Cog releases name the class BaseRunner with a run() method and point at it with a run: key; the older BasePredictor, predict() and predict: names still work for existing models but are deprecated and make Cog print warnings. You will meet both in the wild.
# cog.yaml
build:
gpu: true
python_version: "3.12"
python_requirements: requirements.txt # python_packages is deprecated
system_packages:
- ffmpeg
run: "run.py:Runner" # older models use predict: "predict.py:Predictor"# run.py
from cog import BaseRunner, Input, Path
import torch
class Runner(BaseRunner):
def setup(self) -> None:
# Runs once per instance boot. Everything slow belongs here.
self.pipe = load_pipeline("/src/weights", dtype=torch.float16).to("cuda")
def run(
self,
prompt: str = Input(description="What to generate"),
steps: int = Input(description="Denoising steps", default=28, ge=1, le=80),
seed: int = Input(description="Random seed; -1 for random", default=-1),
) -> Path:
g = None if seed < 0 else torch.Generator("cuda").manual_seed(seed)
image = self.pipe(prompt, num_inference_steps=steps, generator=g).images[0]
out = Path("/tmp/out.png")
image.save(out)
return outThree details in that file decide your production behaviour. First, the split between setup() and run() is the split between cold-start cost and per-request cost: anything in setup() is paid once per instance, anything in run() on every call. Second, Input() constraints such as ge, le and choices become the published schema, so bad inputs are rejected before your code runs. Third, returning a Path makes Replicate upload the file and hand callers a URL. For token streaming, Cog supports returning an iterator; for concurrency within one instance, an async def run() plus the @cog.concurrent(max=N) decorator (cog 0.21.0 and later) replaces the deprecated concurrency.max stanza. Only enable concurrency if the model actually batches or overlaps well on one GPU; otherwise requests simply contend for the same memory.
Image and weight size are cold-start time: time a cold cog predict on a clean machine before you publish.
Billing models, deployments and cold starts
Replicate bills in three different ways, and the difference is not just price: it determines whether a cold boot is your problem.
- Public community models run on shared hardware. You pay only for the time the model is actively processing your request; setup and idle time are free. The price of that generosity is that an unpopular model may have no warm instance, so your first call waits for a boot, and you share a queue with other users.
- Official models are priced per unit of input or output (per token, per image, per second of video, depending on the model) rather than per second of hardware. Cost is predictable per request, and Replicate keeps them warm.
- Private models and deployments run on dedicated instances. You pay for all the time an instance is online: setting up, idle, and active. Replicate documents one exception: fast-booting fine-tunes are billed only for active time whether public or private.
A deployment is how you take control of capacity. It binds a model version to a hardware type and a minimum and maximum instance count, and gives you a private endpoint. Setting the minimum above zero removes cold boots for the first few concurrent requests and puts a floor under your bill; setting it to zero lets the deployment scale to nothing when idle, and the first request after a quiet period pays the boot. Changing the deployment's version rolls new instances without changing the URL your callers use, which makes deployments the right abstraction even for public models you depend on.
Calling it: sync, polling and signed webhooks
There are three ways to wait for a result. Synchronous: send Prefer: wait and the API holds the request open, by default up to 60 seconds, returning the finished prediction if it completes in time and the in-progress one otherwise. Good for fast models in interactive paths. Polling: create, then GET the prediction's URL until it is terminal. Simple, but wasteful at scale. Webhooks: pass a webhook URL and a webhook_events_filter drawn from start, output, logs and completed; Replicate POSTs the prediction object to you as it changes. For anything slow, use webhooks with ["completed"] and keep polling as a fallback sweep.
import os, requests
API = "https://api.replicate.com/v1"
H = {"Authorization": f"Bearer {os.environ['REPLICATE_API_TOKEN']}",
"Content-Type": "application/json"}
def submit(job_id: str, prompt: str) -> str:
body = {
"input": {"prompt": prompt, "steps": 28},
"webhook": f"https://api.example.com/hooks/replicate?job={job_id}",
"webhook_events_filter": ["completed"],
}
r = requests.post(f"{API}/deployments/acme/image-gen/predictions",
json=body, headers=H, timeout=30)
r.raise_for_status()
pred = r.json()
save_job(job_id, prediction_id=pred["id"], status=pred["status"]) # persist BEFORE returning
return pred["id"]Webhooks are signed. Each request carries webhook-id, webhook-timestamp and webhook-signature headers. The signed content is the id, the timestamp and the raw body joined with full stops, HMAC-SHA256 with your signing secret, which you fetch once from GET /v1/webhooks/default/secret. The secret starts with whsec_ and the part after that prefix is base64. The signature header can hold several space-separated v1,<base64> entries; accept if any matches.
import base64, hashlib, hmac, time
def verify(headers, raw_body: bytes, secret: str, tolerance_s: int = 300) -> bool:
wid, ts = headers["webhook-id"], headers["webhook-timestamp"]
if abs(time.time() - int(ts)) > tolerance_s: # reject replays
return False
key = base64.b64decode(secret.split("_", 1)[1]) # strip "whsec_"
signed = f"{wid}.{ts}.".encode() + raw_body # raw bytes, not re-serialised JSON
expected = base64.b64encode(hmac.new(key, signed, hashlib.sha256).digest()).decode()
for entry in headers["webhook-signature"].split():
_, sig = entry.split(",", 1)
if hmac.compare_digest(sig, expected):
return True
return False
def on_webhook(request):
if not verify(request.headers, request.body, SECRET):
return 401
if seen_before(request.headers["webhook-id"]): # deliveries can repeat
return 200
pred = request.json()
if pred["status"] == "succeeded":
enqueue_copy_outputs(pred["id"], pred["output"]) # outputs expire; copy now
mark_job(pred["id"], pred["status"], pred.get("error"))
return 200 # answer fast; work async
Worked example: an on-demand image feature
Suppose a product generates marketing images on demand. Users click a button and expect a result within about 15 seconds; traffic is 2,000 images a day, peaking at around 6 concurrent requests in business hours and near zero overnight. Each image takes about 4 seconds of GPU time once the model is warm, and a cold boot of your fine-tuned container takes about 90 seconds (measure yours; these are illustrative).
| Option | Latency behaviour | Cost behaviour |
|---|---|---|
| Public model, pinned version | fine when warm; a 90 s boot after quiet periods breaks the 15 s budget | pay ~2,000 × 4 s of active time per day |
| Deployment, min 0, max 4 | first request each morning waits for a boot | pay active time plus boot and idle time while instances are up |
| Deployment, min 1, max 4 | one warm instance always; bursts above it boot new instances | pay 24 h of one instance, plus extra instances while scaled up |
| Official equivalent model | warm, per-image price | pay per image; no control over the weights |
The arithmetic that decides it: 2,000 × 4 s is about 2.2 GPU-hours of real work per day, while a minimum of one instance is 24 GPU-hours. Keeping an instance warm around the clock therefore costs roughly ten times the useful work. A common answer is a deployment with min_instances raised during business hours and dropped to zero overnight by a scheduled job, plus a webhook-driven UI that shows progress instead of blocking, so the occasional cold start degrades to a slower result rather than a timeout. If the per-image price of an official model that meets your quality bar is below your effective per-image cost, that beats both.
Failure modes
The failures teams actually hit are mostly about state, not GPUs.
- Lost outputs. API inputs, outputs and logs are deleted after one hour by default. A webhook handler that stores the output URL instead of the bytes ends up with dead links. Copy files to your own storage on
succeeded. - Duplicate side effects. Webhooks are retried, so the same
webhook-idcan arrive twice. Deduplicate on it, and make the job update idempotent. - Orphaned predictions. If you create a prediction and crash before saving its id, you pay for work nobody will collect. Write the job record first, attach the prediction id immediately after creation, and run a sweeper that polls non-terminal jobs older than a threshold.
- Unpinned versions. A community model's owner pushes a new version with a renamed input and your calls start failing validation. Pin versions; upgrade deliberately behind an evaluation.
- Silent cold starts. p50 latency looks fine while p99 is a boot. Record the time spent in
startingseparately fromprocessing(the prediction object carries timestamps) so the two are not blended. - Runaway cost. A retry loop that re-creates predictions on client timeout, rather than re-reading the existing one, multiplies spend. Retry the GET, not the POST, and cancel predictions whose caller has gone.
Trade-offs
Replicate's strength is time to first result: a huge catalogue, a uniform API over very different models, Cog as an honest packaging standard, and per-second billing that makes experiments cheap. Its weaknesses are the mirror image. You do not control the scheduler, the region placement or the kernel-level serving stack, so you cannot apply the optimisations a dedicated LLM server gives you, and at steady high utilisation, renting GPUs and running your own server is usually cheaper. For language models at volume, compare against providers built around optimised LLM serving, such as the approach described in Together AI; for media generation, the serving concerns in diffusion model serving apply whether you host the model or Replicate does.
The decision rule that holds up: use public and official models for prototyping and low, bursty volume; move to a deployment when you need a stable endpoint, a private model or warm capacity; and move off the platform when utilisation is high enough that per-second rental of idle-free capacity costs more than owning the stack. The general architecture of that stack is covered in LLM serving architecture, and how to size it in GPU serving capacity planning.
What to do next
- Pin every community model you call to an explicit version and record it in config.
- Measure cold-boot and warm run time for your model separately, using the prediction timestamps, and set latency budgets on each.
- Switch slow paths from polling to webhooks with
["completed"], verify the HMAC signature, and deduplicate onwebhook-id. - Copy every output file to your own storage inside the webhook flow; the API copy expires after an hour.
- Persist the prediction id before acknowledging the user, and add a sweeper for non-terminal jobs.
- Compute active GPU-hours per day versus a warm minimum, and choose public, official or a deployment with a scheduled
min_instancesfrom that number. - If you publish your own model, keep weight loading in
setup(), use the currentBaseRunner/run()names, and time a coldcog predictbefore shipping.