Replicate is a hosted platform for running machine-learning models behind an HTTP API. You send JSON inputs to a model, Replicate schedules the work onto a GPU or CPU instance running that model's container, and you get back outputs: text, numbers, or URLs to generated files. Its catalogue is large and skewed towards media models (image, video, audio, upscaling) as well as language models, and anyone can publish a model into it by packaging the model with Replicate's open-source tool, Cog. Cloudflare announced in November 2025 that it was acquiring Replicate; Replicate has said its API is unchanged, and everything below describes that API.

The interface looks trivially simple, which is why teams get surprised in production. Every call is really an asynchronous job with a lifecycle; whether the job waits seconds or minutes for a GPU depends on which of three billing models the model falls under; outputs vanish after an hour; and webhooks arrive more than once. This article covers what a prediction is, how Cog packages a model, where cold starts come from, how to receive results safely, and how to choose between public, official and deployed models.

Models, versions and predictions

Three nouns carry the whole system. A model is a named thing, owner/name, such as a public image generator or your company's private fine-tune. A version is an immutable build of that model: a specific container image with specific weights and a specific input schema, identified by a long hash. A prediction is one execution of one version against one set of inputs. Predictions are the unit you create, poll, cancel and pay for.

Pinning matters. Calling a community model by owner/name:version gives you the exact build you tested; anything less than a pinned version can change under you when the owner publishes, including its input names, defaults or output format overnight. Official models are the exception: they are maintained by Replicate with a stable interface and are called through the model-scoped endpoint without a version.

There are three creation endpoints, one per way of addressing compute:

EndpointAddressesTypical use
POST /v1/predictionsa specific version (in the body)community models, pinned
POST /v1/models/{owner}/{name}/predictionsan official modelstable, maintained models
POST /v1/deployments/{owner}/{name}/predictionsyour deploymentproduction traffic on dedicated instances

A prediction moves through starting, processing and one terminal state: succeeded, failed or canceled. starting covers both queueing and any cold boot; processing means the model code is running. Running predictions can be stopped with POST /v1/predictions/{id}/cancel, which you should do whenever the user who asked for the result has gone away, because time-billed models keep billing.

One prediction, end to endYour servicePOST /v1/predictionsReplicate APIid + status=startingQueueper model / deploymentInstanceCog containerinputCold boot (if no warm instance)pull image, setup() loads weightsrun()status=processing, logs streamTerminal statesucceeded / failed / canceledWebhook to your endpointverify HMAC, dedupe on webhook-idcompleted eventCopy outputs to your storageAPI data deleted after 1 houror poll / Prefer: wait
The lifecycle of a prediction. The only step you do not control is the cold boot; everything after the terminal state is your responsibility.

Packaging a model with Cog

Cog is the packaging layer. It reads a cog.yaml describing the environment and a Python class describing how to load and run the model, then builds a container that exposes an HTTP prediction server with an input schema derived from your type hints. Current Cog releases name the class BaseRunner with a run() method and point at it with a run: key; the older BasePredictor, predict() and predict: names still work for existing models but are deprecated and make Cog print warnings. You will meet both in the wild.

# cog.yaml
build:
  gpu: true
  python_version: "3.12"
  python_requirements: requirements.txt   # python_packages is deprecated
  system_packages:
    - ffmpeg
run: "run.py:Runner"                       # older models use predict: "predict.py:Predictor"
# run.py
from cog import BaseRunner, Input, Path
import torch

class Runner(BaseRunner):
    def setup(self) -> None:
        # Runs once per instance boot. Everything slow belongs here.
        self.pipe = load_pipeline("/src/weights", dtype=torch.float16).to("cuda")

    def run(
        self,
        prompt: str = Input(description="What to generate"),
        steps: int = Input(description="Denoising steps", default=28, ge=1, le=80),
        seed: int = Input(description="Random seed; -1 for random", default=-1),
    ) -> Path:
        g = None if seed < 0 else torch.Generator("cuda").manual_seed(seed)
        image = self.pipe(prompt, num_inference_steps=steps, generator=g).images[0]
        out = Path("/tmp/out.png")
        image.save(out)
        return out

Three details in that file decide your production behaviour. First, the split between setup() and run() is the split between cold-start cost and per-request cost: anything in setup() is paid once per instance, anything in run() on every call. Second, Input() constraints such as ge, le and choices become the published schema, so bad inputs are rejected before your code runs. Third, returning a Path makes Replicate upload the file and hand callers a URL. For token streaming, Cog supports returning an iterator; for concurrency within one instance, an async def run() plus the @cog.concurrent(max=N) decorator (cog 0.21.0 and later) replaces the deprecated concurrency.max stanza. Only enable concurrency if the model actually batches or overlaps well on one GPU; otherwise requests simply contend for the same memory.

Image and weight size are cold-start time: time a cold cog predict on a clean machine before you publish.

Billing models, deployments and cold starts

Replicate bills in three different ways, and the difference is not just price: it determines whether a cold boot is your problem.

  • Public community models run on shared hardware. You pay only for the time the model is actively processing your request; setup and idle time are free. The price of that generosity is that an unpopular model may have no warm instance, so your first call waits for a boot, and you share a queue with other users.
  • Official models are priced per unit of input or output (per token, per image, per second of video, depending on the model) rather than per second of hardware. Cost is predictable per request, and Replicate keeps them warm.
  • Private models and deployments run on dedicated instances. You pay for all the time an instance is online: setting up, idle, and active. Replicate documents one exception: fast-booting fine-tunes are billed only for active time whether public or private.

A deployment is how you take control of capacity. It binds a model version to a hardware type and a minimum and maximum instance count, and gives you a private endpoint. Setting the minimum above zero removes cold boots for the first few concurrent requests and puts a floor under your bill; setting it to zero lets the deployment scale to nothing when idle, and the first request after a quiet period pays the boot. Changing the deployment's version rolls new instances without changing the URL your callers use, which makes deployments the right abstraction even for public models you depend on.

Which model type you call decides who pays for idle timePublic community modelshared hardwarepay: processing time onlyOfficial modelmaintained, stable APIpay: per input / output unitPrivate model / deploymentdedicated instancespay: setup + idle + activeRisk: cold boots, shared queueRisk: fewer knobsRisk: paying for idle min_instances
The three billing models trade predictability against control. Who pays for idle time is the question to ask first.

Calling it: sync, polling and signed webhooks

There are three ways to wait for a result. Synchronous: send Prefer: wait and the API holds the request open, by default up to 60 seconds, returning the finished prediction if it completes in time and the in-progress one otherwise. Good for fast models in interactive paths. Polling: create, then GET the prediction's URL until it is terminal. Simple, but wasteful at scale. Webhooks: pass a webhook URL and a webhook_events_filter drawn from start, output, logs and completed; Replicate POSTs the prediction object to you as it changes. For anything slow, use webhooks with ["completed"] and keep polling as a fallback sweep.

import os, requests

API = "https://api.replicate.com/v1"
H = {"Authorization": f"Bearer {os.environ['REPLICATE_API_TOKEN']}",
     "Content-Type": "application/json"}

def submit(job_id: str, prompt: str) -> str:
    body = {
        "input": {"prompt": prompt, "steps": 28},
        "webhook": f"https://api.example.com/hooks/replicate?job={job_id}",
        "webhook_events_filter": ["completed"],
    }
    r = requests.post(f"{API}/deployments/acme/image-gen/predictions",
                      json=body, headers=H, timeout=30)
    r.raise_for_status()
    pred = r.json()
    save_job(job_id, prediction_id=pred["id"], status=pred["status"])  # persist BEFORE returning
    return pred["id"]

Webhooks are signed. Each request carries webhook-id, webhook-timestamp and webhook-signature headers. The signed content is the id, the timestamp and the raw body joined with full stops, HMAC-SHA256 with your signing secret, which you fetch once from GET /v1/webhooks/default/secret. The secret starts with whsec_ and the part after that prefix is base64. The signature header can hold several space-separated v1,<base64> entries; accept if any matches.

import base64, hashlib, hmac, time

def verify(headers, raw_body: bytes, secret: str, tolerance_s: int = 300) -> bool:
    wid, ts = headers["webhook-id"], headers["webhook-timestamp"]
    if abs(time.time() - int(ts)) > tolerance_s:          # reject replays
        return False
    key = base64.b64decode(secret.split("_", 1)[1])        # strip "whsec_"
    signed = f"{wid}.{ts}.".encode() + raw_body             # raw bytes, not re-serialised JSON
    expected = base64.b64encode(hmac.new(key, signed, hashlib.sha256).digest()).decode()
    for entry in headers["webhook-signature"].split():
        _, sig = entry.split(",", 1)
        if hmac.compare_digest(sig, expected):
            return True
    return False

def on_webhook(request):
    if not verify(request.headers, request.body, SECRET):
        return 401
    if seen_before(request.headers["webhook-id"]):         # deliveries can repeat
        return 200
    pred = request.json()
    if pred["status"] == "succeeded":
        enqueue_copy_outputs(pred["id"], pred["output"])   # outputs expire; copy now
    mark_job(pred["id"], pred["status"], pred.get("error"))
    return 200                                             # answer fast; work async

Worked example: an on-demand image feature

Suppose a product generates marketing images on demand. Users click a button and expect a result within about 15 seconds; traffic is 2,000 images a day, peaking at around 6 concurrent requests in business hours and near zero overnight. Each image takes about 4 seconds of GPU time once the model is warm, and a cold boot of your fine-tuned container takes about 90 seconds (measure yours; these are illustrative).

OptionLatency behaviourCost behaviour
Public model, pinned versionfine when warm; a 90 s boot after quiet periods breaks the 15 s budgetpay ~2,000 × 4 s of active time per day
Deployment, min 0, max 4first request each morning waits for a bootpay active time plus boot and idle time while instances are up
Deployment, min 1, max 4one warm instance always; bursts above it boot new instancespay 24 h of one instance, plus extra instances while scaled up
Official equivalent modelwarm, per-image pricepay per image; no control over the weights

The arithmetic that decides it: 2,000 × 4 s is about 2.2 GPU-hours of real work per day, while a minimum of one instance is 24 GPU-hours. Keeping an instance warm around the clock therefore costs roughly ten times the useful work. A common answer is a deployment with min_instances raised during business hours and dropped to zero overnight by a scheduled job, plus a webhook-driven UI that shows progress instead of blocking, so the occasional cold start degrades to a slower result rather than a timeout. If the per-image price of an official model that meets your quality bar is below your effective per-image cost, that beats both.

Failure modes

The failures teams actually hit are mostly about state, not GPUs.

  • Lost outputs. API inputs, outputs and logs are deleted after one hour by default. A webhook handler that stores the output URL instead of the bytes ends up with dead links. Copy files to your own storage on succeeded.
  • Duplicate side effects. Webhooks are retried, so the same webhook-id can arrive twice. Deduplicate on it, and make the job update idempotent.
  • Orphaned predictions. If you create a prediction and crash before saving its id, you pay for work nobody will collect. Write the job record first, attach the prediction id immediately after creation, and run a sweeper that polls non-terminal jobs older than a threshold.
  • Unpinned versions. A community model's owner pushes a new version with a renamed input and your calls start failing validation. Pin versions; upgrade deliberately behind an evaluation.
  • Silent cold starts. p50 latency looks fine while p99 is a boot. Record the time spent in starting separately from processing (the prediction object carries timestamps) so the two are not blended.
  • Runaway cost. A retry loop that re-creates predictions on client timeout, rather than re-reading the existing one, multiplies spend. Retry the GET, not the POST, and cancel predictions whose caller has gone.

Trade-offs

Replicate's strength is time to first result: a huge catalogue, a uniform API over very different models, Cog as an honest packaging standard, and per-second billing that makes experiments cheap. Its weaknesses are the mirror image. You do not control the scheduler, the region placement or the kernel-level serving stack, so you cannot apply the optimisations a dedicated LLM server gives you, and at steady high utilisation, renting GPUs and running your own server is usually cheaper. For language models at volume, compare against providers built around optimised LLM serving, such as the approach described in Together AI; for media generation, the serving concerns in diffusion model serving apply whether you host the model or Replicate does.

The decision rule that holds up: use public and official models for prototyping and low, bursty volume; move to a deployment when you need a stable endpoint, a private model or warm capacity; and move off the platform when utilisation is high enough that per-second rental of idle-free capacity costs more than owning the stack. The general architecture of that stack is covered in LLM serving architecture, and how to size it in GPU serving capacity planning.

What to do next

  1. Pin every community model you call to an explicit version and record it in config.
  2. Measure cold-boot and warm run time for your model separately, using the prediction timestamps, and set latency budgets on each.
  3. Switch slow paths from polling to webhooks with ["completed"], verify the HMAC signature, and deduplicate on webhook-id.
  4. Copy every output file to your own storage inside the webhook flow; the API copy expires after an hour.
  5. Persist the prediction id before acknowledging the user, and add a sweeper for non-terminal jobs.
  6. Compute active GPU-hours per day versus a warm minimum, and choose public, official or a deployment with a scheduled min_instances from that number.
  7. If you publish your own model, keep weight loading in setup(), use the current BaseRunner/run() names, and time a cold cog predict before shipping.
Key takeaway: Replicate turns any packaged model into an asynchronous HTTP job. Treat every call as a job with a lifecycle: pin versions, persist ids, receive results by signed, deduplicated webhooks, and copy outputs before they expire. Choose between public, official and deployed models by asking who pays for idle time and who absorbs cold boots, and do that arithmetic with your own measured boot and run times.