Cloud Run runs your container, scales it with traffic and bills for what it uses. That summary is accurate and hides almost everything that decides whether a production service works: what happens to in-flight requests when an instance is stopped, why background threads go silent, where cold starts come from, how many IP addresses a private service uses, and how one service proves its identity to another.
This is the build-and-operate guide. How the platform is built (the front end, the autoscaler, the sandboxes, revisions) is covered in Cloud Run architecture, and concurrency tuning in Cloud Run concurrency and autoscaling. Here we build one realistic system, an internal orders API plus a nightly reconciliation job, both talking to a private Cloud SQL database, and cover every setting you will have to decide along the way. Limits and flags were checked against Google's Cloud Run documentation on 2026-10-01; confirm them again before relying on a specific number, because they change.
What you owe the platform
Cloud Run's container contract is short, and each line has a consequence. Your service must listen for HTTP on the port in the PORT environment variable, which defaults to 8080, and it must start listening within 4 minutes or the instance is considered failed. It must answer each request within the configured request timeout, which can be up to 60 minutes for services; startup time counts against the first request. The platform sets K_SERVICE, K_REVISION and K_CONFIGURATION so your logs can say which revision produced them.
The filesystem is in memory. Anything you write to disk uses the instance's memory allowance and disappears when the instance stops, so a service that writes temporary files can run out of memory without allocating a single object. Instances are disposable: before stopping one, Cloud Run sends SIGTERM, waits 10 seconds, then sends SIGKILL. Anything not finished in those 10 seconds is lost. There are two execution environments: first generation runs in the gVisor sandbox, second generation offers full Linux compatibility, and jobs always use second generation.
The worked example
The system has two workloads. orders-api is an HTTP service called by other internal services and never by the public internet. reconcile is a batch job that runs every night, compares orders against the payment provider's settlement file and writes corrections. Both use a PostgreSQL database on Cloud SQL reachable only on a private IP, and both read the database password from Secret Manager.
A container that behaves
The service code needs three behaviours the contract implies: read the port from the environment, initialise expensive clients once per instance rather than per request, and shut down cleanly on SIGTERM. Gunicorn already handles the signal by stopping new connections and letting workers finish, so the job is mostly to set its graceful timeout below Cloud Run's 10 seconds and to flush anything buffered.
# app.py (Flask, run by gunicorn)
import json, os, signal, sys
from flask import Flask, request
from sqlalchemy import create_engine
app = Flask(__name__)
db_url = (f"postgresql+psycopg2://orders:{os.environ['DB_PASSWORD']}"
f"@{os.environ['DB_HOST']}/orders")
engine = create_engine(db_url, pool_size=5, max_overflow=2,
pool_pre_ping=True) # one pool per instance
def log(severity, msg, **kw):
# JSON on stdout is parsed by Cloud Logging; "severity" becomes the level.
trace = request.headers.get("X-Cloud-Trace-Context", "").split("/")[0] if request else ""
print(json.dumps({"severity": severity, "message": msg,
"revision": os.environ.get("K_REVISION"),
"logging.googleapis.com/trace":
f"projects/{os.environ['GOOGLE_CLOUD_PROJECT']}/traces/{trace}" if trace else None,
**kw}), flush=True)
@app.get("/healthz")
def healthz():
return "ok"
@app.post("/orders")
def create_order():
body = request.get_json()
with engine.begin() as tx:
tx.exec_driver_sql("INSERT INTO orders (id, payload) VALUES (%s, %s) "
"ON CONFLICT (id) DO NOTHING", (body["id"], json.dumps(body)))
log("INFO", "order stored", order_id=body["id"])
return {"status": "stored"}, 201# Dockerfile
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
# graceful-timeout 8s leaves headroom inside Cloud Run's 10 s SIGTERM window
CMD exec gunicorn --bind :$PORT --workers 2 --threads 8 \
--graceful-timeout 8 --timeout 0 app:appThe exec matters: it makes gunicorn process 1 so it receives the signal directly instead of a shell that ignores it. --timeout 0 disables gunicorn's own worker timeout and leaves request timeouts to Cloud Run. The insert is idempotent on the order id because clients retry after timeouts, and a retried request may land on a different instance.
Billing mode decides whether background work runs
Cloud Run has two billing settings. With request-based billing, the default, CPU is allocated only while the instance is handling requests, starting or shutting down, and you pay only for that time. With instance-based billing, CPU stays allocated for the instance's whole lifetime and you pay for all of it. The gcloud flags are --cpu-throttling for request-based and --no-cpu-throttling for instance-based.
The practical effect: under request-based billing, any work that happens after the response is sent (flushing a metrics buffer, an async task queue, an in-process scheduler) runs on a CPU that is throttled close to zero, so it stalls unpredictably. Either finish the work before responding, move it to a queue consumed by another workload, or switch to instance-based billing. Idle instances are kept for at most 15 minutes after their last request unless minimum instances keep them alive. Instance-based billing combined with minimum instances is the pattern for services that must process in the background, such as a streaming pull consumer.
| Workload | Billing setting | Why |
|---|---|---|
| Request/response API, bursty | request-based | pay only during requests; no background work |
| API that flushes telemetry asynchronously | instance-based, or flush before responding | throttled CPU stalls the flush |
| Pull-based consumer, websockets | instance-based plus min instances | work happens outside requests |
| Steady high traffic | compare both | instance-based is often cheaper when instances are always busy |
Cold starts: where they come from and what helps
A cold start is the time between Cloud Run deciding it needs a new instance and that instance serving its first request: starting the sandbox, pulling the image (cached after the first time), starting your process and running its initialisation. Requests waiting for a new instance are queued for a limited time, which the contract gives as the greater of 10 seconds and 3.5 times the average startup time. Your own initialisation is usually the largest part and the only part fully under your control.
Four levers help. Startup CPU boost (--cpu-boost) gives extra CPU during startup and for 10 seconds afterwards; an instance with a 1 vCPU limit is boosted to 2, a 2 vCPU limit to 4, and 4 or more to 8. You pay for the boosted CPU during that window. Minimum instances keep warm capacity so most requests never wait. Lazy initialisation defers clients that only some requests need. And smaller images with fewer dependencies start faster. Concurrency interacts too: the default is 80 concurrent requests per instance and the maximum is 1,000, and a higher value means fewer instances need to start during a spike.
gcloud run deploy orders-api \
--image=us-docker.pkg.dev/acme/apps/orders-api:1.14.2 \
--region=us-central1 \
--service-account=orders-api@acme.iam.gserviceaccount.com \
--no-allow-unauthenticated \
--ingress=internal \
--cpu=1 --memory=1Gi --concurrency=40 --timeout=30s \
--min-instances=1 --max-instances=20 \
--cpu-boost --cpu-throttling \
--network=prod-vpc --subnet=run-us-central1 --vpc-egress=private-ranges-only \
--set-secrets=DB_PASSWORD=orders-db-password:7 \
--set-env-vars=DB_HOST=10.20.0.5,GOOGLE_CLOUD_PROJECT=acme
Private networking with Direct VPC egress
To reach the database on its private IP, the service needs a path into the VPC. Direct VPC egress places instances directly on a subnet, configured with --network, --subnet, optional --network-tags for firewall rules, and --vpc-egress. private-ranges-only sends only private address ranges through the VPC; all-traffic sends everything, which you need if outbound internet traffic must leave through Cloud NAT with a fixed IP address.
Size the subnet before the first incident, not after. Google's documentation says services use about twice as many IP addresses as they have instances at steady state, so 20 instances need around 40 addresses plus headroom for the overlap during a new revision rollout. Addresses are reserved in blocks of 16, and the subnet must be a /26 or larger. Jobs use one address per running task, held for 7 minutes after the task ends. Each instance can push up to 1 Gbps through the VPC before it is throttled. Keep the firewall rule narrow: allow the network tag to reach the database port and nothing else. The broader design of subnets and firewalls is in VPC architecture.
Identity: service accounts, invokers and secrets
Give each workload its own service account with only the permissions it needs; the default Compute Engine service account is far too broad. orders-api needs to read one secret and connect to Cloud SQL; nothing more. --no-allow-unauthenticated makes Cloud Run reject requests without a valid Google-signed identity token, and callers need the roles/run.invoker role on the service. --ingress=internal additionally rejects traffic that does not come from inside your VPC or project, so a leaked token alone is not enough. Internal means routed through a VPC network: a caller on Cloud Run must send its requests through the VPC, for example with Direct VPC egress set to all-traffic, because run.app addresses are not private ranges. An external Application Load Balancer in front needs internal-and-cloud-load-balancing instead.
A caller running on Cloud Run gets an ID token from the metadata server, with the audience set to the URL of the service it calls, and sends it as a bearer token:
import google.auth.transport.requests
import google.oauth2.id_token
import requests
ORDERS_URL = "https://orders-api-abc123-uc.a.run.app"
def create_order(order):
auth_req = google.auth.transport.requests.Request()
token = google.oauth2.id_token.fetch_id_token(auth_req, ORDERS_URL) # audience
r = requests.post(f"{ORDERS_URL}/orders", json=order, timeout=10,
headers={"Authorization": f"Bearer {token}"})
r.raise_for_status()
return r.json()Secrets are mounted from Secret Manager with --set-secrets. The deploy command above pins version 7 rather than latest: an environment-variable secret is resolved when an instance starts, so latest means different instances of the same revision can hold different passwords during a rotation. Pinning makes the rotation an explicit deploy. See Secret Manager architecture for rotation patterns.
Jobs: sharded batch work that finishes
A job runs a container to completion instead of serving requests. An execution has a number of tasks, runs up to a parallelism limit at once, and retries failed tasks up to a maximum. Each task sees CLOUD_RUN_TASK_INDEX (from 0), CLOUD_RUN_TASK_COUNT and CLOUD_RUN_TASK_ATTEMPT, which is all you need to shard work deterministically. A task can run for up to 168 hours (1 hour with GPUs), and an execution can have up to 10,000 tasks.
# reconcile.py: task i of N handles orders whose id hashes to shard i.
import os, sys, zlib
i = int(os.environ["CLOUD_RUN_TASK_INDEX"])
n = int(os.environ["CLOUD_RUN_TASK_COUNT"])
attempt = int(os.environ["CLOUD_RUN_TASK_ATTEMPT"])
rows = settlement_rows_for(date=os.environ["RUN_DATE"])
mine = [r for r in rows if zlib.crc32(r.order_id.encode()) % n == i]
for r in mine:
upsert_correction(r) # idempotent: retries repeat safely
print(f"task {i}/{n} attempt {attempt}: {len(mine)} rows", file=sys.stderr)gcloud run jobs create reconcile \
--image=us-docker.pkg.dev/acme/apps/reconcile:3.2.0 --region=us-central1 \
--tasks=50 --parallelism=10 --max-retries=3 --task-timeout=30m \
--service-account=reconcile@acme.iam.gserviceaccount.com \
--network=prod-vpc --subnet=run-us-central1 --vpc-egress=private-ranges-only
gcloud run jobs execute reconcile --region=us-central1 --waitBecause retries rerun a whole task, every write must be idempotent. Keep parallelism below what the database can absorb: ten tasks each holding a small connection pool is fine; 200 is a connection storm. Trigger the job from Cloud Scheduler, and alert on failed executions, not only on task retries.
Rollouts and rollback
Every deploy creates an immutable revision. For risky changes, deploy with --no-traffic --tag=canary, test the revision at its tag URL, then shift traffic gradually with gcloud run services update-traffic orders-api --to-tags=canary=10. Rollback is moving traffic back to the previous revision, which takes seconds because it is still there. The traffic mechanics are covered in the architecture article; what matters operationally is that database migrations must be compatible with both revisions, because both serve traffic during the shift.
Failure modes and troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Revision fails to become ready | not listening on $PORT, or startup over 4 minutes | bind to 0.0.0.0:$PORT; move slow init out of startup |
| Requests return 429 or 503 under load | max instances reached, or instances cannot start fast enough | raise max instances, add min instances, reduce startup time |
| 504 after exactly the timeout | request longer than --timeout | raise it, or move the work to a job or queue |
| Instance killed, memory exceeded | files written to the in-memory filesystem | stream to Cloud Storage; raise memory |
| Database 'too many connections' | max instances x pool size exceeds the limit | cap max instances and pool size together |
| Background work stops | request-based billing throttles CPU | finish before responding or use instance-based billing |
| Lost work on scale-in | no SIGTERM handling | drain within 10 seconds; make work idempotent |
The database row is the one that surprises most teams. Autoscaling multiplies connections: 20 instances with a pool of 7 each is 140 connections, and a spike to 100 instances is 700. Set max instances from the database's connection limit backwards; the connection model on the database side is covered in Cloud SQL architecture.
What to do next
- Check your container against the contract: listens on
$PORT, starts in well under 4 minutes, writes nothing large to disk, and drains within 10 seconds ofSIGTERM. - List every piece of work that happens after a response is sent, and either move it before the response or choose instance-based billing deliberately.
- Measure cold start time from logs, then try
--cpu-boostand one minimum instance and measure again. - Create one service account per workload, deploy with
--no-allow-unauthenticatedand the narrowest ingress setting that works. - Size the Direct VPC subnet at twice peak instances plus rollout headroom, and never smaller than /26.
- Set max instances from the database connection budget, and write the calculation down next to the deploy config.
- Convert any cron-in-a-container into a job with sharding by
CLOUD_RUN_TASK_INDEXand idempotent writes.