Every cloud offers several ways to run a container, and they differ less in what they run than in how much of the platform you operate. At one end you hand over an image and a port and the provider does everything else, including scaling to zero. At the other you run Kubernetes and own the scheduling, networking and upgrade model yourself, even when the control plane is managed.
This article covers the services between those poles: Google Cloud Run, Azure Container Apps, Amazon ECS on Fargate or EC2, and Azure Container Instances. It explains the two scaling models underneath them, shows how to size a service with arithmetic rather than guesswork, walks through real deployment definitions, and lists the failure modes that catch teams moving off virtual machines. Managed Kubernetes is the next rung up and has its own article; functions, the rung below, are covered in the serverless article.
The abstraction ladder
The useful way to compare container services is by what you still own. The table orders them from least to most operational surface.
| Rung | Examples | Unit you manage | You still own |
|---|---|---|---|
| Single container group | Azure Container Instances | One group of containers | Restarts, scaling, load balancing |
| Request/event-driven | Cloud Run, Azure Container Apps | Service revision | Image, CPU/memory, concurrency, limits |
| Task orchestrator | ECS on Fargate, ECS on EC2 | Task definition + service | Networking, load balancer, scaling policy, (EC2) hosts |
| Managed Kubernetes | EKS, GKE, AKS | Pods, nodes, cluster add-ons | Almost everything above the control plane |
The bottom rung is easy to overlook. Azure Container Instances runs a container group, one or more containers sharing a network namespace and lifetime, with no cluster at all. It has no built-in autoscaling or rolling deployment, so it fits one-off jobs, build agents, scheduled tasks and burst capacity that another system starts and stops, not long-running services behind a load balancer. If you find yourself scripting restarts, health checks and replica counts around it, you have outgrown it and want the next rung.
Moving up the ladder buys control: sidecars, custom schedulers, GPUs, daemon processes, fine-grained networking. It costs operational time. The common mistake is starting on the top rung because it is the most flexible, then spending a team's quarter running a cluster for a handful of stateless HTTP services that a request-driven platform would have run unattended.
What every service does with your image
Under different names, every service follows the same pipeline. CI builds an image and pushes it to a registry. A versioned definition (an ECS task definition revision, a Cloud Run or Container Apps revision) binds the image to CPU, memory, environment, secrets and an identity. A scheduler places instances, an autoscaler changes their number, and a load balancer routes only to instances that pass health checks. Deploy by image digest rather than a mutable tag, so a revision always means the same bytes and a rollback is exact.
Two scaling models: requests versus tasks
Request-driven platforms (Cloud Run, and Container Apps with HTTP scale rules) sit in the request path. They know how many requests each instance is serving and scale on that. The key setting is concurrency, the maximum simultaneous requests per instance. On Cloud Run the documented default is 80 (new services created with gcloud or Terraform may default to 80 times the vCPU count) and the maximum is 1,000; requests may run up to 60 minutes. With no traffic, a service can scale to zero and cost nothing, at the price of a cold start on the next request. Container Apps scales with KEDA-based rules, so besides HTTP it can scale on queue depth or event backlog, and on the Consumption profile an app at zero replicas incurs no usage charge.
Task orchestrators (ECS) do not see requests. A service keeps a desired count of tasks running, and scaling policies change that count from metrics: CPU, memory, load balancer requests per target, or a custom queue metric. This is slower to react but fits workloads that are not request-shaped: queue consumers, long-lived connections, batch workers and anything that must never scale to zero. For the policy mechanics, see cloud autoscaling.
A cold start is everything between the platform deciding to add an instance and that instance serving a request: pulling the image, starting the runtime, loading configuration, opening database pools and warming caches. Its cost depends almost entirely on your image and start-up code, not on the platform. Keep images small and layered so pulls are cached, defer optional initialisation until after the port is listening, and expose a readiness check that only passes once the instance can really serve. Then measure it: deploy a new revision, send one request, and time the first response against a warm one. If the difference breaks your latency target, keep a minimum number of warm instances rather than tuning further.
Sizing from arrival rate and latency
Little's law says the average number of requests in flight equals arrival rate times time in system. It turns traffic figures into instance counts without guesswork.
Worked example: an API receives 400 requests per second at steady state and 2,000 at peak, with 120 ms average latency. In flight at steady state: 400 x 0.12 = 48 requests. Load testing shows one instance with 1 vCPU keeps its latency up to 40 concurrent requests, so set concurrency to 40 and target 70 percent: 48 / (40 x 0.7) = 1.7, so 2 instances. At peak: 2,000 x 0.12 = 240 in flight, 240 / 28 = 8.6, so 9 instances. Set a maximum instance count comfortably above 9, say 15, so a latency spike (which raises in-flight requests at the same arrival rate) does not hit the ceiling, and set minimum instances to 2 so steady-state traffic never waits for a cold start.
Do not trust the default concurrency. A CPU-heavy handler that saturates a vCPU at 5 concurrent requests will queue at 80 and every request slows down, while autoscaling sees nothing wrong. Measure the knee of the latency curve per instance and set concurrency just below it.
# Cloud Run: deploy by digest with explicit concurrency and scaling bounds
gcloud run deploy orders-api \
--image=us-docker.pkg.dev/acme/apps/orders-api@sha256:3f1c... \
--region=us-central1 \
--cpu=1 --memory=512Mi \
--concurrency=40 \
--min-instances=2 --max-instances=15 \
--timeout=30 \
--service-account=orders-api@acme.iam.gserviceaccount.com
ECS: task definitions and services
ECS splits configuration in two. A task definition describes what runs: containers, images, CPU and memory, ports, health checks, logging, and two roles. The execution role lets the ECS agent pull the image and fetch secrets; the task role is what your code uses to call cloud APIs. Mixing them up is a common source of over-broad permissions. A service then keeps a number of copies of one task definition revision running behind a load balancer.
{
"family": "orders-worker",
"requiresCompatibilities": ["FARGATE"],
"networkMode": "awsvpc",
"cpu": "1024",
"memory": "2048",
"executionRoleArn": "arn:aws:iam::123456789012:role/orders-exec",
"taskRoleArn": "arn:aws:iam::123456789012:role/orders-task",
"containerDefinitions": [{
"name": "worker",
"image": "123456789012.dkr.ecr.eu-west-1.amazonaws.com/orders@sha256:3f1c...",
"essential": true,
"stopTimeout": 120,
"healthCheck": {"command": ["CMD-SHELL", "curl -f http://localhost:8080/healthz || exit 1"],
"interval": 15, "timeout": 5, "retries": 3},
"logConfiguration": {"logDriver": "awslogs", "options": {
"awslogs-group": "/ecs/orders", "awslogs-region": "eu-west-1",
"awslogs-stream-prefix": "worker"}}
}]
}On Fargate, CPU and memory come in fixed combinations (a 16 vCPU task takes 32 to 120 GB), and ephemeral storage ranges from 21 to 200 GiB. The awsvpc network mode gives each task its own network interface and private IP, which makes security groups per task possible and makes subnet IP exhaustion a real capacity limit. ECS on EC2 instead runs tasks on instances you manage, which you choose for GPUs, daemon agents, very large tasks or sustained utilisation high enough that per-task Fargate pricing costs more than reserved hosts.
Capacity providers let one service mix FARGATE and FARGATE_SPOT, for example a base of 2 tasks on regular capacity and the rest weighted towards Spot. Spot tasks get a two-minute warning, delivered as an EventBridge task state change and a SIGTERM to the task; the spot capacity article covers the pattern.
Shutdown, health and deployments
Every one of these platforms stops containers by sending SIGTERM, waiting, then sending SIGKILL. On ECS the wait is stopTimeout, 30 seconds by default and at most 120. Cloud Run's grace period is short, measured in seconds rather than minutes, so check the current documentation rather than assuming. Code must stop accepting work, finish or hand back what it holds, and exit before the deadline.
import signal, sys, threading
stopping = threading.Event()
def on_sigterm(signum, frame):
stopping.set() # readiness check starts failing; LB drains us
signal.signal(signal.SIGTERM, on_sigterm)
def worker_loop(queue):
while not stopping.is_set():
msg = queue.receive(wait_seconds=5) # short poll so we notice SIGTERM
if msg is None:
continue
process(msg) # must be idempotent
queue.delete(msg) # ack only after success
sys.exit(0) # unacked messages become visible againDeployments follow the same pattern with different names. ECS rolling updates are bounded by minimumHealthyPercent and maximumPercent, and the deployment circuit breaker can roll back a revision whose tasks keep failing health checks. Cloud Run and Container Apps keep revisions side by side and split traffic by percentage, which gives canary releases without extra tooling; blue-green deployment covers the cut-over patterns in general. Whichever you use, a health check that returns 200 before dependencies are ready will pass a broken release; check what the service actually needs.
Networking and identity
Request-driven platforms give you a public HTTPS endpoint by default. For private-only services, restrict ingress to internal traffic and route egress into your VPC through the platform's VPC connectivity option, or the service cannot reach private databases. On ECS, tasks sit in your subnets from the start, so the work is subnet sizing, security groups and NAT or private endpoints for registry and log traffic.
Give every service its own identity: a dedicated service account on Cloud Run, a managed identity on Container Apps, a task role on ECS. Pull secrets at start-up from a secrets manager by reference rather than baking them into images or plain environment variables in the definition.
Cost and trade-offs
| Factor | Request-driven (Cloud Run, Container Apps) | ECS on Fargate | ECS on EC2 |
|---|---|---|---|
| Idle cost | Zero when scaled to zero | Per running task | Per running host |
| Spiky HTTP traffic | Best fit | Good, slower scaling | Needs spare hosts |
| Queue workers, long jobs | Possible with event scaling | Good fit | Good fit |
| Steady high utilisation | Often most expensive | Middle | Cheapest with commitments |
| GPUs, daemons, host access | Limited | Limited | Full |
The pattern that works for most teams is to start stateless HTTP services on a request-driven platform, move steady, heavy or host-dependent workloads to ECS on EC2 or to Kubernetes when the numbers justify it, and treat Kubernetes as the choice when you need its ecosystem rather than as the default.
Failure modes
- Default concurrency on CPU-bound code. Requests queue inside instances, latency climbs and the autoscaler sees no pressure.
- Cold starts on the critical path. Scale-to-zero plus a heavy runtime start gives multi-second first requests; set minimum instances for latency-sensitive services.
- Ignoring SIGTERM. In-flight work is killed at every deployment and scale-in; unacked messages are redelivered, so non-idempotent handlers double-process.
- Subnet exhaustion. Every awsvpc task takes an IP, and a scale-out stalls when the subnet runs dry.
- Mutable image tags. Two tasks of the same revision run different code after a tag is re-pushed.
- Execution role used as the task role. Application code inherits permissions meant only for the agent.
What to do next
- List your services and mark each as request-driven, queue-driven, long-running or host-dependent; pick the lowest rung that fits each.
- Load test one instance to find its latency knee and set concurrency just below it.
- Size minimum and maximum instances with Little's law from steady-state and peak traffic.
- Deploy by image digest and give each service its own identity and secrets references.
- Handle SIGTERM, make handlers idempotent and set stopTimeout or the platform grace period deliberately.
- Enable rollback: the ECS circuit breaker or revision traffic splitting.
- Review cost monthly and move steady high-utilisation services up the ladder only when the numbers justify it.