Every cloud offers several ways to run a container, and they differ less in what they run than in how much of the platform you operate. At one end you hand over an image and a port and the provider does everything else, including scaling to zero. At the other you run Kubernetes and own the scheduling, networking and upgrade model yourself, even when the control plane is managed.

This article covers the services between those poles: Google Cloud Run, Azure Container Apps, Amazon ECS on Fargate or EC2, and Azure Container Instances. It explains the two scaling models underneath them, shows how to size a service with arithmetic rather than guesswork, walks through real deployment definitions, and lists the failure modes that catch teams moving off virtual machines. Managed Kubernetes is the next rung up and has its own article; functions, the rung below, are covered in the serverless article.

Advertisement

The abstraction ladder

The useful way to compare container services is by what you still own. The table orders them from least to most operational surface.

RungExamplesUnit you manageYou still own
Single container groupAzure Container InstancesOne group of containersRestarts, scaling, load balancing
Request/event-drivenCloud Run, Azure Container AppsService revisionImage, CPU/memory, concurrency, limits
Task orchestratorECS on Fargate, ECS on EC2Task definition + serviceNetworking, load balancer, scaling policy, (EC2) hosts
Managed KubernetesEKS, GKE, AKSPods, nodes, cluster add-onsAlmost everything above the control plane

The bottom rung is easy to overlook. Azure Container Instances runs a container group, one or more containers sharing a network namespace and lifetime, with no cluster at all. It has no built-in autoscaling or rolling deployment, so it fits one-off jobs, build agents, scheduled tasks and burst capacity that another system starts and stops, not long-running services behind a load balancer. If you find yourself scripting restarts, health checks and replica counts around it, you have outgrown it and want the next rung.

Moving up the ladder buys control: sidecars, custom schedulers, GPUs, daemon processes, fine-grained networking. It costs operational time. The common mistake is starting on the top rung because it is the most flexible, then spending a team's quarter running a cluster for a handful of stateless HTTP services that a request-driven platform would have run unattended.

What every service does with your image

CI buildimage + digestRegistryECR, Artifact Reg., ACRService definitiontask def / revision: image, CPU, memory, env, rolesdeploySchedulerplaces instancesAutoscalerrequests, CPU, queueInstancecontainerInstancecontainerInstancecontainerLoad balancerhealth checksClientsVPC, secrets, logs, identity
The common pipeline: an immutable image in a registry, a versioned service definition, a scheduler that places instances, an autoscaler, and a load balancer that only routes to healthy instances.

Under different names, every service follows the same pipeline. CI builds an image and pushes it to a registry. A versioned definition (an ECS task definition revision, a Cloud Run or Container Apps revision) binds the image to CPU, memory, environment, secrets and an identity. A scheduler places instances, an autoscaler changes their number, and a load balancer routes only to instances that pass health checks. Deploy by image digest rather than a mutable tag, so a revision always means the same bytes and a rollback is exact.

Advertisement

Two scaling models: requests versus tasks

Request-driven platforms (Cloud Run, and Container Apps with HTTP scale rules) sit in the request path. They know how many requests each instance is serving and scale on that. The key setting is concurrency, the maximum simultaneous requests per instance. On Cloud Run the documented default is 80 (new services created with gcloud or Terraform may default to 80 times the vCPU count) and the maximum is 1,000; requests may run up to 60 minutes. With no traffic, a service can scale to zero and cost nothing, at the price of a cold start on the next request. Container Apps scales with KEDA-based rules, so besides HTTP it can scale on queue depth or event backlog, and on the Consumption profile an app at zero replicas incurs no usage charge.

Task orchestrators (ECS) do not see requests. A service keeps a desired count of tasks running, and scaling policies change that count from metrics: CPU, memory, load balancer requests per target, or a custom queue metric. This is slower to react but fits workloads that are not request-shaped: queue consumers, long-lived connections, batch workers and anything that must never scale to zero. For the policy mechanics, see cloud autoscaling.

A cold start is everything between the platform deciding to add an instance and that instance serving a request: pulling the image, starting the runtime, loading configuration, opening database pools and warming caches. Its cost depends almost entirely on your image and start-up code, not on the platform. Keep images small and layered so pulls are cached, defer optional initialisation until after the port is listening, and expose a readiness check that only passes once the instance can really serve. Then measure it: deploy a new revision, send one request, and time the first response against a warm one. If the difference breaks your latency target, keep a minimum number of warm instances rather than tuning further.

Sizing from arrival rate and latency

Little's law says the average number of requests in flight equals arrival rate times time in system. It turns traffic figures into instance counts without guesswork.

Worked example: an API receives 400 requests per second at steady state and 2,000 at peak, with 120 ms average latency. In flight at steady state: 400 x 0.12 = 48 requests. Load testing shows one instance with 1 vCPU keeps its latency up to 40 concurrent requests, so set concurrency to 40 and target 70 percent: 48 / (40 x 0.7) = 1.7, so 2 instances. At peak: 2,000 x 0.12 = 240 in flight, 240 / 28 = 8.6, so 9 instances. Set a maximum instance count comfortably above 9, say 15, so a latency spike (which raises in-flight requests at the same arrival rate) does not hit the ceiling, and set minimum instances to 2 so steady-state traffic never waits for a cold start.

Do not trust the default concurrency. A CPU-heavy handler that saturates a vCPU at 5 concurrent requests will queue at 80 and every request slows down, while autoscaling sees nothing wrong. Measure the knee of the latency curve per instance and set concurrency just below it.

# Cloud Run: deploy by digest with explicit concurrency and scaling bounds
gcloud run deploy orders-api \
  --image=us-docker.pkg.dev/acme/apps/orders-api@sha256:3f1c... \
  --region=us-central1 \
  --cpu=1 --memory=512Mi \
  --concurrency=40 \
  --min-instances=2 --max-instances=15 \
  --timeout=30 \
  --service-account=orders-api@acme.iam.gserviceaccount.com

ECS: task definitions and services

ECS splits configuration in two. A task definition describes what runs: containers, images, CPU and memory, ports, health checks, logging, and two roles. The execution role lets the ECS agent pull the image and fetch secrets; the task role is what your code uses to call cloud APIs. Mixing them up is a common source of over-broad permissions. A service then keeps a number of copies of one task definition revision running behind a load balancer.

{
  "family": "orders-worker",
  "requiresCompatibilities": ["FARGATE"],
  "networkMode": "awsvpc",
  "cpu": "1024",
  "memory": "2048",
  "executionRoleArn": "arn:aws:iam::123456789012:role/orders-exec",
  "taskRoleArn": "arn:aws:iam::123456789012:role/orders-task",
  "containerDefinitions": [{
    "name": "worker",
    "image": "123456789012.dkr.ecr.eu-west-1.amazonaws.com/orders@sha256:3f1c...",
    "essential": true,
    "stopTimeout": 120,
    "healthCheck": {"command": ["CMD-SHELL", "curl -f http://localhost:8080/healthz || exit 1"],
                    "interval": 15, "timeout": 5, "retries": 3},
    "logConfiguration": {"logDriver": "awslogs", "options": {
      "awslogs-group": "/ecs/orders", "awslogs-region": "eu-west-1",
      "awslogs-stream-prefix": "worker"}}
  }]
}

On Fargate, CPU and memory come in fixed combinations (a 16 vCPU task takes 32 to 120 GB), and ephemeral storage ranges from 21 to 200 GiB. The awsvpc network mode gives each task its own network interface and private IP, which makes security groups per task possible and makes subnet IP exhaustion a real capacity limit. ECS on EC2 instead runs tasks on instances you manage, which you choose for GPUs, daemon agents, very large tasks or sustained utilisation high enough that per-task Fargate pricing costs more than reserved hosts.

Capacity providers let one service mix FARGATE and FARGATE_SPOT, for example a base of 2 tasks on regular capacity and the rest weighted towards Spot. Spot tasks get a two-minute warning, delivered as an EventBridge task state change and a SIGTERM to the task; the spot capacity article covers the pattern.

Shutdown, health and deployments

Every one of these platforms stops containers by sending SIGTERM, waiting, then sending SIGKILL. On ECS the wait is stopTimeout, 30 seconds by default and at most 120. Cloud Run's grace period is short, measured in seconds rather than minutes, so check the current documentation rather than assuming. Code must stop accepting work, finish or hand back what it holds, and exit before the deadline.

import signal, sys, threading

stopping = threading.Event()

def on_sigterm(signum, frame):
    stopping.set()               # readiness check starts failing; LB drains us

signal.signal(signal.SIGTERM, on_sigterm)

def worker_loop(queue):
    while not stopping.is_set():
        msg = queue.receive(wait_seconds=5)      # short poll so we notice SIGTERM
        if msg is None:
            continue
        process(msg)                              # must be idempotent
        queue.delete(msg)                         # ack only after success
    sys.exit(0)                                   # unacked messages become visible again

Deployments follow the same pattern with different names. ECS rolling updates are bounded by minimumHealthyPercent and maximumPercent, and the deployment circuit breaker can roll back a revision whose tasks keep failing health checks. Cloud Run and Container Apps keep revisions side by side and split traffic by percentage, which gives canary releases without extra tooling; blue-green deployment covers the cut-over patterns in general. Whichever you use, a health check that returns 200 before dependencies are ready will pass a broken release; check what the service actually needs.

Networking and identity

Request-driven platforms give you a public HTTPS endpoint by default. For private-only services, restrict ingress to internal traffic and route egress into your VPC through the platform's VPC connectivity option, or the service cannot reach private databases. On ECS, tasks sit in your subnets from the start, so the work is subnet sizing, security groups and NAT or private endpoints for registry and log traffic.

Give every service its own identity: a dedicated service account on Cloud Run, a managed identity on Container Apps, a task role on ECS. Pull secrets at start-up from a secrets manager by reference rather than baking them into images or plain environment variables in the definition.

Cost and trade-offs

FactorRequest-driven (Cloud Run, Container Apps)ECS on FargateECS on EC2
Idle costZero when scaled to zeroPer running taskPer running host
Spiky HTTP trafficBest fitGood, slower scalingNeeds spare hosts
Queue workers, long jobsPossible with event scalingGood fitGood fit
Steady high utilisationOften most expensiveMiddleCheapest with commitments
GPUs, daemons, host accessLimitedLimitedFull

The pattern that works for most teams is to start stateless HTTP services on a request-driven platform, move steady, heavy or host-dependent workloads to ECS on EC2 or to Kubernetes when the numbers justify it, and treat Kubernetes as the choice when you need its ecosystem rather than as the default.

Failure modes

  • Default concurrency on CPU-bound code. Requests queue inside instances, latency climbs and the autoscaler sees no pressure.
  • Cold starts on the critical path. Scale-to-zero plus a heavy runtime start gives multi-second first requests; set minimum instances for latency-sensitive services.
  • Ignoring SIGTERM. In-flight work is killed at every deployment and scale-in; unacked messages are redelivered, so non-idempotent handlers double-process.
  • Subnet exhaustion. Every awsvpc task takes an IP, and a scale-out stalls when the subnet runs dry.
  • Mutable image tags. Two tasks of the same revision run different code after a tag is re-pushed.
  • Execution role used as the task role. Application code inherits permissions meant only for the agent.

What to do next

  1. List your services and mark each as request-driven, queue-driven, long-running or host-dependent; pick the lowest rung that fits each.
  2. Load test one instance to find its latency knee and set concurrency just below it.
  3. Size minimum and maximum instances with Little's law from steady-state and peak traffic.
  4. Deploy by image digest and give each service its own identity and secrets references.
  5. Handle SIGTERM, make handlers idempotent and set stopTimeout or the platform grace period deliberately.
  6. Enable rollback: the ECS circuit breaker or revision traffic splitting.
  7. Review cost monthly and move steady high-utilisation services up the ladder only when the numbers justify it.
Key takeaway: Container services differ mainly in how much platform you operate. Request-driven services such as Cloud Run and Azure Container Apps scale on concurrency, can scale to zero and suit stateless HTTP; ECS keeps a desired count of tasks on Fargate or your own EC2 hosts and suits workers, long-lived processes and steady load; Kubernetes is the next rung when you need its ecosystem. Size concurrency from a measured latency knee and instance counts from Little's law, deploy by digest, handle SIGTERM with idempotent handlers, and give each service its own identity.