Azure Container Apps runs containers without asking you to operate Kubernetes. Underneath it uses Kubernetes, KEDA for event-driven scaling, Envoy for ingress and optionally Dapr for service-to-service building blocks, but none of those are exposed as APIs you manage. You describe an app (image, resources, ingress, scale rules, secrets) and the platform turns it into revisions and replicas, including scaling to zero when nothing is happening.

That bargain is attractive for HTTP APIs, queue workers and scheduled jobs, and frustrating when you need something Kubernetes would give you directly. This page explains the model from the environment down to a single replica, gives the scaling algorithm with the numbers the platform actually uses, walks through a Bicep deployment and a queue-driven worked example, and ends with failure modes and a checklist. For the wider landscape of container services, see Cloud Containers.

The architecture

Container Apps environment: shared network boundary, logging and workload profilesClientsHTTPSIngress (Envoy)external or internal80%20%Revision api--v7immutable snapshotRevision api--v8canaryReplicas0 .. maxReplicasService Busqueue depthmetricKEDA scalerpolls every 30 sreplicasworker appDapr sidecar optionalContainer Apps jobmanual, schedule, eventWorkload profilesConsumption (serverless) | Dedicated D and E series | GPU | Flexible (preview)Managed identity pulls images and reads Key Vault secrets; no Kubernetes API is exposed.
Ingress splits traffic across immutable revisions; KEDA turns event-source metrics into replica counts for apps and executions for jobs, all inside one environment.

The object model

Four objects carry the whole design. An environment is the boundary: apps inside share a virtual network, a log destination and a set of workload profiles, and can call each other by name. An app is a long-running service with configuration such as ingress, secrets and registries. A revision is an immutable snapshot of the app's template: images, resources, probes, scale rules. A replica is a running instance of a revision, one or more containers sharing a network namespace, like a pod. Jobs sit beside apps for work that runs to completion.

The split between app configuration and revision template decides what a change does. Editing the template (a new image tag, a changed scale rule, different CPU) creates a new revision. Editing configuration (ingress settings, secret values, the revision mode) does not. That is why changing a secret's value does not reach running replicas until they restart or a new revision is created, a frequent surprise covered under failure modes.

Workload profiles

Every environment has a default Consumption profile, and you can add others. The choice decides how replicas are placed and billed.

ProfileShapeScales to zeroBilling model
Consumption0.25 to 4 vCPU, 0.5 to 8 GiB per replicaYesPer replica resource use
Consumption GPUT4 and A100 serverless GPUs in listed regionsYesPer replica resource use
Dedicated D4 to D32General purpose nodes, 4 to 32 vCPUApps can; nodes have min and max countsPer running node instance
Dedicated E4 to E32Memory optimised, 32 to 256 GiBApps can; nodes have min and max countsPer running node instance
Flexible (preview)0.25 to 4 vCPU, up to 16 GiB, single-tenant poolNoConsumption-style plus a management fee

Consumption fits bursty or unpredictable load because idle apps cost nothing at zero replicas. Dedicated fits steady load, large replicas and workloads that need isolation; several apps share each node, so you must leave headroom because the runtime reserves part of every node. Treat the GPU and Flexible rows as region- and preview-dependent and check availability before designing around them.

How scaling actually works

Scaling is horizontal only; there is no vertical scaling of a running revision. Each revision has a minimum and maximum replica count (default 0 and 10; the maximum can be raised to 1,000) and one or more rules. Rules come in three kinds: HTTP, TCP and custom, where custom means any KEDA scaler such as Service Bus, Event Hubs, Kafka, Redis, CPU or memory. With several rules, scale-out begins when any one of them asks for more replicas.

The HTTP and TCP rules are subtler than their names. Every 15 seconds the platform computes the metric as the requests (or connections) seen in the past 15 seconds divided by 15, and compares it with concurrentRequests or concurrentConnections (default 10). So the target behaves like a per-replica rate, not a count of requests in flight. Custom rules are polled every 30 seconds.

The desired count is ceil(currentMetric / target), but growth is stepped: each decision moves to min(maxReplicas, desired, max(4, 2 * current)), so a cold app goes 1, 4, 8, 16, 32 and so on. Scale-up has no stabilisation window; scale-down waits for the condition to hold for 300 seconds and then removes all surplus replicas at once. The 300-second cooldown applies only to the last step, from one replica to zero. If you need an instance always warm, set minReplicas to 1 or more; replicas kept in memory without work may be billed at a lower idle rate on Consumption. General autoscaling theory is in Cloud Autoscaling.

Revisions, traffic and ingress

In single revision mode, a new revision replaces the old one once it is ready, which is what most workers and simple APIs want. In multiple revision mode several revisions stay active and ingress splits traffic by weight, which gives you canaries and blue-green releases without another load balancer. Each revision has its own replicas and scale rules, so a 20 percent canary scales on its own share of traffic.

Multiple mode has a sharp edge for event-driven apps: an old revision that remains active keeps its queue scaler and keeps consuming messages with old code. Microsoft's guidance is to use single mode when scaling on non-HTTP events. Promote by shifting weight to 100 and then deactivating the old revision, not just by reweighting.

Ingress is off, internal (reachable inside the environment and its network) or external. The transport can be HTTP/1.1, HTTP/2, auto-negotiated, or TCP for non-HTTP protocols. Session affinity, IP restrictions and client certificates are ingress settings. Apps without ingress can still run, but they need a non-HTTP scale rule or a minimum above zero, for reasons the failure modes make clear.

Jobs

Jobs run containers to completion with three trigger types. Manual jobs start on an API call or CLI command, schedule jobs run on a cron expression, and event jobs use a KEDA scaler to start executions as messages arrive, one execution per batch of work rather than a long-running consumer. Each job sets a replica timeout, a retry limit and, for parallel work, how many replicas run and how many must complete. Event jobs suit long, heavy, one-message tasks such as rendering or a model batch, where a crash should retry only that unit. A long-running worker app suits many small messages, where container start-up per message would dominate.

A deployment in Bicep

The template below deploys a queue worker that scales from zero on Service Bus depth, pulls its image and reads a Key Vault secret with a user-assigned managed identity, and declares probes and a shutdown grace period. Scale rules authenticate with the same identity, so no connection string sits in the app.

param location string = resourceGroup().location
param envId string
param uamiId string            // user-assigned managed identity resource ID
param uamiClientId string

resource worker 'Microsoft.App/containerApps@2025-01-01' = {
  name: 'orders-worker'
  location: location
  identity: {
    type: 'UserAssigned'
    userAssignedIdentities: { '${uamiId}': {} }
  }
  properties: {
    environmentId: envId
    workloadProfileName: 'Consumption'
    configuration: {
      activeRevisionsMode: 'Single'           // event scaler: keep one active revision
      registries: [ { server: 'contosoacr.azurecr.io', identity: uamiId } ]
      secrets: [
        {
          name: 'db-conn'
          keyVaultUrl: 'https://contoso-kv.vault.azure.net/secrets/db-conn'
          identity: uamiId
        }
      ]
    }
    template: {
      terminationGracePeriodSeconds: 60
      containers: [
        {
          name: 'worker'
          image: 'contosoacr.azurecr.io/orders-worker:1.8.2'
          resources: { cpu: json('0.5'), memory: '1Gi' }
          env: [
            { name: 'DB_CONN', secretRef: 'db-conn' }
            { name: 'AZURE_CLIENT_ID', value: uamiClientId }
          ]
          probes: [
            { type: 'Readiness', httpGet: { path: '/ready', port: 8080 }, periodSeconds: 5 }
            { type: 'Liveness', httpGet: { path: '/live', port: 8080 }, periodSeconds: 10 }
          ]
        }
      ]
      scale: {
        minReplicas: 0
        maxReplicas: 30
        rules: [
          {
            name: 'orders-queue'
            custom: {
              type: 'azure-servicebus'
              metadata: { namespace: 'contoso-sb', queueName: 'orders', messageCount: '20' }
              identity: uamiId
            }
          }
        ]
      }
    }
  }
}

Grant the identity pull rights on the registry, secret read rights in Key Vault and a receive role on the queue before deploying, or the first revision fails to provision or to scale. Check the API version you pin against the current resource reference; preview versions add fields that may change.

Worked example: 900 orders from zero

The worker above is idle with zero replicas. A batch of 900 orders lands on the queue. On the next 30-second poll the scaler sees a non-empty queue and activates the app with one replica. The desired count is ceil(900 / 20) = 45, capped by maxReplicas at 30, and the stepping rule moves the app 1, 4, 8, 16 and then 30, because min(30, 45, max(4, 2 x 16)) = 30. With one decision per poll, full capacity arrives roughly two to three minutes after the first message, plus image pull and start-up time for each new replica.

Two consequences follow. If orders must be processed within a minute, scale-from-zero cannot meet that; keep a minimum of a few replicas or lower messageCount so the steps reach useful capacity sooner. And when the queue drains, replicas are removed only after the 300-second stabilisation window, then the last one after the cooldown, so a worker must handle SIGTERM by finishing or abandoning its current message within the grace period.

Operate it from the signals the platform gives you. Replica count over time shows whether the steps match what you expected; console and system logs in the environment's log destination show probe failures, image pull errors and revision provisioning problems; queue depth and the age of the oldest message tell you whether capacity keeps up. Alarm on the age of the oldest message rather than depth alone, because a deep queue that drains fast is healthy and a shallow queue that never moves is not.

Failure modes

  • The app that never wakes. Ingress disabled, no scale rule and minimum zero: the default HTTP rule has no traffic to observe and nothing ever starts a replica.
  • Secrets that do not change. Updating a secret value creates no revision; running replicas keep the old value until restarted.
  • Old revision still consuming. Multiple revision mode with a queue scaler leaves an old revision processing messages with old code.
  • Cold-start latency. Scale from zero adds activation, image pull and application start. Large images make it worse; keep images small or keep a warm minimum.
  • Lost work on scale-in. Replicas receive SIGTERM and are killed after the grace period. Workers that ignore the signal lose in-flight messages until lock expiry.
  • Actors and scale to zero. Dapr actors do not support scaling to zero.
  • Replica counts are targets. The platform does not guarantee them; capacity limits and maintenance can briefly change the number you see.

Trade-offs

ChooseWhenYou give up
Container AppsHTTP APIs, queue workers, jobs, scale to zero, no cluster teamKubernetes API, CRDs, operators, DaemonSets, node control
AKSFull Kubernetes, custom controllers, mesh, special nodesYou run upgrades, node pools and add-ons
Azure FunctionsSmall event handlers bound to triggersContainer-level control and long-running processes
App ServiceClassic web apps with deployment slotsEvent-driven scaling and scale to zero

What to do next

  1. Pick the workload profile per app: Consumption for bursty work, Dedicated for steady or large replicas, and confirm GPU or Flexible availability in your region.
  2. Give every app a user-assigned managed identity for registry pulls, Key Vault and scalers.
  3. Set minReplicas deliberately: zero only where cold start is acceptable, and never zero with ingress off and no rule.
  4. Use single revision mode for queue-driven apps and multiple mode with weights for HTTP canaries.
  5. Add readiness and liveness probes and handle SIGTERM within terminationGracePeriodSeconds.
  6. Restart or roll a new revision after changing secrets.
  7. Load-test from zero and record the time to full capacity against the 1, 4, 8, 16 step pattern.
Key takeaway: Container Apps gives you revisions, KEDA scaling and managed ingress without operating Kubernetes. Learn the scaling algorithm, because a cold app grows 1, 4, 8, 16 per decision and scale-in waits 300 seconds. Use managed identity everywhere, single revision mode for queue workers, a warm minimum where latency matters, and handle SIGTERM, and remember that secret changes need a restart.