AWS App Runner runs a web application from a container image or a source repository without you managing load balancers, clusters or scaling policies. You give it an image and a port; it gives you an HTTPS URL, scales instances with request concurrency, and redeploys when the image changes. For small teams shipping HTTP services it removed most of the undifferentiated work of running containers on AWS.

Its status changed in 2026. AWS closed App Runner to new customers. Existing customers can keep using it, including creating new services, and AWS says it continues to invest in security and availability but does not plan new features. AWS recommends Amazon ECS Express Mode as the migration target. This article is therefore written for two readers: the team that runs App Runner services today and needs to operate them well, and the team planning the move. It explains how App Runner works, how to configure scaling, health checks and networking correctly, what fails in practice, and how to migrate with a gradual traffic shift.

Advertisement

What App Runner actually is

An App Runner service is one application with one public or private endpoint. Behind the endpoint, App Runner runs instances of your container on AWS-managed Fargate capacity, which is why your account's Fargate vCPU quota bounds how many instances you can get. The service owns the TLS endpoint, request routing, instance lifecycle, rolling deployments, logs and metrics. You own the container, which must listen for HTTP on the configured port, and the configuration: CPU and memory per instance, environment variables and secrets, auto scaling, health checks, networking and an instance role for calls to other AWS services.

Two source types exist. An image service pulls from Amazon ECR (private or public) and can redeploy automatically when a new image is pushed to the tag it tracks. A source code service connects to a repository, builds with an App Runner managed runtime, and redeploys on commits to a branch. Image services are the more portable choice: the same image runs on ECS, EKS or anywhere else, which matters now that migration is on the table.

App Runner request path: managed ingress in front, your VPC only for outbound callsClientsHTTPSApp Runner ingressTLS, load balancingInstance (active)Instance (active)Instance (provisioned)scale on concurrent requests per instancePublic endpointsdefault egressVPC connectorHyperplane ENIsRDS, cacheprivate subnetsNAT / VPC endpointsfor internet + AWS APIsorEgressType=VPCApp Runner-managed trafficimage pulls, logs, secrets: not via your VPCInbound always arrives through App Runner's endpoint (public, or private through an interface endpoint);a VPC connector changes only where your code's outbound connections go
App Runner request path. Ingress, instances and platform traffic are managed by App Runner; a VPC connector changes only your code's outbound route.

Scaling on concurrency, and what you pay for

App Runner scales on concurrent requests per instance, not on CPU. An auto scaling configuration has three numbers. Max concurrency (default 100, allowed 1 to 200) is how many simultaneous requests one instance should handle; when concurrency exceeds it, App Runner adds instances. Min size (default 1, up to 25) is how many instances are always provisioned. Max size (default 25) caps the active instances and therefore your cost and your blast radius on downstream systems.

The billing model is the subtle part. Provisioned instances that are not serving traffic are a warm reserve: you pay for their memory but not their CPU. Active instances pay for both. So min size buys fast response to bursts at a memory-only price, and a higher min size also spreads the service over more Availability Zones. During deployments App Runner temporarily doubles provisioned instances so old and new code have the same capacity, which briefly doubles that memory cost and counts against the Fargate quota.

Set max concurrency from measurement, not the default. A Python service running two workers with eight threads each serves at most sixteen requests at once; with max concurrency at 100, App Runner will pile 100 requests on an instance that queues 84 of them, latency climbs, and scaling never triggers because the threshold is never crossed. Load-test one instance, find the concurrency where p99 latency starts rising, and set max concurrency a little below it.

# Create a named auto scaling configuration (revision 1), then attach it
aws apprunner create-auto-scaling-configuration \
  --auto-scaling-configuration-name api-steady \
  --max-concurrency 40 --min-size 2 --max-size 10

aws apprunner update-service \
  --service-arn "$SERVICE_ARN" \
  --auto-scaling-configuration-arn "$ASC_ARN" \
  --health-configuration Protocol=HTTP,Path=/healthz,Interval=5,Timeout=2,HealthyThreshold=1,UnhealthyThreshold=3

Configurations are named, versioned resources: up to 10 names per account and Region with up to 5 revisions each, shareable across services. Changing a service's configuration or revision redeploys the service, so treat it like a release.

Advertisement

Deployments and health checks

A deployment starts new instances with the new image, waits for them to pass health checks, shifts traffic, and then drains the old instances. If the new instances never become healthy, the deployment fails and the service keeps running the previous version. That makes the health check the most important setting for safe releases.

The default health check is TCP: App Runner only checks that something accepts connections on the port, every 5 seconds with a 2-second timeout, one success to be healthy and five failures to be unhealthy. A process that binds the port and then fails to reach its database passes a TCP check. Switch to HTTP with a path such as /healthz that checks the process can serve, but keep it cheap and do not make it fail when a non-critical dependency is down, or a brief database blip will mark every instance unhealthy together. All health-check values accept 1 to 20.

Automatic deployments on image push are convenient but couple a registry write to a production release. For production services, many teams disable automatic deployment and call start-deployment from the pipeline after tests pass, with an immutable image tag per release instead of latest.

Networking: inbound and outbound are separate

Inbound traffic always reaches your code through App Runner's endpoint. By default that endpoint is public. For internal services you can make it private, reachable only through an interface VPC endpoint in your VPC, so the service has no public address. Custom domains are associated with the service; App Runner obtains and renews their certificates through ACM once you add the validation records it gives you to DNS.

Outbound traffic is configured separately. By default your code reaches the internet and public AWS endpoints directly. To reach private resources such as an RDS database or ElastiCache cluster, you attach a VPC connector: a set of private subnets, ideally in three Availability Zones, and security groups. App Runner creates Hyperplane network interfaces in those subnets, shared by all services that use the same subnet and security group combination, so IP consumption stays low as you scale. The first service to use a new combination waits two to five minutes during start-up while the interfaces are created.

The trap: once a VPC connector is attached, all outbound traffic from your code goes through the VPC. Calls to third-party APIs, and to AWS APIs such as S3 or Secrets Manager from your code, fail unless the subnets route through a NAT gateway or you add VPC endpoints for those services. App Runner's own traffic, such as pulling the image, shipping logs and fetching the secrets you configured, does not use your VPC and keeps working, which is why a service can start cleanly and then time out on its first outbound call. Subnets must be private and all use the same IP address type. For VPC design itself, see Amazon VPC in depth.

Configuration, secrets and identity

Image services take port, start command, environment variables and secrets in the service configuration. Source services can put the same settings, plus build commands, in an apprunner.yaml at the repository root; the file does not apply to image services. The runtime version pin locks the major version by default; pin major and minor to avoid surprise minor upgrades on the next deployment.

# apprunner.yaml (source-code services only; image services configure this in the service)
version: 1.0
runtime: python3
build:
  commands:
    build:
      - pip install -r requirements.txt
run:
  runtime-version: 3.11          # locks major.minor; patches still update
  command: gunicorn app:app --bind 0.0.0.0:8080 --workers 2 --threads 8
  network:
    port: 8080
  env:
    - name: LOG_LEVEL
      value: info
  secrets:
    - name: DATABASE_URL
      value-from: "arn:aws:secretsmanager:eu-west-1:123456789012:secret:prod/db-url"

Secrets referenced by ARN from Secrets Manager or Parameter Store become environment variables, but App Runner pulls them only during a deployment, so a rotated value reaches the service only after you redeploy. If you rotate database credentials, either redeploy after rotation or have the application fetch the secret itself; see Secrets Manager rotation. Two IAM roles are involved: an access role that lets App Runner pull from private ECR, and an instance role your code assumes to call AWS services. Grant the instance role only what the code calls; IAM in depth covers scoping.

Observability

App Runner sends two log streams to CloudWatch Logs: service logs (deployments, health-check results, scaling events) and application logs (your container's stdout and stderr). Write structured JSON to stdout and include a request ID. Service metrics include request count, latency, HTTP status classes, active instances and CPU and memory utilisation. Tracing can be enabled through an observability configuration that sends traces to AWS X-Ray, as long as your code is instrumented, typically with the OpenTelemetry SDK. Alarms worth setting first: 5xx rate, p99 latency, active instances at max size (you are capped) and failed deployments.

Failure modes

FailureSymptomFix
Max concurrency above real capacitylatency rises under load but instance count stays flatload-test one instance; set max concurrency just below its knee
TCP health check on a broken appdeployment succeeds, requests failHTTP health check on a cheap readiness path
Health check depends on the databasedatabase blip marks every instance unhealthycheck only what the process needs to serve
VPC connector without NAT or endpointsthird-party and AWS API calls time outNAT gateway in the connector subnets or VPC endpoints
Public subnets on the connectorcreate or update fails, or rolls backuse private subnets only
Fargate vCPU quota reachedcannot scale out, or deployments fail at double capacityraise the quota; account for the deployment doubling
Rotated secret not picked upauthentication failures after rotationredeploy after rotation or fetch secrets in code
App listens on the wrong port or 127.0.0.1health checks fail, deployment rolls backbind 0.0.0.0 on the configured port

Migrating to ECS Express Mode

No new features means App Runner will not gain capabilities you may need later, so plan a migration even if nothing is urgent. AWS's recommended target, ECS Express Mode, keeps the one-call experience: you supply a container image and two IAM roles, and ECS provisions a Fargate service, an Application Load Balancer with health checks, auto scaling, security groups and a default URL in your account. There is no charge for Express Mode itself; you pay for the resources it creates. Unlike App Runner, those resources are ordinary ECS and ELB resources you can later tune directly; ECS on Fargate explains what you inherit.

Start by recording the App Runner service's image, port, environment variables, secrets, health check, instance size, scaling settings, networking and custom domain. Source-code services first need a Dockerfile and a pipeline that builds and pushes an image, because Express Mode deploys images only. Then create the ECS service with the same image:

aws ecs create-express-gateway-service \
  --execution-role-arn arn:aws:iam::123456789012:role/ecsTaskExecutionRole \
  --infrastructure-role-arn arn:aws:iam::123456789012:role/ecsInfrastructureRoleForExpressServices \
  --primary-container '{
      "image": "123456789012.dkr.ecr.eu-west-1.amazonaws.com/orders-api:2026-09-30",
      "containerPort": 8080,
      "environment": [{"name": "LOG_LEVEL", "value": "info"}]
  }' \
  --service-name orders-api \
  --health-check-path "/healthz" \
  --scaling-target '{"minTaskCount":2,"maxTaskCount":10}' \
  --monitor-resources

Note the scaling model changes. Express Mode scales tasks with target-tracking policies on metrics such as average CPU, not on per-instance request concurrency, so re-derive the targets from your load test rather than copying numbers. Test the new service on its default URL. If you use a custom domain, add it as a host-header condition on the load balancer's listener rule and attach the ACM certificate to the HTTPS listener, then convert the Route 53 record to weighted routing: App Runner at 100, the ECS load balancer at 0. Shift weights in steps (10, 25, 50, 75, 100), watching error rate and latency at each step and rolling back by restoring weights if needed. Keep App Runner running for a validation period of a day or two after reaching 100, then remove the domain association and delete the service. Services that only use the default App Runner URL have no shared hostname to shift, so clients must be moved to a new endpoint; put a custom domain in front first if you can.

Trade-offs: stay or move

Staying on App Runner is reasonable for stable services that fit its model: it keeps working, receives security and availability investment, and existing customers can still create services. The cost is a frozen feature set and an exit you will eventually take anyway. Moving to Express Mode gives you the ECS feature set (sidecars, task-level tuning, load balancer rules, a broader scaling model) and resources you can inspect and change, at the cost of more moving parts in your account and a load balancer that bills even at low traffic. Teams with many services should move the ones that need new capabilities first and migrate the rest on their normal release cadence. For how ECS is structured underneath, see Amazon ECS in depth.

What to do next

  1. Inventory your App Runner services: source type, image tag, port, instance size, scaling configuration, health check, VPC connector, custom domain.
  2. Load-test one instance per service and set max concurrency below the latency knee; confirm min size matches your burst and AZ needs.
  3. Switch every production service to an HTTP health check on a cheap readiness path.
  4. For services with VPC connectors, verify NAT or VPC endpoints cover every outbound dependency.
  5. Convert source-code services to image builds in CI so they can run anywhere.
  6. Put a custom domain in front of any service still using the default URL, then migrate one low-risk service to ECS Express Mode with weighted DNS and write down what differed.
Key takeaway: App Runner runs containers behind a managed HTTPS endpoint and scales on concurrent requests per instance, charging memory for every provisioned instance and CPU only for active ones. Set max concurrency from a load test, use HTTP health checks, and remember that a VPC connector reroutes all outbound traffic. The service is closed to new customers and frozen for features, so run it well today and plan a gradual, DNS-weighted migration of image-based services to ECS Express Mode.