App Engine is Google Cloud's original platform as a service. You upload code with a configuration file, and Google runs it: it provisions instances, balances requests, scales capacity with load, terminates TLS and collects logs. It predates containers as a mainstream deployment unit, and much of its design shows that heritage, including per-language runtimes, a YAML file per service, and URLs of the form version-dot-service-dot-project.
Many production systems still run on it, and it remains a sensible choice for some workloads. It also has behaviours that surprise teams used to Cloud Run or Kubernetes: an application region that cannot be changed, a scaling model that queues requests before adding instances, and instance classes with fixed memory. This article explains the resource model, the two environments, the three scaling types and their parameters, deploys and traffic splitting, and the operational details that decide cost and latency. It ends with a worked sizing example and a comparison with Cloud Run.
The resource model: application, services, versions, instances
A Google Cloud project holds at most one App Engine application, and the application's region is chosen at creation and cannot be changed afterwards. Choose it next to your database. Moving regions means a new project.
An application contains services, formerly called modules. Each service is an independently deployed and independently scaled component with its own app.yaml, runtime, instance class and scaling settings. Every application has a service named default. A service contains versions: each deploy creates an immutable version, and several versions can run side by side. Traffic for a service is assigned to one or more versions by a split. Versions run instances, the processes that serve requests, and scaling settings decide how many exist.
Requests reach a service through its URL, through custom domains, or through dispatch rules that route URL patterns to services. Traffic splitting applies only to requests that do not target a specific version. A request to the version-specific URL goes to that version, which is useful for smoke tests before a version receives user traffic.
Standard versus flexible environment
App Engine has two environments that share the resource model but differ in how code runs.
| Aspect | Standard | Flexible |
|---|---|---|
| Unit | Language runtime with your source | Docker container on Compute Engine VMs |
| Scale to zero | Yes, with automatic or basic scaling | No, at least one instance runs |
| Instance start | Seconds | Minutes, since a VM must boot |
| Sizing | Fixed instance classes (F and B series) | CPU and memory set per version |
| Local disk and background processes | Restricted, ephemeral /tmp | Container filesystem, any process |
| Typical fit | Web apps and APIs with bursty traffic | Custom binaries, long-lived connections |
Flexible was once the answer for anything a standard runtime could not do. Today Cloud Run covers most of that ground with faster starts and scale to zero, so new container workloads rarely start on flexible. The rest of this article focuses on the standard environment.
A minimal service, done properly
Second-generation standard runtimes run ordinary applications: a Flask or Django app, an Express server, a Spring Boot jar. The runtime starts your entrypoint and sends requests to the port in the PORT environment variable. Two details matter for latency. First, import work and client construction happen on every new instance, so keep them lazy or do them in the warmup handler. Second, shutdown is short: App Engine sends SIGTERM and waits up to two seconds before SIGKILL. Let the server process (gunicorn here) handle the signal rather than overriding it in application code, and keep any flush or cleanup inside that window.
# main.py
from flask import Flask
from google.cloud import firestore
app = Flask(__name__)
db = None
def client():
global db
if db is None: # lazy: first request or warmup pays, not import
db = firestore.Client()
return db
@app.get("/_ah/warmup")
def warmup():
client().collection("config").document("flags").get()
return "", 200
@app.get("/items/<item_id>")
def item(item_id):
doc = client().collection("items").document(item_id).get()
return (doc.to_dict(), 200) if doc.exists else ("not found", 404)# app.yaml (standard environment)
runtime: python312
service: default
instance_class: F2
entrypoint: gunicorn -b :$PORT -w 2 --threads 4 main:app
inbound_services:
- warmup
automatic_scaling:
target_cpu_utilization: 0.65
target_throughput_utilization: 0.6
max_concurrent_requests: 8
min_instances: 2 # never scale to zero for this latency-sensitive service
max_instances: 40 # hard cost and downstream-protection ceiling
min_idle_instances: 1
env_variables:
ITEMS_COLLECTION: "items"The gunicorn line matters. App Engine may route up to max_concurrent_requests requests to one instance at once, so the process must be able to serve that many. Two workers with four threads each gives eight, matching the setting. Each worker is a separate process, so a warmup request warms only the worker that receives it. A single synchronous worker would queue requests inside the instance, where the scheduler cannot see them.
Scaling types and what each one controls
Each version uses exactly one of three scaling types, and the choice changes instance classes, request timeouts and startup behaviour.
Automatic scaling is the default and the only type with warmup requests. The scheduler tracks CPU, throughput and pending requests, and creates instances before queues become visible. target_cpu_utilization (default 0.6) and target_throughput_utilization (default 0.6, range 0.5 to 0.95) set the utilisation at which new instances are started. The throughput target is measured against max_concurrent_requests, which defaults to 10. min_instances and max_instances bound the count. min_idle_instances keeps spare instances warm for spikes, at the cost of paying for them. min_pending_latency and max_pending_latency control how long a request may wait in the pending queue before a new instance is started. Requests time out after 10 minutes. Instance classes are F1 (the default, 384 MB, 600 MHz), F2 (768 MB, 1.2 GHz), F4 (1,536 MB, 2.4 GHz) and F4_1G (3,072 MB, 2.4 GHz). The CPU figures describe relative compute share, not clock speed.
Basic scaling creates an instance only when a request arrives and stops it after idle_timeout with no traffic, up to max_instances. Because it reacts only after requests arrive, bursts queue. Manual scaling runs a fixed number of instances continuously. Both use B-series classes: B1 (384 MB), B2 (the default, 768 MB), B4 (1,536 MB), B4_1G (3,072 MB) and B8 (3,072 MB, 4.8 GHz). Both allow requests to run for 24 hours, which makes them the fit for long batch work behind a queue. Manual-scaling instances receive GET /_ah/start immediately when they start. Basic-scaling instances receive it after their first request. A response of 200 to 299, or 404, counts as a successful start.
Cold starts and warmup
A new instance has to start the runtime, import your code and build clients before it can serve a request. When that happens because a user request arrived, that user waits for all of it. With inbound_services: warmup, automatic scaling first sends GET /_ah/warmup to new instances before giving them user traffic. That covers scale-out ahead of demand, but not a scale from zero, where the first request has nowhere else to go. Basic and manual instances never receive warmup requests.
Three levers reduce cold-start pain. min_instances avoids scaling to zero. min_idle_instances keeps headroom for bursts. Keeping start-up light helps too: fewer imports, lazy clients, and no network calls at import time. Measure start-up by the loading-request latency in request logs, not by average latency, which hides it.
Deploys, traffic splitting and routing
Deploy each release as a new version without promoting it, check it through its version URL, then shift traffic gradually. Splitting by IP hashes the client address to a value from 0 to 999. That is simple, but it is imprecise for mobile clients whose addresses change, and for internal callers. Splitting by cookie uses the GOOGAPPUID cookie, which App Engine sets if it is absent, and gives much finer control. Clients must send the cookie back, and services calling each other must forward it.
# Deploy without moving traffic, smoke test the version URL, then canary by cookie
gcloud app deploy app.yaml --version v42 --no-promote
curl -fsS https://v42-dot-default-dot-PROJECT_ID.REGION_ID.r.appspot.com/items/demo
gcloud app services set-traffic default --splits v41=0.9,v42=0.1 --split-by cookie
# ...watch error rate and latency per version in Cloud Monitoring...
gcloud app services set-traffic default --splits v42=1
# Old versions keep costing money if they have min_instances; delete them once rollback is no longer needed
gcloud app versions delete v41 --service defaultDuring a split, two versions serve at once, so both must tolerate each other's data, including schema changes and cache formats. Set Cache-Control on dynamic responses so that a CDN or browser does not serve one version's content to a user pinned to the other. Routing across services uses dispatch.yaml, which is deployed separately from any service:
# dispatch.yaml: route by URL pattern to services
dispatch:
- url: "*/api/*"
service: api
- url: "*/admin/*"
service: adminScheduled work belongs in cron.yaml, which calls a URL on a schedule. Work that must not happen inside a user request belongs in a queue, and Cloud Tasks can target App Engine handlers directly. Point queue targets at a basic or manual service when tasks run longer than automatic scaling's 10-minute limit.
Identity, networking and data access
By default, code runs as the App Engine default service account, PROJECT_ID@appspot.gserviceaccount.com. Historically that account was granted broad project roles, so give each service a dedicated, least-privilege account through the service_account field in app.yaml, as described in Google Cloud IAM. Ingress settings restrict who can reach the application: all traffic, internal only, or internal plus traffic through a Cloud Load Balancer. Combine the last option with a load balancer and Identity-Aware Proxy for internal tools. To reach private IP resources such as a Cloud SQL private instance or Memorystore, attach a Serverless VPC Access connector with vpc_access_connector.
For data, prefer client libraries against managed services such as Firestore, Cloud SQL and Cloud Storage. The first-generation bundled services (the legacy Datastore API, Memcache, Task Queues, the Users API) are available to second-generation runtimes only through app_engine_apis, or the newer app_engine_bundled_services field. Treat them as a migration aid rather than a foundation. Each one ties the code to App Engine.
Worked example: sizing a catalogue service
A catalogue service receives 300 requests per second at peak. Mean latency is 120 ms, most of it waiting on Firestore. By Little's law, about 300 × 0.12 = 36 requests are in flight. With max_concurrent_requests set to 8 and target_throughput_utilization at 0.6, the scheduler aims for about 4.8 concurrent requests per instance. That means about 8 instances at peak. Each request uses about 15 ms of CPU, so an instance at 4.8 in flight handles about 40 requests per second and 0.6 CPU-seconds per second. That is too much for an F1's 600 MHz share, so F2 is the smallest safe class. Memory is about 250 MB per process with two gunicorn workers, which fits F2's 768 MB. F1's 384 MB would not fit.
Overnight, traffic falls to 5 requests per second. min_instances: 2 keeps two warm instances, so there are no cold starts, and that floor sets the minimum bill. max_instances: 40 caps both cost and the load that a runaway retry loop could push into Firestore. If the traffic were mostly idle with rare bursts, min_instances: 0 plus warmup would be cheaper, at the price of a slow first request.
Failure modes
- Memory exceeded on a small class: the instance is terminated after serving a request; check logs for memory-limit messages and move up a class.
- Concurrency mismatch: max_concurrent_requests above what the process serves, so requests queue invisibly inside instances.
- Automatic scaling for long jobs: work beyond 10 minutes is killed; move it to basic or manual scaling, or to a queue.
- Idle old versions: versions with min_instances keep running after traffic moves away.
- Split without compatible data: the canary writes a format the old version cannot read.
- Region regret: the application region cannot change, so latency to a database in another region is permanent.
App Engine or Cloud Run
For new services, Cloud Run is usually the better default. It runs any container, starts quickly, scales to zero, and lets one project deploy services in several regions. Its concurrency model is covered in Cloud Run concurrency. App Engine standard still makes sense for existing applications that run well, for teams that value source deploys with no Dockerfile, and where its versioned traffic model already fits the release process. Migration is mostly mechanical for second-generation runtimes: the same entrypoint runs in a container, scaling settings map to Cloud Run's minimum and maximum instances and concurrency, and dispatch rules become load-balancer URL maps. The hard part is code that depends on bundled services, which must be replaced with standalone client libraries first.
What to do next
- List every service and version with its scaling type, instance class and minimum instances, and stop versions that no longer receive traffic.
- Check that the process concurrency (workers times threads) matches max_concurrent_requests for each service.
- Enable warmup, move client construction into lazy or warmup code, and measure loading-request latency.
- Give each service a dedicated least-privilege service account, and set ingress to the narrowest option that works.
- Adopt no-promote deploys with a cookie-based canary and a documented rollback command.
- Inventory bundled-service usage; if it is zero, prototype the service on Cloud Run and compare cost and latency.