A cold start happens when Lambda has no idle execution environment for your function and must create one before running your handler. AWS says cold starts typically occur in under 1% of invocations and range from under 100 ms to over a second. That sounds negligible until you look at percentiles: for a user-facing API, the cold 1% is your p99, and a function that imports a heavy library can spend several seconds initializing.
The architecture of a cold start (the microVM, the phases, SnapStart and provisioned concurrency as platform features) is covered in Lambda cold start architecture. This article is the practical companion: how to measure cold starts in your own logs, find out which lines of your code cost the time, remove it, and decide when paying for SnapStart or provisioned concurrency is worth it. A worked example takes a Python API from a slow p99 to an acceptable one.
What you can and cannot change
The Init phase has three parts: Lambda starts the extensions, bootstraps the runtime, and runs your function's static code, meaning everything outside the handler. Before that, the platform has to fetch your code and create the environment. AWS documents that the largest contributor to latency before the handler runs is initialization code, and lists the factors: package size including libraries and layers, the amount of initialization work, and how fast libraries set up connections.
That gives you three levers in order of cost. First, make Function init smaller. Second, make the package smaller so less has to be fetched and loaded. Third, when neither is enough, pay the platform to initialize ahead of time with SnapStart or provisioned concurrency. Most teams jump to the third lever without doing the first, and pay every month for time their own imports waste.
Measuring cold starts from your logs
Every invocation ends with a REPORT line in CloudWatch Logs. On a cold invocation it includes Init Duration; on a warm one the field is absent. CloudWatch Logs Insights exposes it as @initDuration, so two queries answer the first questions: how often do cold starts happen, and how long do they take?
# cold-start rate per 5 minutes
filter @type = "REPORT"
| stats count(*) as invocations,
sum(ispresent(@initDuration)) as cold
by bin(5m)
# cold rate = cold / invocations for each bin
# cold versus warm latency
filter @type = "REPORT"
| fields ispresent(@initDuration) as is_cold, @duration + coalesce(@initDuration, 0) as total_ms
| stats count(*), pct(total_ms, 50), pct(total_ms, 99), max(total_ms) by is_coldLook at both numbers together. A high cold rate with short inits suggests traffic that bursts beyond the warm pool; a low rate with multi-second inits suggests heavy initialization code. Active tracing with X-Ray adds an initialization subsegment to cold invocations, which is useful when you need to line cold starts up against downstream calls.
Two log details prevent misreadings. When initialization fails or times out, Lambda writes an INIT_REPORT line with a status such as timeout or error; functions using SnapStart or provisioned concurrency always emit it. And after an invocation crashes or times out, Lambda resets the environment and re-runs Init inside the next invocation, a suppressed init that is not reported separately but inflates that invocation's duration. If you see isolated slow invocations with no Init Duration right after errors, this is the likely cause.
Finding the expensive lines
Init Duration tells you how long; it does not tell you why. Profile initialization locally with the same runtime version and a similar CPU allocation, because Lambda assigns CPU in proportion to memory and a laptop is much faster than a 256 MB function.
# Python: per-module import cost, largest first
python -X importtime -c "import handler" 2> imports.log
sort -t'|' -k2 -n -r imports.log | head -20
# Node.js: CPU profile of module loading
node --cpu-prof -e "require('./handler.js')"
# Java: class loading volume at startup
java -Xlog:class+load:file=classes.log -cp app.jar com.example.WarmupIn Python the usual offenders are data libraries imported for one code path, a full SDK import where one client would do, and frameworks that scan modules at import time. In Node.js it is unbundled node_modules trees loaded file by file. In Java it is class loading and JIT warm-up for dependency-injection frameworks that scan the classpath and build proxies. Also check what init does over the network: a secrets fetch, a configuration call or a database connection in static code adds a round trip to every cold start, and can be throttled when a burst creates hundreds of environments at once.
Code-level fixes that work in every runtime
- Import only what you use. AWS's own guidance is to require an individual service client instead of the whole SDK. In Node.js with the modular v3 SDK, import
@aws-sdk/client-dynamodbrather than an umbrella package, and bundle with a tree-shaking bundler so only reachable code ships. - Initialize shared clients once, outside the handler. SDK clients, HTTP connection pools and parsed configuration belong in static code so warm invocations reuse them. That does not conflict with the next point: pay once for what every invocation needs.
- Lazily load what only some invocations need. If a PDF renderer serves 2% of requests, import it inside that branch and cache it in a module-level variable. AWS's provisioned-concurrency guidance makes the same point for on-demand functions: defer initialization of capabilities you might not use.
- Keep network calls out of init where you can. Cache secrets with a TTL on first use, or read them through an extension that caches them, instead of calling a remote service unconditionally at import.
- Trim the package. Remove test data, docs and unused native binaries. Layers do not make code free: they are extracted into the same environment and count toward what has to be loaded.
- Test memory settings. CPU scales with memory, so a larger size can shorten init enough to cost less overall. Measure Init Duration at two or three sizes before deciding.
import os, json, boto3
ddb = boto3.client("dynamodb") # needed on every request: pay once
_renderer = None # needed rarely: pay on first use
def _get_renderer():
global _renderer
if _renderer is None:
from reportlib import Renderer # heavy import moved off the cold path
_renderer = Renderer()
return _renderer
def handler(event, context):
if event.get("format") == "pdf":
return _get_renderer().render(event)
item = ddb.get_item(TableName=os.environ["TABLE"], Key={"pk": {"S": event["id"]}})
return {"statusCode": 200, "body": json.dumps(item.get("Item", {}))}
Runtime-specific options
Java. The JVM pays for class loading and interpretation before JIT compilation catches up. Reduce framework reflection and classpath scanning, prefer frameworks with build-time dependency injection, and consider SnapStart, which snapshots the initialized environment when you publish a version. AWS's Java performance guidance also discusses limiting tiered compilation through JAVA_TOOL_OPTIONS for short-lived functions; measure it, because the effect depends on the workload.
Python. Lazy imports and smaller dependency sets are the main tools. SnapStart supports Python 3.12 and later; runtime hooks come from the snapshot_restore_py library included in the managed runtime, with @register_before_snapshot and @register_after_restore decorators.
.NET. SnapStart supports .NET 8 and later. The AWS_LAMBDA_DOTNET_PREJIT environment variable controls ahead-of-time JIT: its default applies it only to provisioned-concurrency environments, Always applies it everywhere, Never disables it.
Node.js. SnapStart is not supported for Node.js runtimes, so bundling, tree-shaking and lazy requires carry the load, with provisioned concurrency for strict latency targets.
SnapStart: fast restores with sharp edges
With SnapStart, Lambda runs Init when you publish a version, snapshots memory and disk, and restores new environments from the snapshot. It works only on published versions and aliases pointing at them, not $LATEST, and it cannot be combined with provisioned concurrency, EFS, or ephemeral storage above 512 MB. Restore and after-restore hooks must complete within 10 seconds or the invocation fails with SnapStartTimeoutException.
The sharp edge is uniqueness. Anything generated in init is shared by every environment restored from the same snapshot: a random seed, a UUID, a cached timestamp, temporary credentials. Generate unique values in the handler or in an after-restore hook. Network connections opened in init may not survive restore; AWS SDK connections usually resume, but validate and re-open others after restore.
from snapshot_restore_py import register_before_snapshot, register_after_restore
import uuid, db
pool = db.connect_pool() # built during Init, captured in the snapshot
instance_id = None
@register_before_snapshot
def close_connections():
pool.close() # do not snapshot live sockets
@register_after_restore
def reopen():
global instance_id
pool.reopen()
instance_id = str(uuid.uuid4()) # unique per restored environmentOn cost, AWS states that SnapStart has no additional charge for Java managed runtimes; for Python and .NET there are caching and restore charges, and duration charges for init and hooks. Delete old published versions you no longer need, because each active SnapStart version keeps a cached snapshot.
Provisioned concurrency: buying warm environments
Provisioned concurrency keeps a configured number of environments initialized on a version or alias. AWS positions it as the option for strict cold-start requirements, with double-digit millisecond response times. It is billed while configured, whether or not it serves traffic, and initialization of those environments is billed too. When traffic exceeds the provisioned amount, the overflow runs on on-demand environments and can cold start again; AWS_LAMBDA_INITIALIZATION_TYPE reads provisioned-concurrency or on-demand inside the environment, so logging it shows which kind served each request.
Size it from measured concurrency, which is roughly requests per second multiplied by average duration in seconds; AWS suggests adding about 10% on top of the typical peak. For predictable daily peaks use scheduled scaling; for variable traffic, target tracking on provisioned-concurrency utilization. Make sure the event source invokes the alias that has provisioned concurrency; pointing API Gateway at $LATEST silently bypasses it.
Worked example: a Python API with a 2-second p99
An API function at 512 MB serves about 20 requests per second with short, spiky bursts. The Logs Insights queries show a cold rate of about 1.5% and a cold p50 total of 2.1 seconds against a warm p50 of 80 ms. The figures in this example are illustrative, but the procedure is the one to follow.
python -X importtimeshows a data-frame library and a PDF library accounting for most of the import time, although only the export endpoint uses them. Moving both imports into that branch removes them from the cold path of every other request.- Init also fetches three secrets synchronously. Replacing that with a cached lookup on first use, or a caching extension, removes three network round trips from Function init.
- Rebuilding the package without test fixtures and unused wheels shrinks it, and a trial at 1,024 MB shows Init Duration falls enough that the larger size costs about the same per request.
- After these changes, re-run the queries. If the cold p99 now meets the target, stop. If it does not, compare SnapStart, since the runtime supports Python 3.12 or later, with a small provisioned-concurrency floor sized to the steady baseline, and choose by monthly cost at your traffic.
Failure modes
- Init timeout. On-demand init is limited to 10 seconds. If it is exceeded, Lambda retries init at the first invocation under the function timeout, which appears as occasional very slow requests. Move work out of init or use SnapStart or provisioned concurrency, where init may run for up to 130 seconds or the function timeout, whichever is higher.
- Burst-time throttling in init. Hundreds of new environments call a secrets or config service at once and are throttled. Cache, add jitter, or use an extension.
- Duplicated randomness after restore. Identical IDs or tokens across environments under SnapStart. Generate them after restore.
- Provisioned concurrency bypassed. The trigger targets
$LATESTor another alias. Check the qualifier on every event source. - Warming pings. A scheduled ping keeps one environment warm but does nothing for a burst that needs fifty. Use real capacity controls instead.
- Misattributed latency. A suppressed init after an error looks like a slow handler. Correlate slow invocations with preceding errors.
Related reading
See Lambda cold start architecture for the platform view, AWS Lambda for the execution model and concurrency, API Gateway for the front door most latency-sensitive functions sit behind, and serverless on AWS for the wider architecture.
What to do next
- Run the cold-rate and cold-versus-warm latency queries for your five most latency-sensitive functions and record the baseline.
- Profile initialization with
python -X importtime,node --cpu-profor JVM class-load logging, and list the three most expensive items. - Move rarely used imports into the branches that need them, keep shared SDK clients global, and remove network calls from unconditional init.
- Shrink the package and test two or three memory sizes against Init Duration and cost.
- Only then evaluate SnapStart (Java 11+, Python 3.12+, .NET 8+) after auditing init for uniqueness and connection handling, or provisioned concurrency sized from measured concurrency plus about 10%.
- Verify every event source targets the alias you configured, log
AWS_LAMBDA_INITIALIZATION_TYPE, and alert onINIT_REPORTtimeouts.