Many workloads are not services. A nightly rendering run, a genomics pipeline, a Monte Carlo risk calculation or a backfill over a year of data is a pile of finite tasks that should start, use as much compute as is economical, finish and release the machines. AWS Batch provides the queue, the scheduler and the autoscaled compute and charges nothing extra beyond the EC2, Fargate or EKS resources the jobs consume.
Batch is easy to start and easy to misuse. Jobs sit in RUNNABLE for hours because no instance can ever fit them; Spot interruptions burn through retries meant for real failures; a thousand-way array job fails because one child did. This article explains the model from first principles, builds a four-stage pipeline in boto3, and then covers the operational side: retries that distinguish infrastructure failures from bugs, fair sharing between teams, the stuck-RUNNABLE diagnosis, and when Batch is the wrong tool.
The model: five objects
Compute environment. A pool of capacity with a type, a vCPU range and allowed instance types. In a managed environment Batch launches and terminates instances itself; the types are EC2 (on-demand), SPOT, FARGATE, FARGATE_SPOT and, more recently, ECS_MANAGED_INSTANCES, and Batch can also run on an EKS cluster you own. Setting minvCpus to zero lets the environment scale to nothing between runs.
Job queue. Where jobs wait. A queue has a priority and an ordered list of compute environments; the scheduler places jobs from higher-priority queues first when queues share environments. Scheduling policy. An optional fair-share policy attached to a queue that divides capacity between share identifiers instead of running strictly first in, first out. Job definition. A versioned template: container image, command, resource requirements, IAM job role, retry strategy and timeout. Job. One submission of a definition to a queue, with optional overrides, parameters, dependencies and array size.
A job moves through SUBMITTED (accepted), PENDING (waiting on dependencies), RUNNABLE (ready and waiting for capacity), STARTING (placed; image pulling), RUNNING, and finally SUCCEEDED or FAILED. Knowing which state a job is stuck in is most of debugging: PENDING is a dependency problem, RUNNABLE is a capacity or sizing problem, STARTING is an image or network problem.
Compute environments and allocation strategies
The allocation strategy decides which instance types Batch launches when it scales out.
| Strategy | Capacity | Behaviour |
|---|---|---|
| BEST_FIT | On-demand or Spot | Default. Picks the best-fitting, lowest-cost type and waits if it is unavailable; lowest cost, weakest scaling |
| BEST_FIT_PROGRESSIVE | On-demand | Moves on to other suitable types when the preferred ones are unavailable |
| BEST_FIT_PROGRESSIVE_ORDERED | On-demand | Tries types in the order you list them |
| SPOT_CAPACITY_OPTIMIZED | Spot | Prefers pools least likely to be interrupted |
| SPOT_PRICE_CAPACITY_OPTIMIZED | Spot | Balances interruption likelihood and price; the usual Spot choice |
| SPOT_CAPACITY_OPTIMIZED_PRIORITIZED | Spot | Capacity first, your type order honoured where it costs little capacity |
Give Batch room to choose. Listing whole families, several generations and several Availability Zone subnets makes both on-demand scaling and Spot far more reliable than naming one instance size. With the strategies other than BEST_FIT, Batch may exceed maxvCpus by at most one instance, so treat the maximum as approximate when budgeting. Fargate environments take no instance types or allocation strategy at all: you trade control, GPUs and multi-node jobs for having no hosts to manage. All environments attached to one queue must be of compatible kinds, so a queue cannot mix Fargate and EC2 environments.
import boto3
batch = boto3.client("batch")
batch.create_compute_environment(
computeEnvironmentName="spot-ce", type="MANAGED", state="ENABLED",
computeResources={
"type": "SPOT",
"allocationStrategy": "SPOT_PRICE_CAPACITY_OPTIMIZED",
"minvCpus": 0, "maxvCpus": 2048,
"instanceTypes": ["c6i", "c7i", "m6i", "m7i"], # families, not one size
"subnets": ["subnet-aaa", "subnet-bbb", "subnet-ccc"],
"securityGroupIds": ["sg-batch"],
"instanceRole": "ecsInstanceRole",
})
# ...a second, on-demand "ondemand-ce" with type EC2 and BEST_FIT_PROGRESSIVE...
batch.create_job_queue(
jobQueueName="nightly", state="ENABLED", priority=10,
computeEnvironmentOrder=[{"order": 1, "computeEnvironment": "spot-ce"},
{"order": 2, "computeEnvironment": "ondemand-ce"}],
jobStateTimeLimitActions=[{"state": "RUNNABLE", "maxTimeSeconds": 3600,
"action": "CANCEL",
"reason": "stuck RUNNABLE 1h: check CE size and job requirements"}])Environment order gives a fallback, but understand what triggers it: the scheduler moves to the next environment when the earlier one cannot take the job, for example because it has reached its maximum. It is not a guarantee that jobs leave a Spot environment quickly when Spot capacity is scarce. The jobStateTimeLimitActions entry is a safety net: a job at the head of the queue that stays RUNNABLE longer than the limit (600 seconds to 24 hours) is cancelled with your reason, so a misconfiguration fails loudly instead of blocking the queue overnight.
Job definitions and sizing
batch.register_job_definition(
jobDefinitionName="tile-render", type="container",
containerProperties={
"image": "123456789012.dkr.ecr.us-east-1.amazonaws.com/render:1.8.2",
"command": ["python", "worker.py", "Ref::manifest"],
"jobRoleArn": "arn:aws:iam::123456789012:role/render-job",
"resourceRequirements": [{"type": "VCPU", "value": "4"},
{"type": "MEMORY", "value": "7000"}], # MiB, below 8 GiB on purpose
},
retryStrategy={"attempts": 3, "evaluateOnExit": [
{"onStatusReason": "Host EC2*", "action": "RETRY"}, # instance or Spot termination
{"onExitCode": "75", "action": "RETRY"}, # our "transient, try again" code
{"onReason": "*", "action": "EXIT"}, # everything else fails fast
]},
timeout={"attemptDurationSeconds": 3600},
propagateTags=True)Resource requirements are hard reservations used for placement, so they are the most common cause of trouble. Memory is the classic trap: the operating system and the ECS agent keep part of each instance's memory, so a job that asks for exactly 8,192 MiB will not fit on an instance with 8 GiB and waits forever if nothing larger is allowed. Ask for a little less than the instance size, or allow bigger types. Oversized requests waste money through bin-packing, and undersized ones get killed for running out of memory.
Ref::manifest in the command is substituted from the job's parameters map at submission, which keeps one definition reusable. The job role gives the container its own IAM permissions; scope it to the buckets and prefixes the job needs, as described in the IAM article, rather than reusing the instance role. The timeout applies to each attempt, and to each child of an array job separately.
Worked example: a rendering pipeline
A studio renders a nightly batch of 1,000 tiles. A splitter job writes a manifest; an array job renders one tile per child; a thumbnail array job processes each tile as soon as its own render finishes; and a merge job runs when all thumbnails are done.
split = batch.submit_job(jobName="split-0930", jobQueue="nightly",
jobDefinition="splitter", parameters={"manifest": "s3://in/0930/"})
render = batch.submit_job(jobName="render-0930", jobQueue="nightly",
jobDefinition="tile-render",
parameters={"manifest": "s3://work/0930/manifest.json"},
arrayProperties={"size": 1000},
dependsOn=[{"jobId": split["jobId"]}])
thumbs = batch.submit_job(jobName="thumbs-0930", jobQueue="nightly",
jobDefinition="thumbnail", arrayProperties={"size": 1000},
dependsOn=[{"jobId": render["jobId"], "type": "N_TO_N"}])
merge = batch.submit_job(jobName="merge-0930", jobQueue="nightly",
jobDefinition="merger",
dependsOn=[{"jobId": thumbs["jobId"]}]) # waits for all 1000The array job is one submission that expands into 1,000 children, each seeing its index in AWS_BATCH_JOB_ARRAY_INDEX, starting at 0; array sizes run from 2 to 10,000. An N_TO_N dependency links child i of one array to child i of another, so thumbnail 17 starts as soon as render 17 succeeds instead of waiting for the slowest tile. A plain dependency on an array job waits for every child.
import json, os, sys, boto3
s3 = boto3.client("s3")
# exists(), render() and TransientError are your own helpers; they are not shown here.
def main(manifest_uri):
idx = int(os.environ["AWS_BATCH_JOB_ARRAY_INDEX"]) # 0..999
attempt = int(os.environ.get("AWS_BATCH_JOB_ATTEMPT", "1"))
bucket, key = manifest_uri[5:].split("/", 1)
tiles = json.load(s3.get_object(Bucket=bucket, Key=key)["Body"])
tile = tiles[idx]
out_key = f"out/0930/{tile['id']}.png"
if exists(out_key): # idempotent: a retried child skips finished work
print(f"tile {idx} already done (attempt {attempt})"); return 0
try:
data = render(tile)
except TransientError:
return 75 # matches the RETRY rule in the job definition
s3.put_object(Bucket="work", Key=out_key, Body=data) # S3 never exposes a partial object
return 0
if __name__ == "__main__":
sys.exit(main(sys.argv[1]))The worker encodes three rules that make large Batch runs reliable. It is idempotent: it checks whether its output already exists, so a retried child or a rerun of the whole array does not redo finished work. Its output write is atomic: S3 never exposes a partially written object, so a child killed mid-upload leaves nothing rather than a truncated tile; on EFS or FSx get the same guarantee by writing a temporary file and renaming it. And it chooses its exit code: transient failures return 75, which the retry strategy retries, while bugs return anything else and fail fast. Large shared inputs are often better on Amazon EFS or FSx for Lustre than re-downloaded by every child.
Retries and Spot
A retry strategy allows 1 to 10 attempts; each failed attempt puts the job back into RUNNABLE, and the container sees its attempt number in AWS_BATCH_JOB_ATTEMPT, starting at 1. The evaluateOnExit rules match the exit code, the reason or the status reason with patterns that may end in an asterisk, and choose RETRY or EXIT. When an instance is terminated, including by a Spot interruption, the status reason begins with Host EC2, which is why the example retries on Host EC2*.
The rule that catches people out: if none of the rules match, the job is retried. Without the final catch-all EXIT rule, a job with a bug in it runs all of its attempts, multiplying cost and delaying the failure report. Always end the list with an EXIT rule for everything else. Cancelled and terminated jobs are not retried, and neither are jobs whose definition is invalid. For long jobs on Spot, retries alone are not enough: write checkpoints to S3 and resume from the latest one when the attempt number is greater than 1, so an interruption costs minutes rather than hours. The Spot capacity article covers interruption behaviour and diversification in more depth.
Queues, priorities and fair share
Two teams sharing one environment through one first-in, first-out queue means whoever submits 50,000 jobs first owns the cluster until they finish. Separate queues with different priorities help when one workload should always win. When teams should share, attach a fair-share scheduling policy and submit each job with a share identifier. The policy's shareDistribution assigns weights to identifiers, with unlisted identifiers weighing 1.0; shareDecaySeconds sets how far back usage counts, from the minimum of 600 seconds up to a week; and computeReservation holds back capacity for identifiers that are not yet active, reserving (computeReservation/100) raised to the number of active identifiers. With a value of 50 and one active team, half the maximum vCPUs stay free for a second team to arrive; with two active teams, a quarter.
Fair share only reorders jobs that are waiting. It does not preempt running jobs, so very long jobs can still hold capacity; keep individual jobs short, or put long ones in their own environment.
Multi-node parallel jobs
Some work needs several instances cooperating at once: MPI simulations or distributed training. A multi-node parallel job starts a group of nodes together and tells each container its role through AWS_BATCH_JOB_NODE_INDEX, AWS_BATCH_JOB_NUM_NODES, AWS_BATCH_JOB_MAIN_NODE_INDEX and, on the other nodes, AWS_BATCH_JOB_MAIN_NODE_PRIVATE_IPV4_ADDRESS. They are not supported on Spot, and their environments should use a cluster placement group so the nodes share a low-latency network. If one node fails, the job fails, so checkpoint the same way as for Spot.
Failure modes and diagnosis
- Stuck in RUNNABLE. Check in order: does any allowed instance type fit the job's vCPU, memory and GPU request after overhead; is the environment at maxvCpus or disabled; is the account's EC2 quota for those types exhausted; can new instances reach the ECS and ECR endpoints through a NAT gateway or VPC endpoints; does the instance role exist. Instances that launch but never register with ECS almost always point to networking or the instance role.
- Stuck in STARTING. Image pull failures, registry throttling or a very large image. Keep images small and in ECR in the same Region.
- Cascading dependency failure. If a job fails, jobs that depend on it fail too; if any array child fails, the parent fails once all children finish. Design merges to tolerate partial results or rerun only the failed indexes.
- Retry storms. A missing catch-all EXIT rule, or a bug that exits with a retried code, multiplies spend. Alert on attempt counts.
- Silent cost. A nonzero minvCpus keeps instances running all night with an empty queue.
Container output goes to CloudWatch Logs, by default in the /aws/batch/job log group, and every state change is published to EventBridge with the detail type Batch Job State Change. Route FAILED events to a notification, and record per-queue counts of RUNNABLE jobs and their age as your primary health metric.
Batch or something else
Choose Batch for many independent containerized tasks that need more compute than Lambda allows, that benefit from Spot, and that you want to run without operating a scheduler. Choose Step Functions when the hard part is the workflow logic itself, with branching, human approvals or calls to many services; it can submit Batch jobs as steps, and the combination is common. Choose Kubernetes jobs when you already run EKS and want one control plane, possibly with Batch on EKS providing the queueing. Choose Lambda for tasks that finish in seconds to minutes.
What to do next
- Create one Spot and one on-demand environment with minvCpus of 0, several instance families and subnets in several Availability Zones.
- Write retry rules that retry
Host EC2*and a chosen transient exit code, and end with a catch-all EXIT rule. - Make your workers idempotent, write outputs atomically, and checkpoint anything that runs longer than a few minutes.
- Measure peak memory for a sample of jobs and set requests just below instance sizes.
- Add a RUNNABLE time limit action to every queue, and alert on FAILED events and on the age of the oldest RUNNABLE job.
- If more than one team shares capacity, attach a fair-share policy and require a share identifier on submission.