Azure Batch is a managed scheduler for running large numbers of independent command-line tasks on a pool of virtual machines that it creates, scales and repairs for you. You do not run a cluster manager, a job queue or a node bootstrapper; you describe a pool, submit a job, add tasks, and Batch places each task on a free slot, retries it when it fails and reports its exit code. Typical workloads are rendering frames, Monte Carlo runs, genomics pipelines, nightly ETL shards, and offline ML inference or embedding over millions of documents.

This article explains the object model from first principles, then builds a real job with the current Python SDK (azure-batch 15.x, whose classes are named BatchPoolCreateOptions, BatchJobCreateOptions and BatchTaskCreateOptions; older samples use different names). It covers autoscale formulas, Spot capacity, a worked sizing example, failure modes and trade-offs. SDK names were checked against the 15.1.0 package and the autoscale details against Microsoft Learn in October 2026.

The object model

Five objects make up Batch, and each owns a different decision:

ObjectWhat it isWhat you decide
AccountRegional container with quotas (cores, pools, jobs)region, pool allocation mode, identity
PoolA set of identical compute nodes (VMs)VM size, OS image, node count or autoscale formula, task slots per node, start task
JobA group of tasks bound to one poolpriority, constraints, what happens when tasks finish or fail
TaskOne command line run on one nodecommand, inputs, outputs, retries, wall-clock limit, dependencies
NodeOne VM in a poolmostly nothing: Batch creates, repairs and removes nodes

A task is a process, not a container scheduler's pod. Batch downloads its resource_files into the task's working directory, runs command_line (directly, or inside a container if the pool has container support), captures stdout.txt and stderr.txt, uploads any output_files that match, and records the exit code. Exit code 0 is success. Anything else is a failure, which triggers a retry if max_task_retry_count allows one. Each task can use required_slots of a node's task_slots_per_node, which is how several small tasks share one large VM.

Jobs can also carry three special tasks. A job preparation task runs on each node before the first task of the job lands there, which suits downloading a shared model or reference dataset once per node rather than once per task. A job release task runs on each node that ran the job when the job ends, to clean up. A job manager task is a task that itself adds the job's other tasks, so very large or data-dependent fan-outs can be generated inside the pool instead of from your laptop.

A pool can hold dedicated nodes and Spot nodes. Spot nodes use spare Azure capacity at a discount and can be preempted. Batch's older "low-priority" VMs were retired on 30 September 2025 in favour of Spot. The API field and formula variable keep the old name: target_low_priority_nodes and $TargetLowPriorityNodes are documented as the Spot targets.

Architecture and data flow

Azure Batch: you submit tasks, the service schedules them onto a poolYour clientazure-batch SDKBlob Storageinputs / outputsBatch accountjobs, tasks, schedulerAutoscale formulaevery 5 min - 168 hcreate job / tasksPool (VM size, image, start task)Dedicated 14 task slotsSpot 14 task slotsDedicated 24 task slotsSpot 24 task slotsDedicated 34 task slotsSpot 3preemptedTask: download inputs, run commandupload output_files, exit codescheduleresizeoutputsresource_filesPreemption puts a Spot node's running tasks back in the queue (requeue_count), separately from failure retries (retry_count).Nodes cost money while they exist, busy or idle, so the formula and the deallocation option are cost controls.
The client only talks to the Batch account. The account scales the pool, schedules tasks into node slots, and tasks move data through Blob Storage.

The client never talks to nodes. It makes control-plane calls (create pool, add tasks, read task state) and moves data through Blob Storage. That has two consequences. Your submission code can exit as soon as the tasks are added, because Batch holds the queue. And everything a task needs must be reachable from the node: storage containers through a managed identity or a SAS URL, container images through a registry the pool can authenticate to, packages through a start task or a custom image.

Building a job with the Python SDK

The sketch below creates a pool, a job and 2,000 shard tasks for an embedding run. It uses Microsoft Entra authentication via DefaultAzureCredential rather than account keys. Instead of hard-coding an OS image, it asks Batch which images it supports, because the image and the node agent SKU must match.

import datetime as dt
from azure.identity import DefaultAzureCredential
from azure.batch import BatchClient, models
# FORMULA, INPUT_CONTAINER_URL, OUTPUT_CONTAINER_URL and UAMI_ID (a user-assigned
# managed identity attached to the pool) are defined elsewhere.
ident = models.BatchNodeIdentityReference(resource_id=UAMI_ID)

client = BatchClient(endpoint="https://<account>.<region>.batch.azure.com",
                     credential=DefaultAzureCredential())

# Pick a supported Ubuntu 22.04 image and its matching node agent SKU.
img = next(i for i in client.list_supported_images()
           if i.os_type == "linux" and "22.04" in i.node_agent_sku_id)

client.create_pool(pool=models.BatchPoolCreateOptions(
    id="embed-pool",
    vm_size="Standard_D8s_v5",
    virtual_machine_configuration=models.VirtualMachineConfiguration(
        image_reference=img.image_reference,
        node_agent_sku_id=img.node_agent_sku_id),
    task_slots_per_node=4,
    task_scheduling_policy=models.BatchTaskSchedulingPolicy(node_fill_type="pack"),
    enable_auto_scale=True,
    auto_scale_formula=FORMULA,                       # see the next section
    auto_scale_evaluation_interval=dt.timedelta(minutes=5),
    start_task=models.BatchStartTask(
        # setup/ blobs keep their path: they land at ./setup/ in the start task's directory
        command_line="/bin/bash -c 'apt-get update && apt-get install -y python3-pip"
                     " && python3 -m pip install -r setup/requirements.txt"
                     " && cp setup/embed.py $AZ_BATCH_NODE_SHARED_DIR/'",
        resource_files=[models.ResourceFile(storage_container_url=INPUT_CONTAINER_URL,
                                            blob_prefix="setup/", identity_reference=ident)],
        user_identity=models.UserIdentity(auto_user=models.AutoUserSpecification(
            scope="pool", elevation_level="admin")),
        wait_for_success=True, max_task_retry_count=2)))

client.create_job(job=models.BatchJobCreateOptions(
    id="embed-20261003",
    pool_info=models.BatchPoolInfo(pool_id="embed-pool"),
    constraints=models.BatchJobConstraints(max_wall_clock_time=dt.timedelta(hours=12)),
    all_tasks_complete_mode="noaction"))              # an empty job counts as complete

def out_for(n):                                       # one blob folder per shard
    return models.OutputFileDestination(container=models.OutputFileBlobContainerDestination(
        container_url=OUTPUT_CONTAINER_URL, path=f"embeddings/{n:05d}", identity_reference=ident))

tasks = [models.BatchTaskCreateOptions(
            id=f"shard-{n:05d}",                       # deterministic ids make resubmits safe
            # no shell unless you start one: bash expands the env var
            command_line=f"/bin/bash -c 'python3 $AZ_BATCH_NODE_SHARED_DIR/embed.py"
                         f" --in shards/{n:05d} --out out'",
            resource_files=[models.ResourceFile(storage_container_url=INPUT_CONTAINER_URL,
                                                blob_prefix=f"shards/{n:05d}/",
                                                identity_reference=ident)],
            output_files=[models.OutputFile(
                file_pattern="out/*.parquet",
                destination=out_for(n),
                upload_options=models.OutputFileUploadConfiguration(upload_condition="tasksuccess"))],
            constraints=models.BatchTaskConstraints(max_task_retry_count=3,
                                                    max_wall_clock_time=dt.timedelta(minutes=45)))
         for n in range(2000)]

client.create_tasks(job_id="embed-20261003", task_collection=tasks, max_concurrency=4)
client.update_job(job_id="embed-20261003",
                  job=models.BatchJobUpdateOptions(all_tasks_complete_mode="terminatejob"))

Several details matter here. create_tasks splits the list into service-sized requests for you and raises CreateTasksError (carrying pending_tasks, failure_tasks and errors) if some tasks could not be added. Because the IDs are deterministic, resubmitting the failed ones is safe: a duplicate ID is rejected rather than run twice. The job is created with all_tasks_complete_mode="noaction" and switched to "terminatejob" only after the last task is added, because a job with no tasks counts as complete and would end at once. The start task runs as a pool-wide admin auto-user, installs pip and the dependencies system-wide, and copies the script to the node's shared directory, because the start task's working directory is not visible to tasks. Batch does not run command lines in a shell, so environment variables such as $AZ_BATCH_NODE_SHARED_DIR only expand inside /bin/bash -c. Inputs fetched with blob_prefix keep their blob paths under the working directory, and each shard uploads to its own folder so retries and parallel tasks never overwrite each other. With node_fill_type="pack", Batch fills one node's slots before using the next, so idle nodes empty out and the formula can remove them. A "spread" policy keeps every node partly busy.

To watch progress, read client.get_job_task_counts(job_id).task_counts (active, running, completed, succeeded, failed) instead of listing 2,000 tasks in a loop.

Autoscale formulas

An autoscale formula is a small program that Batch evaluates every auto_scale_evaluation_interval (default 15 minutes, minimum 5 minutes, maximum 168 hours). It reads service-defined metrics such as $PendingTasks (running plus queued tasks), $ActiveTasks and $CurrentLowPriorityNodes, and sets $TargetDedicatedNodes, $TargetLowPriorityNodes and $NodeDeallocationOption. Once a pool autoscales, the fixed node counts are ignored.

// Slots-aware: 4 task slots per node, at most 25 nodes, 5 of them dedicated.
slots = 4;
maxNodes = 25;
floorDedicated = 5;
pct = $PendingTasks.GetSamplePercent(TimeInterval_Minute * 5);
pending = pct < 70 ? max(0, $PendingTasks.GetSample(1)) : max($PendingTasks.GetSample(TimeInterval_Minute * 5));
want = min(maxNodes, ceil(pending / slots));
$TargetDedicatedNodes = pending > 0 ? min(want, floorDedicated) : 0;
$TargetLowPriorityNodes = max(0, want - $TargetDedicatedNodes);
$NodeDeallocationOption = taskcompletion;

Three ideas carry the formula. It divides pending tasks by slots per node, because $PendingTasks counts tasks, not nodes. It guards against sparse metrics: when fewer than 70% of the last five minutes' samples are present, it uses the latest sample, and otherwise the peak over the window, so a brief dip does not shrink the pool. And it uses taskcompletion so shrinking waits for running tasks to finish. The alternatives are requeue (the default: kill and reschedule, which is fast but wastes work), terminate (kill and drop) and retaineddata (also wait for retained task data to be cleaned up). To test a formula, enable autoscale on the pool with a trivial one such as $TargetDedicatedNodes = 0, then call client.evaluate_pool_auto_scale with the real formula; it returns the computed values without applying them. Formula syntax errors are the most common reason a pool "never scales".

Worked example: sizing a 2,000-shard run

Take the job above: 2,000 shards, each needing 2 vCPUs for about 6 minutes on a Standard_D8s_v5 (8 vCPUs, so 4 slots per node). That is 2,000 × 6 = 12,000 task-minutes, or 200 task-hours.

QuantityCalculationResult
Concurrent slots at the 25-node cap25 × 4100
Waves of tasks2,000 / 10020
Ideal wall-clock time20 × 6 min120 min
Node-hours billed (ideal)25 nodes × 2 h50 (10 dedicated + 40 Spot)
Formula target at the startmin(25, ceil(2,000 / 4) = 500)25 nodes
Formula target at 60 pending tasksceil(60 / 4)15 nodes

Real runs are slower. Nodes take minutes to allocate and run the start task, the last wave is ragged, and Spot nodes get preempted. Suppose 3 of the 20 Spot nodes are preempted at the 40-minute mark. Their 12 running tasks go back in the queue; Batch tracks this in each task's execution_info.requeue_count, a separate counter from retry_count, so preemption does not use up the failure retries you configured. If Spot capacity does not return, the remaining 22 nodes give 88 slots. The formula's want stays at 25 until 96 or fewer tasks are pending, so Batch keeps asking for Spot nodes it cannot get. That is why a few dedicated nodes are a useful floor: they guarantee the job finishes even with zero Spot capacity. For guidance on mixing the two, see Spot capacity strategies.

Failure modes

  • Start task fails. With wait_for_success=True the node goes to state starttaskfailed and takes no tasks. A broken package mirror can leave a 25-node pool idle while it bills. Alert on nodes in starttaskfailed and unusable, and bake heavy dependencies into a custom image.
  • Quota stops growth. A pool cannot exceed the account's dedicated or Spot core quota for that VM family. The formula's target is then simply not reached, with no task error. Check quotas before a big run.
  • Output upload errors. Uploads use the node's identity or the SAS URL you supply, and a missing role assignment surfaces as a file-upload error after the command exited 0. Judge tasks by execution_info.result and failure_info, not the exit code alone, and use ExitConditions.file_upload_error to decide what the job does next.
  • Non-idempotent tasks. Retries and preemption mean a task can run more than once, possibly after partly writing its output. Write to a temporary name and rename on success, or make writes keyed by shard.
  • Hung tasks. Without max_wall_clock_time a stuck process holds a slot forever. Set it on every task and the job.
  • Forgotten pools. Nodes bill while they exist, busy or not. Autoscale to zero, or delete the pool when the job ends. Tasks not finished within 180 days of being added are terminated by the service.

Trade-offs

Batch suits embarrassingly parallel work measured in hours, where each task is a command line and the main concern is cheap, elastic capacity. It is not an interactive service platform, and it does not orchestrate long-running services; AKS does that, with its own batch tooling if you already run Kubernetes. For model training with experiment tracking, Azure Machine Learning compute (Azure ML Studio) is usually the better fit. Batch is priced as the underlying VMs and storage, so its cost discipline is your formula, your VM choice and your Spot mix; Azure VMs in depth helps choose sizes. The price of the managed scheduler is control: you cannot customise its placement logic beyond slots, fill type, job priority and affinity.

What to do next

  1. Install the current azure-batch package, and use DefaultAzureCredential with a managed identity instead of account keys.
  2. Pick the image and node agent SKU from list_supported_images(), and move heavy setup from the start task into a custom image once it stabilises.
  3. Give every task a deterministic ID, max_wall_clock_time, max_task_retry_count, and output files uploaded on success only.
  4. Make tasks idempotent: write to temporary paths, then rename, and key outputs by shard.
  5. Write a slots-aware autoscale formula with a small dedicated floor, test it with evaluate_pool_auto_scale, and use taskcompletion to avoid wasting work.
  6. Check core quotas for both dedicated and Spot in the VM family before large runs.
  7. Monitor get_job_task_counts, unusable nodes and requeue counts, and make sure every pool can scale to zero or is deleted when the job ends.
Key takeaway: Azure Batch turns a list of command lines into a scheduled, retried, autoscaled run on VMs you never manage directly. Model work as idempotent tasks with deterministic IDs, wall-clock limits and success-only outputs, size pools by task slots, autoscale on pending tasks with a small dedicated floor under Spot capacity, deallocate on task completion, and make sure every pool scales to zero when the work is done.