First, the name. On 3 December 2024 AWS renamed the machine-learning service to Amazon SageMaker AI and reused the name Amazon SageMaker for a broader data and AI platform that includes SageMaker Unified Studio. The APIs did not change: the service is still sagemaker in boto3 and the CLI, and training jobs, models and endpoints work as before. This article is about that ML service, which the queue still calls SageMaker.

SageMaker AI is easiest to understand as a scheduler for containers you supply. You describe a job or an endpoint through an API call; AWS launches ML instances in a service-managed account, pulls your image, mounts or copies your data according to a fixed filesystem contract, runs your code under an IAM role you pass, and moves results back to S3. Studio notebooks, the Python SDK, Pipelines and JumpStart are all layers that end in those same calls. Learn the contract and the rest is wrappers. A cross-cloud comparison of this model lives in cloud-native ML compared; here we go one level down.

Advertisement

The resource model

Five resources carry nearly all production work. A training job runs a container to completion and writes a model artifact. A model binds an inference image to an artifact and a role. An endpoint configuration lists production variants (model, instance type, count). An endpoint serves one configuration at a time and can be moved to a new one. A model package, inside a model package group, is a versioned, approvable entry in the registry. Processing jobs and batch transform jobs are siblings of the training job with different contracts.

Every create call takes a RoleArn that SageMaker assumes to read data, pull images, write artifacts and logs. The caller needs iam:PassRole on that role, which is the control that stops any developer from launching a job with an administrator role. Scope one role per workload to its own S3 prefixes and ECR repository; the policy mechanics are in AWS IAM.

SageMaker AI: your containers, AWS-managed instances, S3 at both endsYour codeboto3 / SDK / pipelineSageMaker APICreateTrainingJobPassRoleTraining instancesyour image from ECRS3 / FSx / EFSchannels indatamodel.tar.gzfrom /opt/ml/modelCheckpointsS3, syncedspotModel registrypackage + approvalModel + configCreateModel, EndpointConfigapprovedReal-time endpointinstances, autoscalingServerless endpointspiky, small payloadsAsync endpointqueue, scale to 0Batch transformS3 in, S3 outEvery box is an API resource with an IAM role, a VPC choice and CloudWatch logs and metrics.
Training writes an artifact to S3; the registry versions it; a model and endpoint configuration deploy it to one of four inference options.

The training container contract

Your training image needs no SageMaker library. SageMaker runs the image with the argument train (or your configured entrypoint) and arranges this filesystem:

/opt/ml/
  input/
    config/hyperparameters.json   # HyperParameters from the API, values are strings
    config/resourceconfig.json    # current_host, hosts: who you are in a cluster
    data/<channel>/               # one directory per channel (File and FastFile modes)
  model/                          # write the final model here; uploaded as model.tar.gz
  output/failure                  # write a reason here before exiting non-zero
  output/data/                    # other outputs, uploaded alongside the model
  checkpoints/                    # default CheckpointConfig.LocalPath, synced to S3

On exit, everything in /opt/ml/model is packed into model.tar.gz under the output path, and the job status follows the exit code. Hyperparameters arrive as strings, at most 100 of them; parse them yourself. Anything written to stdout goes to CloudWatch Logs, and regexes in MetricDefinitions turn matching log lines into CloudWatch metrics. AWS's framework containers add a toolkit that exposes the same information as environment variables such as SM_MODEL_DIR and SM_CHANNEL_TRAIN; a bring-your-own image should read the files, which are always there.

Advertisement

Launching a job with boto3

The SDK has changed shape several times (version 3 replaced Estimator, Model and Predictor with ModelTrainer and ModelBuilder), but the API underneath is stable, so production code that must survive SDK upgrades often calls it directly:

import boto3, time
sm = boto3.client("sagemaker")

sm.create_training_job(
    TrainingJobName=f"ticket-clf-{int(time.time())}",
    RoleArn="arn:aws:iam::111122223333:role/sm-train-ticket-clf",
    AlgorithmSpecification={
        "TrainingImage": "111122223333.dkr.ecr.eu-west-1.amazonaws.com/ticket-clf:1.4.0",
        "TrainingInputMode": "FastFile",
        "MetricDefinitions": [{"Name": "val_f1", "Regex": "val_f1=([0-9.]+)"}],
    },
    HyperParameters={"epochs": "4", "lr": "3e-5"},          # strings, always
    InputDataConfig=[
        {"ChannelName": "train", "DataSource": {"S3DataSource": {
            "S3DataType": "S3Prefix", "S3Uri": "s3://ml-data/tickets/train/",
            "S3DataDistributionType": "FullyReplicated"}}},
        {"ChannelName": "val", "DataSource": {"S3DataSource": {
            "S3DataType": "S3Prefix", "S3Uri": "s3://ml-data/tickets/val/",
            "S3DataDistributionType": "FullyReplicated"}}},
    ],
    OutputDataConfig={"S3OutputPath": "s3://ml-artifacts/ticket-clf/",
                      "KmsKeyId": "alias/ml-artifacts"},
    ResourceConfig={"InstanceType": "ml.g5.2xlarge", "InstanceCount": 1,
                    "VolumeSizeInGB": 100},
    EnableManagedSpotTraining=True,
    CheckpointConfig={"S3Uri": "s3://ml-artifacts/ticket-clf/checkpoints/"},
    StoppingCondition={"MaxRuntimeInSeconds": 4 * 3600,
                       "MaxWaitTimeInSeconds": 10 * 3600},   # must be >= MaxRuntime
    VpcConfig={"SecurityGroupIds": ["sg-0abc"], "Subnets": ["subnet-0a", "subnet-0b"]},
    EnableNetworkIsolation=True,
)

A job accepts up to 20 input channels. MaxRuntimeInSeconds defaults to one day and can be raised to 28 days; when it is reached, or when you stop the job, the container receives SIGTERM and has 120 seconds before it is killed. Network isolation blocks all outbound calls from the container except peer traffic in a distributed job, which shuts the door on exfiltration from a compromised dependency; SageMaker still moves data through the VPC you specify.

Input modes and where data lives

ModeWhat happensUse when
FileChannel copied to the instance volume before your code startsDataset fits on disk and is read many times
FastFileS3 prefix exposed as a POSIX filesystem, streamed on readLarge datasets read sequentially; skips the download wait
PipeData streamed through a Unix pipeLegacy record streaming; most new code uses FastFile
FSx for Lustre / EFSFilesystem mounted in your VPCMany epochs over huge data, many jobs sharing one dataset

File mode is simple but its download time is billed and scales with dataset size; a 2 TB channel can spend a long time copying before the first step. FastFile avoids that wait but rewards large sequential files and punishes millions of tiny ones, so pack small samples into shards. For multi-node jobs, S3DataDistributionType set to ShardedByS3Key gives each instance a different subset of objects instead of a full copy. For repeated high-throughput reads, a Lustre filesystem linked to S3 is the usual answer; see FSx for Lustre.

Managed spot training and checkpoints

Setting EnableManagedSpotTraining runs the job on spare capacity, which AWS says can cut training cost by up to 80 percent. Spot capacity can be reclaimed. SageMaker then waits for capacity and restarts the job, and MaxWaitTimeInSeconds (which must be at least MaxRuntimeInSeconds) bounds waiting plus running. A restart only helps if your code resumes: SageMaker syncs CheckpointConfig.LocalPath (default /opt/ml/checkpoints/) to the S3 URI you give, and restores it on restart.

import json, os, signal, sys, glob, torch

CKPT = "/opt/ml/checkpoints"
hp = json.load(open("/opt/ml/input/config/hyperparameters.json"))
epochs, lr = int(hp["epochs"]), float(hp["lr"])
stop = False
signal.signal(signal.SIGTERM, lambda *_: globals().update(stop=True))  # StopTrainingJob / MaxRuntime: 120 s grace

def latest():
    files = sorted(glob.glob(f"{CKPT}/epoch-*.pt"))
    return torch.load(files[-1]) if files else None

def main():
    model = build_model()
    opt = torch.optim.AdamW(model.parameters(), lr=lr)
    start = 0
    if (s := latest()):                       # resumed after interruption
        model.load_state_dict(s["model"]); opt.load_state_dict(s["opt"]); start = s["epoch"] + 1
    for epoch in range(start, epochs):
        train_one_epoch(model, opt, "/opt/ml/input/data/train", should_stop=lambda: stop)
        f1 = evaluate(model, "/opt/ml/input/data/val")
        print(f"val_f1={f1:.4f}", flush=True)  # matched by MetricDefinitions
        torch.save({"model": model.state_dict(), "opt": opt.state_dict(), "epoch": epoch},
                   f"{CKPT}/epoch-{epoch:03d}.pt")
        if stop:
            sys.exit(1)                       # SageMaker resumes from the checkpoint on spot
    torch.save(model.state_dict(), "/opt/ml/model/model.pt")

try:
    main()
except Exception as e:
    open("/opt/ml/output/failure", "w").write(repr(e)[:1000])
    raise

Three rules make this reliable. Save model and optimizer state and the epoch or step, not just weights, or the resumed run silently changes its learning-rate schedule. Checkpoint at an interval you can afford to lose: hourly on a long run, every epoch on a short one. And do not depend on a signal: the 120-second SIGTERM window is documented for stops and the runtime limit, not as a guarantee for spot reclamation, so treat the periodic checkpoint as the only one you can count on and anything written on the signal as a bonus.

Four ways to serve

The inference image contract mirrors training: the model artifact is extracted under /opt/ml/model, and the container must answer GET /ping and POST /invocations on port 8080. What differs is the capacity model around it.

OptionDocumented limitsScaling and costFits
Real-time25 MB payload; 60 s response, 8 min streamingInstances always on; Application Auto ScalingSteady interactive traffic, GPUs
Serverless4 MB payload; 60 sPer request; cold startsSpiky low-volume CPU models
Asynchronous1 GB payload; up to 1 hourInternal queue; can scale to 0Long or large requests, tolerant clients
Batch transformDatasets of GBs; runs for daysInstances for the job onlyScoring a whole table offline

Real-time endpoints scale through Application Auto Scaling, usually on the predefined metric SageMakerVariantInvocationsPerInstance. Load-test one instance to find the invocations per instance it sustains at your latency objective and target about 70 percent of that. GPU endpoints take minutes to add capacity, so set the minimum for your morning ramp, not your overnight trough.

Safe deployment and the registry

Updating an endpoint to a new configuration can use deployment guardrails: a blue/green update that shifts all traffic at once, a canary slice first, or linear steps, with CloudWatch alarms that roll back automatically if they fire during the bake period.

sm.create_model(ModelName="ticket-clf-1-4-0", ExecutionRoleArn=SERVE_ROLE,
    PrimaryContainer={"Image": SERVE_IMAGE,                    # listens on 8080: GET /ping, POST /invocations
                      "ModelDataUrl": "s3://ml-artifacts/ticket-clf/JOB/output/model.tar.gz"})
sm.create_endpoint_config(EndpointConfigName="ticket-clf-1-4-0",
    ProductionVariants=[{"VariantName": "main", "ModelName": "ticket-clf-1-4-0",
                         "InstanceType": "ml.g5.xlarge", "InitialInstanceCount": 2}])
sm.update_endpoint(EndpointName="ticket-clf", EndpointConfigName="ticket-clf-1-4-0",
    DeploymentConfig={
        "BlueGreenUpdatePolicy": {
            "TrafficRoutingConfiguration": {"Type": "CANARY", "WaitIntervalInSeconds": 600,
                "CanarySize": {"Type": "CAPACITY_PERCENT", "Value": 10}},
            "TerminationWaitInSeconds": 300},
        "AutoRollbackConfiguration": {"Alarms": [{"AlarmName": "ticket-clf-5xx"},
                                                 {"AlarmName": "ticket-clf-p99"}]}})

The registry gives the artifact a lifecycle. A training pipeline registers each candidate as a model package with metrics attached and status PendingManualApproval; a reviewer or an automated gate sets Approved; an EventBridge rule on that change triggers the deployment above. That chain is how you answer, months later, which data and code produced the model serving traffic now.

Worked example: a support-ticket classifier

A team fine-tunes a transformer classifier on 4 million labelled tickets, about 12 GB of sharded JSON lines in S3, then serves 40 requests per second at a p99 under 300 ms. Training on one GPU instance takes about 3 hours. They use FastFile so training starts within minutes, spot with a 10-hour wait bound, and a checkpoint per epoch. One run is reclaimed after 2 hours; it resumes from the epoch-1 checkpoint and finishes 4 hours later, still inside the bound. Billable time is only the time instances actually ran.

Load tests show one ml.g5.xlarge instance sustains about 30 requests per second within the latency objective, so they target 21 invocations per second per instance, run a minimum of two instances across Availability Zones and a maximum of six. New versions go out as a 10 percent canary for ten minutes with 5xx and p99 alarms attached. A nightly re-score of all open tickets runs as batch transform rather than through the endpoint, so it never competes with interactive traffic.

Failure modes

  • Job fails with no useful message. Write the reason to /opt/ml/output/failure; it becomes the job's failure reason instead of a bare exit code.
  • Spot job restarts from zero. Checkpoints written outside the local path, or a resume path never exercised. Test it by stopping a job deliberately.
  • Hours of File-mode download. Switch to FastFile, shard the data, or mount FSx.
  • Endpoint health check fails on deploy. The container loads a large model before answering /ping; load lazily or raise the startup health-check timeout.
  • Accelerator mismatch. Inferentia and Trainium instances need models compiled with the Neuron SDK, a different toolchain from CUDA; see the Neuron SDK.
  • Idle endpoints. Real-time endpoints bill while idle. Tag them with an owner and alarm on zero invocations for a week.

What to do next

  1. Build your training image to the /opt/ml contract and run it locally with the same directories mounted.
  2. Create one IAM role per workload, scoped to its S3 prefixes, ECR repository and KMS key, and grant PassRole only for it.
  3. Launch the first job with boto3, FastFile input, MetricDefinitions and a failure file.
  4. Turn on managed spot with checkpoints, then stop a job on purpose and confirm it resumes.
  5. Pick the inference option from the limits table, load-test one instance and set autoscaling from the result.
  6. Register models in a package group and deploy with a canary and rollback alarms.
  7. Keep training data in S3 with lifecycle rules and versioning, as in S3 architecture.
Key takeaway: SageMaker AI runs your containers on managed instances under a role you pass, with a fixed filesystem contract for data in, model out and checkpoints. Use the API directly to stay stable across SDK changes, choose File, FastFile or a mounted filesystem by dataset shape, pair spot with checkpoints you have tested, pick real-time, serverless, asynchronous or batch serving from their documented limits, and deploy through the registry with canaries and alarm-driven rollback.