A SageMaker real-time endpoint is a managed HTTPS front door in front of a fleet of instances, each running your model container. You call it with InvokeEndpoint, it routes the request to an instance, and the container answers. That sounds simple, and the first deployment usually is. The trouble comes later: endpoints that fail to start because a model took nine minutes to load, scaling that reacts after the latency target is already blown, retries that turn a model bug into a traffic storm, and GPUs paid for around the clock to serve a model used twice an hour.

This article is about running endpoints well. It covers the resource model, the request path, the container lifecycle with its exact timeouts, capacity sizing from first principles, scaling on concurrency, inference components and scale to zero, variants and data capture, and a client error policy. If you want the overview of all four SageMaker inference options and model training, start with Amazon SageMaker in depth; this page goes deeper on the real-time option.

Models, configurations and variants

Three resources make up an endpoint. A model names a container image, the S3 location of artifacts and an execution role. An endpoint configuration lists one or more production variants, each binding a model to an instance type, an initial instance count and a traffic weight, plus optional settings such as data capture. The endpoint is the running thing with a stable name that clients call.

Endpoint configurations are immutable. To change anything, such as the model version, instance type or data capture, you create a new configuration and call UpdateEndpoint; SageMaker brings up new capacity, shifts traffic and retires the old, so the endpoint name never changes for clients. Only weights and instance counts can be changed in place, with UpdateEndpointWeightsAndCapacities. Name configurations after the model version, keep old ones until the next release is stable, and roll back by updating to the previous configuration.

The request path

ClientSigV4, InvokeEndpointSageMaker runtimeauth, variant weightsRouterleast outstanding or randomInstance 1container on :8080Instance 2container on :8080/invocations60 s to respond/ping2 s timeout, readinessApplication Auto Scalingmetrics every 60 s or 10 sCloudWatchlatency, errors, logsadd or remove
The request path. The runtime authenticates and chooses a variant by weight; the router picks an instance; the container must accept the connection quickly and answer within 60 seconds. Scaling acts on CloudWatch metrics.

Calls are signed with AWS Signature Version 4, so access is controlled entirely by IAM: grant sagemaker:InvokeEndpoint on the endpoint's ARN to the calling role and nothing else. Endpoints are not public; to serve browsers or partners, put an API layer in front. The runtime chooses a variant by weight, unless the caller sets TargetVariant, then the router picks an instance. The variant's routing configuration sets the strategy: LEAST_OUTSTANDING_REQUESTS sends work to the instance with the most spare capacity, RANDOM picks uniformly, and PREFIX_AWARE keeps requests with the same prompt prefix on the same instance, which helps language models reuse cached prefixes. For models whose request cost varies widely, least outstanding requests usually beats random noticeably at the tail.

One documented inconsistency is worth knowing: the InvokeEndpoint API reference caps the request body at 6,291,456 bytes, while the inference options guide states 25 MB for real-time payloads. Design for the smaller number, and send large inputs as S3 references that the container fetches.

The container lifecycle contract

SageMaker runs your image as docker run IMAGE serve, downloads and extracts the artifacts to read-only /opt/ml/model, and expects a web server on port 8080 answering GET /ping and POST /invocations. The numbers that govern its life are these:

MomentDocumented ruleWhat it means for you
StartupMust pass /ping consistently within 8 minutes, or the launch failsLarge models need ContainerStartupHealthCheckTimeoutInSeconds, 60-3600
Artifact downloadModelDataDownloadTimeoutInSeconds, 60-3600Raise it for multi-gigabyte artifacts
ConnectionAccept sockets within 250 msNever block the accept loop with model work
RequestRespond within 60 secondsSet client socket timeout above 60 s, about 70 s
Health check/ping times out after 2 secondsCheck readiness cheaply; no inference in /ping
ShutdownSIGTERM, then SIGKILL 30 seconds laterDrain in-flight work on SIGTERM; use exec-form ENTRYPOINT

A static 200 from /ping is the most common container bug. SageMaker uses ping as its main health signal and replaces instances that fail it, except on endpoints that use inference components. If ping always succeeds, an instance whose model failed to load, or whose GPU ran out of memory, keeps receiving traffic and returning errors. Make ping report real readiness:

import signal, threading
from fastapi import FastAPI, Request, Response

app = FastAPI()
state = {"model": None, "healthy": False, "draining": False}

def load():
    state["model"] = load_model("/opt/ml/model")      # may take minutes
    state["model"].predict(WARMUP_INPUT)               # compile kernels before traffic
    state["healthy"] = True

threading.Thread(target=load, daemon=True).start()     # keep the server accepting sockets

@app.get("/ping")
def ping():
    ok = state["healthy"] and not state["draining"]
    return Response(status_code=200 if ok else 503)

@app.post("/invocations")
async def invocations(request: Request):
    body = await request.body()
    return Response(content=state["model"].predict_bytes(body),
                    media_type="application/json")

signal.signal(signal.SIGTERM, lambda *_: state.update(draining=True))

Run it with a server whose worker count matches the hardware, for example one worker per GPU or a few per CPU core set, and with an exec-form ENTRYPOINT so signals reach the process.

Sizing from first principles

Sizing starts from Little's law: the number of requests in flight equals arrival rate times time in the system. Measure two things on one instance under a load test: the latency at your objective, and the concurrency at which latency starts to climb, the instance's useful concurrency. Suppose a text classifier on one GPU instance handles 8 concurrent requests at 120 ms p50 and 250 ms p99 before latency degrades. Peak traffic is 300 requests per second. In-flight work at peak is 300 times 0.12, about 36 requests, so you need 36 divided by 8, rounded up, which is 5 instances at full use. Add headroom for imbalance and for losing one Availability Zone; SageMaker spreads instances across zones when there are at least two, so 7 is a reasonable floor at peak.

Overnight traffic of 20 requests per second needs only 2.4 requests in flight, one instance, but keep two for availability. That spread, 2 to 7, is the scaling range. Write those numbers down; they are what you will check autoscaling against.

Scaling on concurrency

Scaling is done by Application Auto Scaling, with the variant as scalable target. The classic metric, SageMakerVariantInvocationsPerInstance, is emitted once a minute and counts requests rather than work, so it reacts late and misleads when request cost varies. The high-resolution metric SageMakerVariantConcurrentRequestsPerModelHighResolution tracks requests in flight inside the container, including queued ones and streaming responses until their last token, and is emitted every 10 seconds, so scale-out reacts much faster. Scale-in proceeds at the normal pace. Target it at the useful concurrency you measured, minus a margin:

import boto3
aas = boto3.client("application-autoscaling")
rid = "endpoint/ticket-clf/variant/main"

aas.register_scalable_target(
    ServiceNamespace="sagemaker", ResourceId=rid,
    ScalableDimension="sagemaker:variant:DesiredInstanceCount",
    MinCapacity=2, MaxCapacity=8)

aas.put_scaling_policy(
    PolicyName="concurrency-target", ServiceNamespace="sagemaker", ResourceId=rid,
    ScalableDimension="sagemaker:variant:DesiredInstanceCount",
    PolicyType="TargetTrackingScaling",
    TargetTrackingScalingPolicyConfiguration={
        "TargetValue": 6.0,   # measured useful concurrency 8, minus margin
        "PredefinedMetricSpecification": {
            "PredefinedMetricType":
                "SageMakerVariantConcurrentRequestsPerModelHighResolution"},
        "ScaleInCooldown": 600,
        "ScaleOutCooldown": 60})

Remember what scaling cannot fix: a new GPU instance must provision, download artifacts, start the container and pass ping before it takes traffic, often several minutes. Set the minimum for your predictable ramp, schedule capacity ahead of known peaks, and alarm on queueing before latency. For metrics, alarms and dashboards, see CloudWatch in depth.

Inference components and scale to zero

Inference components separate models from instances. You create an endpoint whose variant has managed instance scaling, then deploy one or more inference components, each a model with explicit resource requirements and a copy count. SageMaker packs copies onto instances, and callers name the component with InferenceComponentName. This lets several models share a GPU fleet, and with managed instance scaling's MinInstanceCount set to 0 and component copies allowed to scale to zero, an idle endpoint can release all instances.

sm = boto3.client("sagemaker")
sm.create_endpoint_config(
    EndpointConfigName="shared-gpu-v1",
    ExecutionRoleArn=ROLE,
    ProductionVariants=[{
        "VariantName": "main", "InstanceType": "ml.g5.2xlarge", "InitialInstanceCount": 1,
        "ManagedInstanceScaling": {"Status": "ENABLED", "MinInstanceCount": 0, "MaxInstanceCount": 4},
        "RoutingConfig": {"RoutingStrategy": "LEAST_OUTSTANDING_REQUESTS"}}])

sm.create_inference_component(
    InferenceComponentName="summarizer-v3", EndpointName="shared-gpu", VariantName="main",
    Specification={"ModelName": "summarizer-v3",
                   "ComputeResourceRequirements": {"NumberOfAcceleratorDevicesRequired": 1,
                                                   "MinMemoryRequiredInMb": 16384}},
    RuntimeConfig={"CopyCount": 1})

The trade is latency for cost. A request arriving when no copy exists waits for an instance and a model load, minutes rather than milliseconds, so scale to zero suits internal tools and batch-like traffic, not user-facing paths. Scale components on SageMakerInferenceComponentConcurrentRequestsPerCopyHighResolution. And remember the ping caveat: on component endpoints a failing ping does not trigger automatic instance replacement, so alarm on component errors yourself.

Variants and data capture

Multiple variants on one endpoint give you A/B tests: weights split traffic, TargetVariant pins a request, and the response header reports which variant served it, so log it next to business outcomes. Data capture records sampled requests and responses to S3 for monitoring and for building evaluation sets:

DataCaptureConfig={
    "EnableCapture": True,
    "InitialSamplingPercentage": 10,
    "DestinationS3Uri": "s3://ml-capture/ticket-clf/",
    "KmsKeyId": CAPTURE_KEY_ARN,
    "CaptureOptions": [{"CaptureMode": "Input"}, {"CaptureMode": "Output"}]}

Captured payloads are customer data: encrypt them with a key you control, restrict the bucket, set lifecycle expiry, and pass InferenceId from the caller so captured records join to your application logs. For new versions, prefer deployment guardrails with canary or linear traffic shifting and alarm-based rollback; the canary example is in the SageMaker overview linked above.

Client errors and retries

Clients decide whether a bad minute stays small. InvokeEndpoint errors fall into groups that need different handling:

ErrorHTTPMeaningRetry?
ValidationError400Bad request, missing endpoint or variantNo; fix the caller
ModelError424Your container returned 4xx or 5xx; OriginalStatusCode and LogStreamArn are includedOnly if the original code is transient, such as 503 during load
ModelNotReadyException429Serverless capacity or a multi-model target still loadingYes, with backoff
InternalFailure500Service-side failureYes, with backoff and jitter
ServiceUnavailable503Service temporarily unavailableYes, with backoff and jitter
from botocore.config import Config
rt = boto3.client("sagemaker-runtime",
                  config=Config(read_timeout=70, connect_timeout=5,
                                retries={"mode": "adaptive", "max_attempts": 3}))
try:
    r = rt.invoke_endpoint(EndpointName="ticket-clf", ContentType="application/json",
                           Body=payload, InferenceId=request_id)
except rt.exceptions.ModelError as e:
    log.error("model %s: %s", e.response.get("OriginalStatusCode"), e.response.get("LogStreamArn"))
    raise   # a model bug retried three times is three times the load

Cap total retries across the call chain, not just per client, and add a circuit breaker so a broken model version fails fast instead of tripling its own traffic.

Failure modes

  • Endpoint stuck creating, then failed. The model loaded slower than the startup window. Raise the startup and download timeouts and load in a background thread.
  • Errors with healthy-looking instances. Ping returns a static 200. Report real readiness.
  • Latency spikes before scaling. Invocation-count scaling is minute-grained. Use the high-resolution concurrency metric and a higher minimum.
  • Timeouts at exactly 60 seconds. Long requests belong in asynchronous inference or streaming responses.
  • Retry storms. ModelError retried blindly. Classify errors as above.
  • Bill for idle GPUs. Low-traffic models on dedicated endpoints. Consolidate with inference components or scale to zero where latency allows.
  • IAM too broad. Callers with sagemaker:* can update endpoints. Grant invoke only; see AWS IAM.

What to do next

  1. Load-test one instance per model to find useful concurrency and latency at your objective, and write down the scaling range.
  2. Replace any static ping with a readiness check, load models in the background, and set startup and download timeouts to measured load time plus margin.
  3. Switch target tracking to the high-resolution concurrency metric, with cooldowns and a minimum that covers your morning ramp.
  4. Set client read timeouts above 60 seconds, classify errors, and stop retrying ModelError by default.
  5. Enable data capture at a low sampling rate with KMS encryption, lifecycle expiry and an InferenceId from every caller.
  6. Find models below a few requests per minute and move them onto a shared inference-component endpoint.
  7. Rehearse a rollback by updating the endpoint back to the previous configuration in staging.
Key takeaway: A SageMaker endpoint is an immutable configuration of variants behind a stable name, routing requests to containers that must accept sockets within 250 ms, answer within 60 seconds and report real readiness on ping. Size it from measured useful concurrency, scale on the high-resolution concurrency metric, and use inference components to share or release GPUs when latency allows. Give clients timeouts above 60 seconds and retry only transient errors, never blind ModelError retries.