A SageMaker real-time endpoint is a managed HTTPS front door in front of a fleet of instances, each running your model container. You call it with InvokeEndpoint, it routes the request to an instance, and the container answers. That sounds simple, and the first deployment usually is. The trouble comes later: endpoints that fail to start because a model took nine minutes to load, scaling that reacts after the latency target is already blown, retries that turn a model bug into a traffic storm, and GPUs paid for around the clock to serve a model used twice an hour.
This article is about running endpoints well. It covers the resource model, the request path, the container lifecycle with its exact timeouts, capacity sizing from first principles, scaling on concurrency, inference components and scale to zero, variants and data capture, and a client error policy. If you want the overview of all four SageMaker inference options and model training, start with Amazon SageMaker in depth; this page goes deeper on the real-time option.
Models, configurations and variants
Three resources make up an endpoint. A model names a container image, the S3 location of artifacts and an execution role. An endpoint configuration lists one or more production variants, each binding a model to an instance type, an initial instance count and a traffic weight, plus optional settings such as data capture. The endpoint is the running thing with a stable name that clients call.
Endpoint configurations are immutable. To change anything, such as the model version, instance type or data capture, you create a new configuration and call UpdateEndpoint; SageMaker brings up new capacity, shifts traffic and retires the old, so the endpoint name never changes for clients. Only weights and instance counts can be changed in place, with UpdateEndpointWeightsAndCapacities. Name configurations after the model version, keep old ones until the next release is stable, and roll back by updating to the previous configuration.
The request path
Calls are signed with AWS Signature Version 4, so access is controlled entirely by IAM: grant sagemaker:InvokeEndpoint on the endpoint's ARN to the calling role and nothing else. Endpoints are not public; to serve browsers or partners, put an API layer in front. The runtime chooses a variant by weight, unless the caller sets TargetVariant, then the router picks an instance. The variant's routing configuration sets the strategy: LEAST_OUTSTANDING_REQUESTS sends work to the instance with the most spare capacity, RANDOM picks uniformly, and PREFIX_AWARE keeps requests with the same prompt prefix on the same instance, which helps language models reuse cached prefixes. For models whose request cost varies widely, least outstanding requests usually beats random noticeably at the tail.
One documented inconsistency is worth knowing: the InvokeEndpoint API reference caps the request body at 6,291,456 bytes, while the inference options guide states 25 MB for real-time payloads. Design for the smaller number, and send large inputs as S3 references that the container fetches.
The container lifecycle contract
SageMaker runs your image as docker run IMAGE serve, downloads and extracts the artifacts to read-only /opt/ml/model, and expects a web server on port 8080 answering GET /ping and POST /invocations. The numbers that govern its life are these:
| Moment | Documented rule | What it means for you |
|---|---|---|
| Startup | Must pass /ping consistently within 8 minutes, or the launch fails | Large models need ContainerStartupHealthCheckTimeoutInSeconds, 60-3600 |
| Artifact download | ModelDataDownloadTimeoutInSeconds, 60-3600 | Raise it for multi-gigabyte artifacts |
| Connection | Accept sockets within 250 ms | Never block the accept loop with model work |
| Request | Respond within 60 seconds | Set client socket timeout above 60 s, about 70 s |
| Health check | /ping times out after 2 seconds | Check readiness cheaply; no inference in /ping |
| Shutdown | SIGTERM, then SIGKILL 30 seconds later | Drain in-flight work on SIGTERM; use exec-form ENTRYPOINT |
A static 200 from /ping is the most common container bug. SageMaker uses ping as its main health signal and replaces instances that fail it, except on endpoints that use inference components. If ping always succeeds, an instance whose model failed to load, or whose GPU ran out of memory, keeps receiving traffic and returning errors. Make ping report real readiness:
import signal, threading
from fastapi import FastAPI, Request, Response
app = FastAPI()
state = {"model": None, "healthy": False, "draining": False}
def load():
state["model"] = load_model("/opt/ml/model") # may take minutes
state["model"].predict(WARMUP_INPUT) # compile kernels before traffic
state["healthy"] = True
threading.Thread(target=load, daemon=True).start() # keep the server accepting sockets
@app.get("/ping")
def ping():
ok = state["healthy"] and not state["draining"]
return Response(status_code=200 if ok else 503)
@app.post("/invocations")
async def invocations(request: Request):
body = await request.body()
return Response(content=state["model"].predict_bytes(body),
media_type="application/json")
signal.signal(signal.SIGTERM, lambda *_: state.update(draining=True))Run it with a server whose worker count matches the hardware, for example one worker per GPU or a few per CPU core set, and with an exec-form ENTRYPOINT so signals reach the process.
Sizing from first principles
Sizing starts from Little's law: the number of requests in flight equals arrival rate times time in the system. Measure two things on one instance under a load test: the latency at your objective, and the concurrency at which latency starts to climb, the instance's useful concurrency. Suppose a text classifier on one GPU instance handles 8 concurrent requests at 120 ms p50 and 250 ms p99 before latency degrades. Peak traffic is 300 requests per second. In-flight work at peak is 300 times 0.12, about 36 requests, so you need 36 divided by 8, rounded up, which is 5 instances at full use. Add headroom for imbalance and for losing one Availability Zone; SageMaker spreads instances across zones when there are at least two, so 7 is a reasonable floor at peak.
Overnight traffic of 20 requests per second needs only 2.4 requests in flight, one instance, but keep two for availability. That spread, 2 to 7, is the scaling range. Write those numbers down; they are what you will check autoscaling against.
Scaling on concurrency
Scaling is done by Application Auto Scaling, with the variant as scalable target. The classic metric, SageMakerVariantInvocationsPerInstance, is emitted once a minute and counts requests rather than work, so it reacts late and misleads when request cost varies. The high-resolution metric SageMakerVariantConcurrentRequestsPerModelHighResolution tracks requests in flight inside the container, including queued ones and streaming responses until their last token, and is emitted every 10 seconds, so scale-out reacts much faster. Scale-in proceeds at the normal pace. Target it at the useful concurrency you measured, minus a margin:
import boto3
aas = boto3.client("application-autoscaling")
rid = "endpoint/ticket-clf/variant/main"
aas.register_scalable_target(
ServiceNamespace="sagemaker", ResourceId=rid,
ScalableDimension="sagemaker:variant:DesiredInstanceCount",
MinCapacity=2, MaxCapacity=8)
aas.put_scaling_policy(
PolicyName="concurrency-target", ServiceNamespace="sagemaker", ResourceId=rid,
ScalableDimension="sagemaker:variant:DesiredInstanceCount",
PolicyType="TargetTrackingScaling",
TargetTrackingScalingPolicyConfiguration={
"TargetValue": 6.0, # measured useful concurrency 8, minus margin
"PredefinedMetricSpecification": {
"PredefinedMetricType":
"SageMakerVariantConcurrentRequestsPerModelHighResolution"},
"ScaleInCooldown": 600,
"ScaleOutCooldown": 60})Remember what scaling cannot fix: a new GPU instance must provision, download artifacts, start the container and pass ping before it takes traffic, often several minutes. Set the minimum for your predictable ramp, schedule capacity ahead of known peaks, and alarm on queueing before latency. For metrics, alarms and dashboards, see CloudWatch in depth.
Inference components and scale to zero
Inference components separate models from instances. You create an endpoint whose variant has managed instance scaling, then deploy one or more inference components, each a model with explicit resource requirements and a copy count. SageMaker packs copies onto instances, and callers name the component with InferenceComponentName. This lets several models share a GPU fleet, and with managed instance scaling's MinInstanceCount set to 0 and component copies allowed to scale to zero, an idle endpoint can release all instances.
sm = boto3.client("sagemaker")
sm.create_endpoint_config(
EndpointConfigName="shared-gpu-v1",
ExecutionRoleArn=ROLE,
ProductionVariants=[{
"VariantName": "main", "InstanceType": "ml.g5.2xlarge", "InitialInstanceCount": 1,
"ManagedInstanceScaling": {"Status": "ENABLED", "MinInstanceCount": 0, "MaxInstanceCount": 4},
"RoutingConfig": {"RoutingStrategy": "LEAST_OUTSTANDING_REQUESTS"}}])
sm.create_inference_component(
InferenceComponentName="summarizer-v3", EndpointName="shared-gpu", VariantName="main",
Specification={"ModelName": "summarizer-v3",
"ComputeResourceRequirements": {"NumberOfAcceleratorDevicesRequired": 1,
"MinMemoryRequiredInMb": 16384}},
RuntimeConfig={"CopyCount": 1})The trade is latency for cost. A request arriving when no copy exists waits for an instance and a model load, minutes rather than milliseconds, so scale to zero suits internal tools and batch-like traffic, not user-facing paths. Scale components on SageMakerInferenceComponentConcurrentRequestsPerCopyHighResolution. And remember the ping caveat: on component endpoints a failing ping does not trigger automatic instance replacement, so alarm on component errors yourself.
Variants and data capture
Multiple variants on one endpoint give you A/B tests: weights split traffic, TargetVariant pins a request, and the response header reports which variant served it, so log it next to business outcomes. Data capture records sampled requests and responses to S3 for monitoring and for building evaluation sets:
DataCaptureConfig={
"EnableCapture": True,
"InitialSamplingPercentage": 10,
"DestinationS3Uri": "s3://ml-capture/ticket-clf/",
"KmsKeyId": CAPTURE_KEY_ARN,
"CaptureOptions": [{"CaptureMode": "Input"}, {"CaptureMode": "Output"}]}Captured payloads are customer data: encrypt them with a key you control, restrict the bucket, set lifecycle expiry, and pass InferenceId from the caller so captured records join to your application logs. For new versions, prefer deployment guardrails with canary or linear traffic shifting and alarm-based rollback; the canary example is in the SageMaker overview linked above.
Client errors and retries
Clients decide whether a bad minute stays small. InvokeEndpoint errors fall into groups that need different handling:
| Error | HTTP | Meaning | Retry? |
|---|---|---|---|
| ValidationError | 400 | Bad request, missing endpoint or variant | No; fix the caller |
| ModelError | 424 | Your container returned 4xx or 5xx; OriginalStatusCode and LogStreamArn are included | Only if the original code is transient, such as 503 during load |
| ModelNotReadyException | 429 | Serverless capacity or a multi-model target still loading | Yes, with backoff |
| InternalFailure | 500 | Service-side failure | Yes, with backoff and jitter |
| ServiceUnavailable | 503 | Service temporarily unavailable | Yes, with backoff and jitter |
from botocore.config import Config
rt = boto3.client("sagemaker-runtime",
config=Config(read_timeout=70, connect_timeout=5,
retries={"mode": "adaptive", "max_attempts": 3}))
try:
r = rt.invoke_endpoint(EndpointName="ticket-clf", ContentType="application/json",
Body=payload, InferenceId=request_id)
except rt.exceptions.ModelError as e:
log.error("model %s: %s", e.response.get("OriginalStatusCode"), e.response.get("LogStreamArn"))
raise # a model bug retried three times is three times the loadCap total retries across the call chain, not just per client, and add a circuit breaker so a broken model version fails fast instead of tripling its own traffic.
Failure modes
- Endpoint stuck creating, then failed. The model loaded slower than the startup window. Raise the startup and download timeouts and load in a background thread.
- Errors with healthy-looking instances. Ping returns a static 200. Report real readiness.
- Latency spikes before scaling. Invocation-count scaling is minute-grained. Use the high-resolution concurrency metric and a higher minimum.
- Timeouts at exactly 60 seconds. Long requests belong in asynchronous inference or streaming responses.
- Retry storms. ModelError retried blindly. Classify errors as above.
- Bill for idle GPUs. Low-traffic models on dedicated endpoints. Consolidate with inference components or scale to zero where latency allows.
- IAM too broad. Callers with
sagemaker:*can update endpoints. Grant invoke only; see AWS IAM.
What to do next
- Load-test one instance per model to find useful concurrency and latency at your objective, and write down the scaling range.
- Replace any static ping with a readiness check, load models in the background, and set startup and download timeouts to measured load time plus margin.
- Switch target tracking to the high-resolution concurrency metric, with cooldowns and a minimum that covers your morning ramp.
- Set client read timeouts above 60 seconds, classify errors, and stop retrying ModelError by default.
- Enable data capture at a low sampling rate with KMS encryption, lifecycle expiry and an InferenceId from every caller.
- Find models below a few requests per minute and move them onto a shared inference-component endpoint.
- Rehearse a rollback by updating the endpoint back to the previous configuration in staging.