AWS X-Ray is AWS's distributed tracing backend. It receives timing data from every service a request touches, assembles it into a trace, draws a map of which services call which, and lets you search for slow or failing requests. If a checkout takes four seconds, X-Ray is how you find out that 3.2 of them were spent waiting on one DynamoDB query from the inventory service.

The way you feed X-Ray has changed. On February 25, 2026 the X-Ray SDKs and the X-Ray daemon entered maintenance mode, and AWS now ships them only for security issues. AWS recommends OpenTelemetry for instrumentation, through the upstream OTel SDKs or the AWS Distro for OpenTelemetry (ADOT), with the CloudWatch agent or an OpenTelemetry Collector in place of the daemon. This article treats X-Ray as the backend and OpenTelemetry as the way in. It covers the data model, the pipeline, instrumenting a service, sampling (where most of the surprises are), querying, and failure modes. For generic OpenTelemetry concepts, read distributed tracing with OpenTelemetry first.

The data model: traces, segments, annotations

X-Ray's native unit is the segment: a JSON document describing the work one service did for one request, with a name, start and end time, HTTP details, and flags for error (4xx), fault (5xx) and throttle (429). Work inside it, such as a call to DynamoDB or another service, is a subsegment. All segments for one request share a trace ID of the form 1-5759e988-bd862e3fe1be46a994272793: a version, eight hex digits of Unix epoch seconds, and 24 random hex digits. Between services the context travels in the X-Amzn-Trace-Id header, as Root=...;Parent=...;Sampled=1.

OpenTelemetry maps onto this cleanly. A server span becomes a segment and every other span becomes a subsegment. Span attributes become segment metadata by default, which is stored but not searchable. To make an attribute searchable, it must become an annotation: list its key in the aws.xray.annotations span attribute. Annotations are what you filter on later, and they are limited, so choose them on purpose. The quotas that shape design are:

QuotaDefaultAdjustable
Segment document size64 KBNo
Indexed annotations per trace50No
Trace and service graph retention30 daysNo
Segments per second2,600No
Custom sampling rules per Region25Yes
Groups per account25No

The 64 KB limit is the one that bites. A span carrying a full SQL statement or a large response body as attributes can exceed it, and the segment is rejected rather than truncated in a helpful way. Keep attributes small and keep payloads out of spans.

The pipeline: from span to X-Ray

A production pipeline has three tiers. The application creates spans with an OTel SDK. A local collector, either the CloudWatch agent (version 1.300025.0 or later collects OTel traces) or an OTel Collector, receives them over OTLP on ports 4317 (gRPC) and 4318 (HTTP), batches them, and forwards them. The same collector also proxies sampling-rule requests on port 2000, which is how SDKs fetch central sampling rules. The backend receives either X-Ray segments, via the Collector's awsxray exporter calling PutTraceSegments, or OTLP directly at https://xray.<region>.amazonaws.com/v1/traces. That endpoint takes HTTP only (no gRPC), needs SigV4 signing (the Collector's sigv4auth extension), and accepts up to 10,000 spans or 5 MB uncompressed per request.

If you enable Transaction Search in CloudWatch, the spans X-Ray receives are also stored as structured logs in a log group called aws/spans, and a configurable percentage is indexed as X-Ray trace summaries. That lets you keep all spans for log-style analytics while indexing only some for trace search. It moves part of the cost to CloudWatch Logs, so check the CloudWatch pricing page before turning it on for a high-volume service.

Tracing to X-Ray in 2026: OpenTelemetry in the app, X-Ray as the backendService (OTel SDK / ADOT)spans, X-Ray propagatorLambda / API Gatewayactive tracing, ADOT layerCloudWatch agentor OTel CollectorOTLP :4317 / :4318 insampling proxy :2000X-RayPutTraceSegmentsOTLP traces endpointxray.region/v1/tracesTransaction Searchaws/spans log groupSampling rulesreservoir + rate, centralTrace map / searchfilter expressions, groupsOTLPsegmentsOTLP + SigV4if enabledrules/targetsqueryInstrumentation (left) is OpenTelemetry; the X-Ray SDKs and daemon are in maintenance mode.The collector tier is the CloudWatch agent or an OTel Collector; X-Ray (right) stores, indexes and maps.
OpenTelemetry instruments the code, a local agent or collector batches and proxies sampling, and X-Ray stores and indexes.

This is the Collector configuration from AWS's migration guide, the minimum that replaces the daemon:

extensions:
  awsproxy:                 # serves X-Ray sampling rules to SDKs on :2000
    endpoint: 127.0.0.1:2000
  health_check:
receivers:
  otlp:
    protocols:
      grpc: { endpoint: 127.0.0.1:4317 }
      http: { endpoint: 127.0.0.1:4318 }
processors:
  batch:
exporters:
  awsxray:
    region: 'us-east-1'
service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [batch]
      exporters: [awsxray]
  extensions: [awsproxy, health_check]

The credentials the collector runs with need xray:PutTraceSegments, plus xray:GetSamplingRules and xray:GetSamplingTargets for central sampling. On ECS, run the agent as a sidecar and point the application at it with OTEL_EXPORTER_OTLP_TRACES_ENDPOINT=http://cwagent:4318/v1/traces.

Instrumenting a service with OpenTelemetry

Here is a Python service instrumented with the upstream OTel SDK and AWS's extension packages. It uses the X-Ray ID generator, so trace IDs carry the epoch prefix X-Ray's segment format expects, and the X-Ray propagator, so context crosses API Gateway, Lambda and other AWS services that read X-Amzn-Trace-Id. It promotes two attributes to annotations.

# pip install opentelemetry-sdk opentelemetry-exporter-otlp-proto-http \
#             opentelemetry-sdk-extension-aws opentelemetry-propagator-aws-xray
from opentelemetry import trace, propagate
from opentelemetry.sdk.resources import Resource
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.sdk.trace.sampling import ParentBased, TraceIdRatioBased
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.extension.aws.trace import AwsXRayIdGenerator
from opentelemetry.propagators.aws import AwsXRayPropagator

provider = TracerProvider(
    resource=Resource.create({"service.name": "checkout"}),
    id_generator=AwsXRayIdGenerator(),
    sampler=ParentBased(TraceIdRatioBased(0.05)),   # ratio only, no reservoir; central
                                                    # X-Ray rules need ADOT's remote sampler
)
provider.add_span_processor(BatchSpanProcessor(
    OTLPSpanExporter(endpoint="http://localhost:4318/v1/traces")))
trace.set_tracer_provider(provider)
propagate.set_global_textmap(AwsXRayPropagator())

tracer = trace.get_tracer("checkout")

def place_order(order):
    with tracer.start_as_current_span("place_order", kind=trace.SpanKind.SERVER) as span:
        span.set_attribute("tenant.tier", order.tier)        # searchable
        span.set_attribute("order.items", len(order.lines))  # searchable
        span.set_attribute("order.id", order.id)             # metadata only
        span.set_attribute("aws.xray.annotations", ["tenant.tier", "order.items"])
        with tracer.start_as_current_span("reserve_stock"):
            reserve(order)                                   # becomes a subsegment

In real code you would add the OTel instrumentation packages for your web framework, HTTP client and botocore rather than hand-writing spans, or run the ADOT auto-instrumentation agent, which wires the ID generator, propagator and X-Ray remote sampling for you. The annotation keys may be rewritten when stored (dots are not valid in annotation keys), so check how they appear in the console before writing saved queries. For Lambda, AWS recommends its OpenTelemetry Lambda layer. Set OTEL_AWS_APPLICATION_SIGNALS_ENABLED=false if you want tracing only, and see AWS Lambda in depth for the cold-start cost of adding a layer.

Sampling: reservoir, rate and the parent decision

Sampling decides cost and usefulness, and X-Ray's model has two parts per rule: a reservoir, a fixed number of requests per second to trace, and a rate, the percentage of requests traced after the reservoir is used up. The default rule is a reservoir of 1 and a rate of 5 percent. Rules are matched in priority order (1 to 9999) on service name, service type, host, HTTP method, URL path and resource ARN, with * and ? wildcards.

When rules are fetched centrally, X-Ray shares each rule's reservoir across all instances. When rules are local, each instance applies its own reservoir. A worked example shows the difference. Eight instances serve 400 requests per second in total under the default rule. Centrally: 1 + 0.05 × 399 ≈ 21 traces per second, about 1.8 million per day. Locally: each instance traces 1 + 0.05 × 49 ≈ 3.45, so 8 × 3.45 ≈ 27.6 per second, about 2.4 million per day, and the gap grows with instance count. Autoscaling to 40 instances with local rules raises the traced volume even though traffic did not change.

The pitfall AWS warns about most is that sampling is parent-based. The first instrumented service (often API Gateway, a load balancer's first hop, or the edge service) decides, and every downstream service honours that decision. A rule written for an internal service B that is only ever called by A never fires, because A has already decided. Put rules on the root service. Asynchronous workers that start new traces are roots too, and need their own rules. For debugging, add a temporary high-priority rule, for example priority 1, reservoir 1, rate 100 percent for PUT /history/*, then delete it. Rules change without a redeploy. If you need to keep every error regardless of the head decision, that is tail sampling, which the OTel Collector supports and X-Ray rules do not.

Finding traces: filter expressions and groups

Traces are found with filter expressions in the CloudWatch console's trace search. Useful patterns are service("checkout") { fault } for 5xx inside one service, responsetime > 2 for slow requests, and annotation.tenant_tier = "gold" AND responsetime > 1 for your most valuable customers. Annotations are why you choose attributes carefully: they are the only custom fields you can filter traces on. A group saves a filter expression as a named view with its own trace map and CloudWatch metrics. With only 25 per account, use them for long-lived views such as per-tier latency, not ad hoc questions. Pair X-Ray metrics with alarms as described in CloudWatch in depth, and use traces to explain an alarm rather than to raise one.

A worked investigation shows how the pieces fit. The p99 latency alarm on checkout fires. In the trace map, the checkout node's edge to inventory has turned orange while every other edge is normal. You filter with service("checkout") { responsetime > 3 } and open the slowest trace. Its timeline shows the server segment for checkout at 4.1 seconds, a subsegment calling inventory at 3.4 seconds, and inside inventory a DynamoDB Query subsegment at 3.2 seconds, marked as an error with ProvisionedThroughputExceededException and stretched by the SDK's retries. Next you filter on annotation.tenant_tier = "gold" to check whether your largest customers are affected, and they are. The fix belongs to the inventory table's capacity, not to checkout, and you know that within minutes because the context crossed every hop and the right attribute was an annotation.

Failure modes

  • Broken traces. One service uses W3C traceparent only and the next reads only X-Amzn-Trace-Id, so the trace splits in two. Configure both propagators wherever you cross between OTel-native and AWS-native hops.
  • Rejected segments. Oversized attributes push a segment over 64 KB. Watch the collector's exporter error logs and keep payloads out of spans.
  • No data, no error. The collector lacks xray:PutTraceSegments, or the SDK exports to a port nothing listens on. Check the collector's own metrics for spans received and sent.
  • Port 2000 conflict. The old daemon and the new agent both try to bind it. Stop the daemon first, as AWS's migration guide says.
  • Rules that never fire. They target a downstream service; move them to the root.
  • Cost surprises. Local rules multiply with instance count, and Transaction Search adds log ingestion. Review traced volume after every scaling change.

Trade-offs

X-Ray's strengths are integration and zero operations. API Gateway, Lambda, Step Functions and AppSync trace natively, the trace map comes free, and there is nothing to run except a collector. The limits are 30-day retention, 50 annotations, and a query model built around filter expressions rather than arbitrary analytics, although Transaction Search narrows that gap. A self-run backend such as Jaeger or Tempo (see distributed tracing with Jaeger) gives you full control of retention, storage and sampling, at the cost of running it. Because instrumentation is now OpenTelemetry either way, the choice is reversible: change the collector's exporter, not your code.

What to do next

  1. Inventory every service still using an X-Ray SDK or the daemon and plan its move to OTel or ADOT. Those components get security fixes only.
  2. Replace the daemon with the CloudWatch agent or an OTel Collector, keeping port 2000 for sampling and confirming the IAM permissions.
  3. Pick up to a handful of annotations per service (tenant tier, route, region) and list them in aws.xray.annotations.
  4. Move sampling rules to the root services, switch to central rules, and recompute daily traced volume with the reservoir and rate arithmetic above.
  5. Send a test request through every hop and confirm one unbroken trace in the trace map.
  6. Decide whether Transaction Search's all-spans storage is worth its log cost for each high-volume service.
Key takeaway: X-Ray is now best understood as a tracing backend fed by OpenTelemetry: the X-Ray SDKs and daemon have been in maintenance mode since February 2026, and the CloudWatch agent or an OTel Collector replaces the daemon. Choose annotations deliberately, keep segments under 64 KB, put sampling rules on root services and use central rules so volume does not grow with instance count, and verify one unbroken trace end to end.