Amazon OpenSearch Service runs OpenSearch, the open-source search and analytics engine that forked from Elasticsearch 7.10, as a managed cluster called a domain. AWS provisions the nodes, replaces failed ones, applies software updates and takes automated snapshots. It does not choose your shard count, design your mappings, size your heap pressure or stop a wildcard query from taking down a node. Most OpenSearch incidents on AWS come from those decisions, not from the managed parts.

This article explains how a domain is built, how to size it with the formulas AWS publishes, how data should flow in and age out, and how to operate it. The numbers below come from the OpenSearch Service developer guide as read on 2026-10-03; where AWS documents a value only for some versions, the article says so. For the engine's internals, read Elasticsearch cluster architecture, which applies almost unchanged to OpenSearch.

How a domain is built

A domain is an OpenSearch cluster with three kinds of node. Dedicated master nodes, called cluster manager nodes in current OpenSearch, hold cluster state: which indices exist, their mappings, and where every shard lives. They do no indexing or searching, which is why they keep working when data nodes are overloaded. Use three, so that losing one still leaves a quorum of two. Data nodes hold shards on EBS volumes or instance storage and do all the indexing and query work. UltraWarm nodes serve read-only indices whose data lives in S3 with a local cache, and cold storage keeps indices in S3 detached from any node until you attach them to query.

Each index is split into primary shards, each with zero or more replicas. A shard is a full Lucene index, so it costs heap, file handles and CPU even when idle. OpenSearch Service gives the JVM heap half of an instance's memory, capped at 32 GiB, and everything else goes to the operating system's page cache, which Lucene relies on for fast reads. That cap is why instances above 64 GiB of memory add cache, not heap, and why scaling out eventually beats scaling up.

Production domains should use Multi-AZ with Standby. AWS's guide describes it as three zones, three dedicated masters, data nodes in a multiple of three and at least two replicas, giving each zone a full copy of the data. One zone's nodes are held as standby and do not serve searches; on an infrastructure failure AWS activates them in under a minute, with no shard reshuffling. AWS states it is available from OpenSearch 1.3, at no extra cost, and lists 99.99 percent availability for it against 99.9 percent without standby. It is also restrictive: a shard may not exceed 65 GB, a node may hold at most 1,000 shards and the cluster 75,000, and only certain instance families qualify.

Producersapps, Kinesis, agentsOpenSearch Ingestionor Lambda / FirehoseDomain in your VPC, three Availability ZonesManager AZ-adedicatedManager AZ-bdedicatedManager AZ-cdedicatedData AZ-ahot, gp3 EBSData AZ-bhot, gp3 EBSData AZ-cstandby copyUltraWarm nodesread-only, S3-backed, cachedCold storageS3, attach to query_bulkClients and DashboardsIAM / SigV4 or FGAC usersISM policyhot (rollover) -> warm_migration -> cold or deleteWrites land on hot nodes; ISM moves aging indices to cheaper tiers; managers hold cluster state.
A three-zone domain with standby: dedicated managers hold cluster state, hot data nodes take writes, and ISM moves aging indices to UltraWarm and cold storage.

Sizing: storage, shards, compute

Size a domain in three passes: storage, shards, then compute. AWS's guide gives a formula for each of the first two.

Storage. Indexes are larger than the source data, replicas multiply it, and the service and operating system reserve space. AWS's rule of thumb is: minimum storage = source data x (1 + replicas) x 1.45. Then keep disk usage under roughly 75 to 80 percent so recovery and merges have room.

Shards. Keep shards between 10 and 30 GiB when search latency matters and between 30 and 50 GiB for write-heavy log analytics. The guide's formula is: (source data + room to grow) x (1 + indexing overhead, about 0.1) / desired shard size = approximate primary shards. Also keep no more than 25 shards per GiB of heap on any node. The per-node shard ceiling depends on the engine version: 1,000 per node up to OpenSearch 2.15, and for 2.17 and later 1,000 per 16 GiB of heap up to 4,000. Set the primary count explicitly in an index template, because changing it on an existing index means a shrink, a split or a reindex.

Compute. Choose memory-optimized instances for heavy aggregations and search, and general-purpose ones for balanced workloads, then load-test with real queries. CPU and heap pressure under a realistic query mix decide the node count far more often than disk does.

Worked example: 200 GiB of logs a day

Worked example. A platform team ingests 200 GiB of application logs per day, queries mostly the last three days, keeps 14 days searchable and must retain 90 days for audits. Using one daily index:

QuantityCalculationResult
Primaries per daily index200 x 1.1 / 40 GiB target5.5, round to 6
Hot storage, 3 days, 2 replicas (standby)200 x 3 x (1 + 2) x 1.452,610 GiB
Warm data, days 4 to 14200 x 11 x 1.1, one copy in S3about 2,420 GiB
Hot shards on the cluster3 days x 6 x 3 copies54 shards

With six hot data nodes, two per zone, each node needs about 435 GiB of index data, so a 600 GiB gp3 volume keeps it near 72 percent full. Fifty-four shards over six nodes is nine per node, far below any limit. Six primaries on six nodes also lets one day's writes use every node. Days 4 to 14 move to UltraWarm, which stores one copy in S3 rather than replicas on EBS, and indices older than 14 days go to cold storage until the 90-day mark. Remember that with standby, one zone's nodes do not serve searches, so the query load test must pass on the active nodes alone.

Notice what the arithmetic rejected. A per-service daily index for 40 services, each with the AWS guide's default of five primaries and one replica, would create 400 shards a day, mostly tiny. That pattern, not data volume, is the most common reason log domains run out of heap.

Getting data in and aging it out

Writes should arrive in bulk. Each _bulk request carries many documents, and a few megabytes per request is a reasonable starting point to tune from. Managed pipelines, either OpenSearch Ingestion or Firehose, handle batching and retries for you; Kinesis in front of them absorbs bursts. Custom writers should sign requests with SigV4 and back off on HTTP 429, which means a node's write queue is full:

import boto3
from opensearchpy import OpenSearch, RequestsHttpConnection, AWSV4SignerAuth, helpers

region = "eu-west-1"
auth = AWSV4SignerAuth(boto3.Session().get_credentials(), region, "es")
client = OpenSearch(
    hosts=[{"host": "vpc-logs-abc123.eu-west-1.es.amazonaws.com", "port": 443}],
    http_auth=auth, use_ssl=True, verify_certs=True,
    connection_class=RequestsHttpConnection, timeout=30,
)

def actions(events):
    for e in events:
        yield {"_index": "logs-write", "_source": e}   # write alias, not a dated name

ok, errors = helpers.bulk(client, actions(read_events()), chunk_size=2000,
                          max_retries=5, initial_backoff=2, raise_on_error=False)

Writing to an alias rather than a dated index name lets Index State Management roll the index over by size or age. An index template fixes shard counts and mappings, and an ISM policy moves indices through the tiers. The warm_migration action is the OpenSearch Service-specific step that moves an index to UltraWarm:

PUT logs-000001
{ "aliases": { "logs-write": { "is_write_index": true } } }

PUT _index_template/logs
{ "index_patterns": ["logs-*"],
  "template": { "settings": { "number_of_shards": 6, "number_of_replicas": 2,
                              "plugins.index_state_management.rollover_alias": "logs-write" },
                "mappings": { "dynamic": "strict",
                              "properties": { "@timestamp": {"type": "date"},
                                              "service": {"type": "keyword"},
                                              "level": {"type": "keyword"},
                                              "message": {"type": "text"} } } } }

PUT _plugins/_ism/policies/logs
{ "policy": { "description": "hot 3d, warm to 14d, delete at 90d",
    "default_state": "hot",
    "states": [
      { "name": "hot", "actions": [{ "rollover": { "min_size": "240gb", "min_index_age": "1d" } }],
        "transitions": [{ "state_name": "warm", "conditions": { "min_index_age": "3d" } }] },
      { "name": "warm", "actions": [{ "warm_migration": {} }],
        "transitions": [{ "state_name": "delete", "conditions": { "min_index_age": "90d" } }] },
      { "name": "delete", "actions": [{ "delete": {} }] } ],
    "ism_template": [{ "index_patterns": ["logs-*"] }] } }

Create the first index, logs-000001, with the write alias before any writer runs. Rollover needs a numeric suffix to increment, and if a writer sends to logs-write first, OpenSearch creates a concrete index with that name and rollover fails. The policy is deliberately simplified: it keeps warm indices until deletion rather than adding the cold step, which uses its own ISM action. min_size is the total size of the index's primaries, so 240 GiB over six primaries targets 40 GiB shards. dynamic: strict rejects unknown fields, which prevents mapping explosions from free-form log payloads; the alternative is to map them as a flattened field.

Security, changes and monitoring

Access control has two layers. The domain's resource policy and its VPC security groups decide who can reach the endpoint. Fine-grained access control, when enabled, adds users, roles, and index-, document- and field-level permissions inside the cluster, and maps IAM roles to OpenSearch roles. Put domains in a VPC, sign requests with IAM, and use fine-grained access control when several teams or tenants share one domain. See AWS VPC for subnet and security-group design.

Configuration changes such as instance type or count, and many version upgrades, are applied as a blue/green deployment: AWS builds new nodes, copies shards across and removes the old ones. That needs spare capacity and takes longer on big domains, so make changes when the cluster is healthy and not during an incident if you can avoid it. Take manual snapshots to your own S3 bucket before risky changes; the automated snapshots belong to the service and are meant for its own recovery.

Watch a short list of CloudWatch metrics: ClusterStatus.red and ClusterStatus.yellow, FreeStorageSpace, JVMMemoryPressure, CPUUtilization, ClusterIndexWritesBlocked, and thread-pool rejection counts. AWS publishes a recommended alarm list with thresholds; start from it rather than inventing your own.

Failure modes

  • Red cluster. At least one primary shard is unassigned, so some data cannot be searched or written. Use GET _cluster/allocation/explain to see why; the usual causes are lost nodes with no replica and full disks.
  • Writes blocked. Low free storage or sustained memory pressure makes the service block writes, which shows as ClusterIndexWritesBlocked and ClusterBlockException. Delete or migrate old indices or add storage before turning off the block.
  • Heap exhaustion from queries. Large terms aggregations on high-cardinality fields, leading wildcards and deep pagination fill the heap and trip circuit breakers. Limit them in the application and use search_after instead of large from offsets.
  • Too many shards. Per-tenant or per-service daily indices create thousands of small shards and slow every cluster state update. Consolidate indices and use rollover.
  • Mapping explosion. Dynamic mapping of arbitrary JSON keys hits the default 1,000-field limit, and the heap cost arrives well before that.
  • Uneven zones. Without standby, data node counts that are not a multiple of three with two replicas overload one node.
  • Bulk rejections. HTTP 429 from oversized or too-parallel bulk writers; retrying immediately makes it worse, so back off.

Trade-offs

Provisioned domains give full control over shards, plugins, tiers and cost, and reward teams that plan capacity. OpenSearch Serverless removes node and shard management and bills in OpenSearch Compute Units, at the cost of fewer knobs and a different feature set, so check that every API your application uses is supported before choosing it. Self-managed OpenSearch on EC2 or Kubernetes gives every setting and every plugin, but you own upgrades, snapshots and recovery.

For pure log retention with rare queries, compare the bill with CloudWatch Logs or Athena over S3, which charge mainly for storage and per query. For vector search, OpenSearch's k-NN features are a strong option when you also need keyword filters; the OpenSearch vector math article covers the memory arithmetic. The choice is about who runs capacity: you, AWS's provisioned service with your sizing, or Serverless.

What to do next

  1. Write down data volume per day, retention per tier, query patterns and latency targets.
  2. Compute storage with source x (1 + replicas) x 1.45 and primaries with the shard formula.
  3. Create a Multi-AZ with Standby domain with three dedicated masters and data nodes in a multiple of three.
  4. Place it in a VPC, sign requests with IAM, and enable fine-grained access control for shared domains.
  5. Add an index template with explicit shard counts and strict or controlled mappings.
  6. Write through an alias and add an ISM policy for rollover, warm migration and deletion.
  7. Load-test with real queries on the active nodes only and watch heap pressure, not just CPU.
  8. Set AWS's recommended CloudWatch alarms and register an S3 repository for manual snapshots.
  9. Review shard count and size every month as volumes change.
Key takeaway: Amazon OpenSearch Service manages nodes, patching and recovery, but shard counts, mappings and query behaviour remain your job, and they cause most incidents. Run production on Multi-AZ with Standby with three dedicated masters and data nodes in multiples of three. Size storage as source x (1 + replicas) x 1.45 and choose primaries for 10 to 30 GiB shards for search or 30 to 50 GiB for logs, staying under 25 shards per GiB of heap. Write in bulk through an alias, let ISM roll indices over and move them to UltraWarm and cold storage, keep mappings controlled, and alarm on cluster status, free storage, JVM pressure and blocked writes.