Amazon Neptune is AWS's managed graph database. It stores data as a property graph, queried with Apache TinkerPop Gremlin or openCypher, or as an RDF graph, queried with SPARQL, and it is built for workloads where the questions are about relationships: which accounts share a device with a known fraudster, what a user's friends bought, how a drug relates to a gene through papers, which network hosts can reach a sensitive one. Those questions are awkward in SQL because each hop is another join; in a graph engine they are a traversal.
This article explains how Neptune is built, how to choose a data model and query language, how clients connect and authenticate, how to bulk load from S3, how serverless capacity works, and what goes wrong in production. A worked fraud-ring example runs through the middle, and the article ends with a checklist. Facts and limits were checked against the Neptune user guide on 2026-10-04; versions and limits change, so confirm them against the docs for your engine version.
Architecture: compute, storage and endpoints
Neptune separates compute from storage, the same design used by Aurora (see Aurora storage). A cluster has one primary instance that performs all writes and up to 15 Neptune replicas that serve reads. They all attach to one cluster volume, which keeps six copies of the data across three Availability Zones and grows to a maximum of 128 TiB. Because replicas read the same volume rather than replaying a separate copy, adding one does not mean copying the database, and if the primary fails Neptune promotes a replica.
Clients reach instances through endpoints. The cluster endpoint always points at the current primary; use it for writes. The reader endpoint spreads connections across replicas; use it for read-only queries. Each instance also has its own endpoint, and you can define custom endpoints over a subset of instances, for example to isolate analytics queries. The default port is 8182.
Neptune Database is not the only engine. Neptune Analytics is a separate in-memory engine for analysing large graphs with graph algorithms and low-latency analytic queries, loaded from a Neptune database or from data in S3. Use Neptune Database for transactional reads and writes; consider Analytics when the job is whole-graph computation such as PageRank or community detection.
Data models and query languages
Pick the model before anything else, because a property graph and an RDF graph in Neptune are different data. A property graph can be queried with both Gremlin and openCypher; an RDF graph is queried with SPARQL.
| Model | Query languages | Good fit |
|---|---|---|
| Property graph | openCypher (declarative), Gremlin (traversal steps) | Application data: users, accounts, devices, transactions, with properties on nodes and edges |
| RDF | SPARQL | Knowledge graphs that use shared vocabularies, ontologies and URIs, or must interoperate with linked-data sources |
Within a property graph, openCypher is usually easier for teams coming from SQL, because it describes the pattern to match; Gremlin is a sequence of traversal steps, which gives precise control over how the traversal runs. Neptune's Gremlin implementation has documented differences from the reference TinkerPop server, so read the compliance page before porting code. Neither model has a fixed schema enforced by the database; constraints such as unique ids are your application's job.
Connecting and authenticating
Neptune is normally deployed inside a VPC with no access from outside it (see VPC design). From engine release 1.4.6.x you can enable public endpoints per instance, but IAM authentication is then required and access is still limited by security groups. Every connection uses HTTPS with TLS 1.2. Neptune has no username and password authentication: when IAM database authentication is enabled, each request is signed with AWS Signature Version 4 under the service name neptune-db, and IAM policies decide who may connect and, from release 1.2.0.0, which query actions they may run (see IAM condition keys).
Queries go over HTTPS endpoints such as /openCypher and /gremlin, or over WebSockets for Gremlin drivers. The AWS SDKs include a neptunedata client that signs requests for you:
import boto3
from botocore.config import Config
client = boto3.client(
"neptunedata",
endpoint_url="https://mygraph.cluster-ro-abc123.eu-west-1.neptune.amazonaws.com:8182",
config=Config(read_timeout=None, retries={"total_max_attempts": 1}),
)
resp = client.execute_open_cypher_query(
openCypherQuery="MATCH (a:Account {acctId: $id})-[:USES]->(d:Device) RETURN d.devId",
parameters='{"id": "acct-1001"}',
)
print(resp["results"])AWS's own examples disable client retries and the client read timeout and rely on Neptune's server-side query timeout instead; the engine-status example in the docs shows clusterQueryTimeoutInMs at 120000. Use the reader endpoint for reads as above, and parameterise queries rather than concatenating strings, both for safety and so the engine can reuse plans.
Worked example: finding a fraud ring
A payments company wants to flag accounts linked to known fraud. Nodes are Account, Device, Card and IpAddress; relationships are USES (account to device), HOLDS (account to card) and LOGGED_IN_FROM (account to IP). When an account is confirmed fraudulent, analysts want every account within two shared identifiers of it.
// openCypher: accounts sharing a device or card with a known-fraud account
MATCH (bad:Account {status: 'fraud'})-[:USES|HOLDS]->(x)<-[:USES|HOLDS]-(other:Account)
WHERE other.status <> 'fraud'
RETURN other.acctId AS account, collect(DISTINCT labels(x)[0]) AS shared, count(*) AS links
ORDER BY links DESC
LIMIT 50// Gremlin: the same pattern as explicit traversal steps
g.V().has('Account', 'status', 'fraud').
out('USES', 'HOLDS').
in('USES', 'HOLDS').hasLabel('Account').
not(has('status', 'fraud')).
groupCount().by('acctId').
order(local).by(values, desc).limit(local, 50)Data comes from S3 in Neptune's openCypher CSV format, one file per node or relationship type. Each node file needs an :ID column and usually a :LABEL column. The id is the node's element id, not a property, so the account file writes its header as acctId:ID, which also stores the value as the acctId property the queries above match; the device file does the same with devId:ID. Relationship files use :ID, :START_ID, :END_ID and :TYPE. Typed property columns follow, such as status:String.
The trap in this data is the shared identifier. A public Wi-Fi IP address or a device reset to factory defaults can be linked to thousands of accounts, and a two-hop query through it returns everyone. That is the supernode problem: exclude IP addresses from the pattern, as above, or filter identifiers whose degree is above a threshold you store as a property and refresh in batch. Bound the result with LIMIT, and test the query against the worst hub in your data, not the average.
Bulk loading from S3
For large imports, do not issue millions of individual write queries. Neptune's bulk loader reads files from an S3 bucket in the same Region. The setup is: put the files in S3, create an IAM role with read and list access to the bucket and associate it with the cluster, create an S3 VPC endpoint, then POST a load request to the cluster.
curl -X POST https://mygraph.cluster-abc123.eu-west-1.neptune.amazonaws.com:8182/loader \
-H 'Content-Type: application/json' \
-d '{
"source": "s3://acme-graph/fraud/2026-10-04/",
"format": "opencypher",
"userProvidedEdgeIds": "TRUE",
"iamRoleArn": "arn:aws:iam::123456789012:role/NeptuneLoadFromS3",
"region": "eu-west-1",
"mode": "AUTO",
"failOnError": "FALSE",
"parallelism": "MEDIUM",
"queueRequest": "TRUE"
}'The response contains a loadId to poll for status and errors. Format values are csv (Gremlin CSV), opencypher, ntriples, nquads, rdfxml and turtle. Mode AUTO resumes a previous load from the same source if one exists, so a fixed and resubmitted job skips files already loaded; providing relationship ids lets it resume relationship files too. Parallelism ranges from LOW to OVERSUBSCRIBE, with HIGH the default; if openCypher loads fail with LOAD_DATA_DEADLOCK, lower it and retry. Up to 64 jobs can queue, and dependencies chains a job behind others. The loader is not ACID: with failOnError set to TRUE it stops at the first error but keeps what it already loaded, so design loads to be re-runnable.
Capacity: provisioned and serverless
Provisioned clusters use memory-optimised instance classes; graph queries are memory-hungry, so size the primary so the frequently traversed part of the graph stays cached. Neptune Serverless uses the instance class db.serverless and scales in Neptune Capacity Units. One NCU is 2 GiB of memory with proportional CPU and networking. You set a minimum of at least 1.0 NCU and a maximum between 2.5 and 128 NCUs in the cluster's scaling configuration.
The docs give useful sizing guidance. A minimum set too low means slower scale-up, because Neptune scales in increments based on current capacity, and it means the buffer cache is evicted when capacity shrinks, so the first queries after a quiet period read from storage. Readers in promotion tiers 0 and 1 scale with the writer so they are ready for failover; readers in tiers 2 to 15 scale independently and can lag under heavy writes if their minimum is low. Set the maximum to cap cost, accepting that peaks beyond it will see higher latency or timeouts.
Failure modes
- Query timeouts. Unbounded traversals through hubs hit the cluster timeout. Add LIMITs, exclude hub labels and use explain output to find the expensive step.
- Concurrent write conflicts. Neptune can reject writes that conflict with another transaction, returning a concurrent modification error; retry these with exponential backoff and keep write transactions small.
- Errors after a 200 response. openCypher results stream in chunks, so a failure partway through arrives after a 200 status. From engine release 1.4.5.0 you can send
TE: trailersto get the final status in anX-Neptune-Statustrailer. - Request too large. Gremlin and SPARQL HTTP requests must be under 150 MB, and a single property value is limited to 55 MB. Batch writes and store large objects in S3 with a pointer in the graph.
- Dropped WebSockets. Idle Gremlin WebSocket connections close after about 20 to 25 minutes, and with IAM auth every WebSocket is disconnected after about 10 days. Drivers must reconnect.
- Stale reads from replicas. Replicas can lag the primary briefly; read-your-own-write flows should use the cluster endpoint.
- Null characters in strings. Neptune rejects them; sanitise input before loading.
Operating it
Start with the engine status endpoint. aws neptunedata get-engine-status --endpoint-url https://<cluster-endpoint>:8182 reports the engine version, the instance role (writer or reader), the Gremlin, SPARQL and openCypher versions, and which features are on: slow query logs, the audit log, Neptune Streams, IAM authentication, plus settings such as the cluster query timeout. Check it after every engine upgrade and parameter change, because some parameter changes only apply after a reboot.
Turn on slow query logging and the audit log in non-trivial environments; the first tells you which traversals to fix, the second who ran them. For serverless clusters, chart the CloudWatch metrics ServerlessDatabaseCapacity and NCUUtilization per instance: an instance that sits near its maximum is being capped, and one that drops to the minimum every night will start the morning with a cold cache. Neptune Streams, a change log of graph mutations, is the usual way to feed changes to a search index or downstream consumers without dual writes.
For loads, remember the housekeeping limits: Neptune tracks only the most recent 1,024 load jobs and keeps the last 10,000 error details per job, so record load ids and copy error summaries into your own logs when a job finishes.
When Neptune fits, and what you trade
Neptune earns its keep when traversals of three or more hops, variable-length paths or pattern matching are core to the product. If your queries are one hop, such as "list this user's followers", a key-value or relational store is usually cheaper, and a recursive CTE in Postgres or Aurora handles occasional hierarchies. Full-text search over graph properties is not Neptune's strength; AWS documents an integration that pairs it with OpenSearch. For retrieval-augmented generation, graphs complement vector search rather than replace it; GraphRAG versus vector search compares them. Operationally, you trade control for a managed service: there is no username and password auth, access is through IAM and the VPC, and you cannot install plugins, in exchange for not running replication, backups or failover yourself.
What to do next
- Write the five queries your product needs and count their hops; confirm that a graph engine is actually warranted.
- Choose property graph or RDF, and for property graphs pick openCypher or Gremlin as the team's default.
- Create a small cluster with IAM database authentication enabled, in private subnets across at least two AZs, and connect with the neptunedata client.
- Export a sample of real data to Neptune's CSV format, include relationship ids, and run the bulk loader with mode AUTO.
- Run each query against your worst hub and add LIMITs, hub exclusions and degree properties until it stays within the timeout.
- Add retry with backoff for conflicting writes, use the reader endpoint for reads, and enable TE trailers for streamed results.
- Decide provisioned or serverless; for serverless, set a minimum NCU that keeps the hot graph cached.