AWS Graviton is a family of Arm-based server processors that AWS designs and runs in its own EC2 instances. Graviton instances are usually listed at a lower hourly price than the comparable Intel or AMD instance and, for many server workloads, deliver similar or better throughput per vCPU, which is where the price-performance claim comes from. The catch is the instruction set: Graviton runs arm64 code, so every binary, container image, native library and agent on the box must exist for that architecture.

This article explains what changes when you move to Arm, how the Graviton generations map to instance families, how to port a service methodically, how to build images for both architectures from one pipeline, and how to decide with numbers rather than marketing whether a given workload should move. It assumes you know EC2 basics; the underlying virtualisation layer is covered in the Nitro System article.

Generations and instance families

Instance names encode the processor. The letter g after the generation number means Graviton; extra letters add features, such as d for local NVMe instance storage and n for higher network bandwidth. The table lists the generations with the facts AWS has published about each.

ProcessorCore designExample familiesNotes
Graviton (2018)Arm Cortex-A72a1First generation; mostly of historical interest
Graviton2Arm Neoverse N1, 64 coresm6g, c6g, r6g, t4gWhere most production migrations started
Graviton3 / 3EArm Neoverse V1, 64 coresm7g, c7g, r7g, c7gn, hpc7gDDR5 memory, wider vector units (SVE)
Graviton4Arm Neoverse V2, 96 coresm8g, c8g, r8g, x8g12 DDR5-5600 memory channels per chip
Graviton5192 cores per chipm9gAnnounced in preview at re:Invent 2025; check Regional availability

The same processors sit behind many managed services, so you can also adopt Graviton without touching an instance: RDS and ElastiCache offer Graviton node classes, Lambda functions can select the arm64 architecture, and Fargate tasks can request an ARM64 runtime platform. Those are often the easiest first wins because AWS has already done the porting.

To see which Graviton types exist in your Region and what each offers, ask EC2 rather than a blog post, because availability differs by Region and changes over time:

aws ec2 describe-instance-types \
  --filters Name=processor-info.supported-architecture,Values=arm64 \
  --query "InstanceTypes[].[InstanceType,VCpuInfo.DefaultVCpus,MemoryInfo.SizeInMiB]" \
  --output table

# On a running instance, confirm the architecture the OS sees
uname -m        # aarch64 on Graviton, x86_64 on Intel or AMD

AMIs are architecture-specific too. Public AMI parameters in Systems Manager usually come in pairs with x86_64 and arm64 suffixes, and a launch fails if the AMI architecture does not match the instance type.

A vCPU is a core on Graviton

On most current Intel and AMD instances a vCPU is one hardware thread, and two threads share each physical core through simultaneous multithreading. Graviton has no SMT: each vCPU is a whole physical core. A large instance has two vCPUs in both cases, but on Graviton those are two cores with their own execution units and private caches, while on x86 they are usually two threads competing for one core.

This changes how utilisation behaves. A thread-per-core design shows more linear scaling as load rises, because there is no hidden contention between sibling threads; an x86 instance at 50 percent CPU may already be close to its real limit if both hyperthreads of each core are busy. It also means CPU percentage is a poor basis for comparison across architectures. Compare the amount of useful work done at an acceptable latency, which is what the worked example below does.

Two hardware details matter to software. Arm has a weaker memory ordering model than x86, so code with data races that happened to work on x86 can fail on Arm; correct code using proper atomics and locks is unaffected. And Graviton2 onwards supports the Large System Extensions atomic instructions, which scale much better under contention than the older load-exclusive and store-exclusive loops, but only if the compiler emits them.

Auditing a workload before porting

Porting is mostly an inventory exercise. Sort everything that runs on the instance into four groups, because each has a different cost.

GroupExamplesTypical effort
Interpreted or JIT codeJava, Python, Node.js, Ruby, PHP, .NETOften none for your own code; check the runtime version
Native extensions pulled in by packagesPython wheels, Node add-ons, JNI libraries, Ruby gems with C codeFind arm64 builds or compile them; the most common blocker
Compiled servicesGo, Rust, C, C++Rebuild for the new target; review intrinsics and assembly
Third-party binaries and agentsAPM agents, security scanners, log shippers, sidecarsVendor must ship arm64; check before you plan dates

Two free AWS resources speed this up. The AWS Graviton Technical Guide, published on GitHub as aws/aws-graviton-getting-started, collects language-specific advice and known pitfalls. Porting Advisor for Graviton, aws/porting-advisor-for-graviton, scans a source tree and dependency manifests and reports libraries and code patterns that may not work on arm64. Treat its report as a starting list, not a verdict.

Language notes that save time: for Java, newer JDK releases contain years of arm64 tuning, so a port is a good time to leave very old JDKs behind. For Python, pip installs a prebuilt manylinux aarch64 wheel if the package publishes one; if not, it silently falls back to compiling from source, which fails in slim images without a compiler. Go needs only GOARCH=arm64; Rust needs the aarch64-unknown-linux-gnu target. For C and C++, target the core you run on, for example -mcpu=neoverse-v1 for Graviton3, and make sure LSE atomics are used, either by targeting a CPU that has them or with -moutline-atomics, which recent GCC versions enable by default.

Building for two architectures

Porting a service to Graviton: one pipeline, two architecturesSource + depsaudit native codeCI buildbuildx: amd64 + arm64Registryone tag, manifest listTest on arm64unit + load testMixed fleetASG or node group, both archspromotex86 instancesm7i: pulls amd64 layerGraviton instancesm7g / m8g: pulls arm64 layerCompare at equal p99cost per request, not CPU %The runtime picks the matching image from the manifest list; the decision is made on measured cost per unit of work.
One pipeline produces a multi-architecture image. Each host pulls the layer for its own architecture, so a mixed fleet runs the same tag everywhere and the comparison is fair.

Containers make a mixed fleet manageable because a registry can store a manifest list: one tag that points at an amd64 image and an arm64 image. Docker, containerd and Kubernetes pull the variant that matches the node. Build both from one Dockerfile with Buildx.

# One-time: a builder that can target several platforms
docker buildx create --name multi --use

# Build and push both architectures under a single tag
docker buildx build \
  --platform linux/amd64,linux/arm64 \
  -t 123456789012.dkr.ecr.us-east-1.amazonaws.com/api:1.42.0 \
  --push .

# Verify the tag really contains both
docker buildx imagetools inspect \
  123456789012.dkr.ecr.us-east-1.amazonaws.com/api:1.42.0

Buildx can build the foreign architecture under QEMU emulation, which is fine for interpreted code but slow for heavy compilation. For compiled projects, run the arm64 build on an arm64 runner, or cross-compile: Go and Rust cross-compile cleanly, and a multi-stage Dockerfile can use the BUILDPLATFORM and TARGETARCH build arguments to compile on the fast native platform for the target one. Run the test suite on real arm64 hardware, not under emulation, because emulation hides performance problems and some concurrency bugs.

On Kubernetes, label-based scheduling keeps things safe during the transition. Nodes carry the well-known kubernetes.io/arch label, so a workload that is not yet ported can require amd64 with a node selector while ported workloads float across both. On EC2 Auto Scaling, a mixed instances policy whose Arm overrides reference an arm64 launch template does the same for VM-based services; see the Auto Scaling groups guide.

Worked example: cost per request

Suppose an HTTP API runs on ten m7i.large instances in us-east-1 and you want to know whether m7g.large is cheaper for the same service level. At the time of writing, third-party price trackers list Linux on-demand prices of $0.1008 per hour for m7i.large and $0.0816 per hour for m7g.large, both with 2 vCPUs and 8 GiB; confirm current prices on the AWS pricing page before using these numbers.

Quantitym7i.largem7g.large
Price per hour$0.1008$0.0816
Monthly cost per instance (730 h)$73.58$59.57
Ten instances per month$735.84$595.68
Break-even instance ratio1.000.1008 / 0.0816 = 1.235

The last row is the useful one. Graviton is cheaper as long as you need fewer than 1.235 Graviton instances per x86 instance to serve the same load at the same latency. So the question to answer is not whether Graviton is faster but how many requests per second one instance of each type sustains while p99 latency stays inside your objective. Measure it like this:

  1. Deploy the same image tag to one instance of each type behind identical configuration.
  2. Drive both with the same open-loop load generator, stepping the request rate up until p99 latency crosses your objective or errors appear.
  3. Record the highest rate that met the objective. Suppose m7i.large holds 1,150 requests per second and m7g.large 1,050.
  4. Compute cost per million requests: x86 is $0.1008 / (1,150 x 3,600) x 10^6 = $0.0243; Graviton is $0.0816 / (1,050 x 3,600) x 10^6 = $0.0216. Graviton is about 11 percent cheaper per request even though it served fewer requests per instance.
  5. Repeat with production-like data and payload sizes before believing the result.

The example numbers in steps 3 and 4 are illustrative, not a benchmark claim; your workload will produce its own. The method is what transfers. If the Graviton figure had been below 931 requests per second, 1,150 divided by 1.235, x86 would have won at these prices.

Failure modes

Most failed migrations hit one of these, roughly in this order.

  • exec format error at container start. An image was built for amd64 only, or a multi-stage build copied a downloaded x86 binary. Inspect the manifest list and grep the Dockerfile for hard-coded download URLs containing x86_64 or amd64.
  • Builds that suddenly need a compiler. A Python or Node dependency has no arm64 prebuilt artifact, so installation compiles from source and fails or takes many minutes. Pin versions that publish arm64 artifacts or add a build stage with the toolchain.
  • Agents that are x86 only. Monitoring or security agents installed from user data fail quietly, and the instance looks healthy while sending no telemetry. Alert on missing telemetry, not only on errors.
  • Slow code paths. Libraries with hand-written x86 SIMD often have a generic fallback for other architectures that works but is much slower. Profile hot paths on arm64, and prefer library versions that ship Arm-optimised kernels.
  • Concurrency bugs. Latent data races surface under the weaker memory model. Run the test suite with the race detector or thread sanitizer on arm64.
  • Capacity and Spot pools. A new family can have less capacity in some zones. Keep x86 overrides in the same group so scaling still succeeds.

Trade-offs and sequencing

Graviton is a strong default for stateless services, JVM and Go backends, caches and many databases, and the managed-service versions are close to free to adopt. It is a weaker fit when a workload depends on closed-source x86 binaries, on licences priced per core in a way that penalises a thread-per-core design, or on code tuned with x86 intrinsics where the Arm path has not been optimised. A mixed fleet costs something too: two architectures in CI, two sets of performance baselines, and occasional bugs that reproduce on only one of them.

A sensible sequence is to move managed services first, then stateless services that build cleanly, then stateful and native-heavy workloads after you have measured. Reserve capacity and Savings Plans after the move, not before: Compute Savings Plans apply across instance families, but EC2 Instance Savings Plans and Standard Reserved Instances are tied to a family. Storage for the new fleet is covered in the EBS volume types guide.

What to do next

  1. Switch one managed service, such as a Lambda function, an ElastiCache cluster or a non-critical RDS instance, to its Graviton class and watch it for a week.
  2. Run Porting Advisor for Graviton on your largest service and list its native dependencies and agents.
  3. Make your CI produce multi-architecture images for every service, even those not yet moving.
  4. Run the cost-per-request test above for one service and record the break-even ratio.
  5. Add Graviton overrides with their own arm64 launch template to one Auto Scaling group or node group, and shift traffic gradually.
  6. Alert on missing telemetry from new instances so x86-only agents cannot fail silently.
  7. Revisit Savings Plans and reservations once the fleet mix is stable.
Key takeaway: Graviton trades an instruction-set change for lower prices and a thread-per-core design. Inventory native code and agents, build multi-architecture images from one pipeline, test on real arm64 hardware, and decide each workload on measured cost per request at your latency objective rather than on CPU percentage or headline claims.