Every large cloud now sells an AI platform, and every one of them has been renamed at least once in the past two years. Google announced at Cloud Next in April 2026 that Vertex AI's services now ship as the Gemini Enterprise Agent Platform. Microsoft renamed Azure AI Foundry to Microsoft Foundry at Ignite in November 2025. AWS renamed the original SageMaker to SageMaker AI in December 2024, placed it inside a broader next-generation SageMaker for data and analytics, and runs generative models through Bedrock. The names change faster than the architecture underneath, which is why comparing brand pages is a poor way to choose.

This article compares the platforms by what they do rather than what they are called. All three provide the same six layers, drawn below, and differ in how those layers are packaged, priced and governed. Knowing the layers lets you map a workload onto whichever cloud you already run, see where lock-in actually sits, and design the thin seam that keeps you portable. Product details move quickly, so treat names and limits here as a map to check against the current documentation, not as a contract.

Advertisement

The six layers

The six layers every hyperscaler AI platform providesGovernanceidentity, private networking, quotas, content safety, audit, evaluationAgent runtimemanaged agent hosting, memory, tool gateway, observabilityModel catalogue + managed APIsfirst- and third-party models, per tokenPipelines and registryorchestration, lineage, model versionsCustom trainingjobs on GPU or TPU clusters you sizeHostingreal-time, serverless, async, batch endpointsShared infrastructureaccelerators, object storage, VPC, key management, billingGoogle CloudGemini Enterprise Agent Platform(the evolution of Vertex AI)AWSBedrock + Bedrock AgentCoreSageMaker AIMicrosoft AzureMicrosoft FoundryAzure Machine Learning
The layers are the same on every cloud; the product names are not. Most workloads touch only two or three of them.

Model catalogue and managed APIs let you call a foundation model by ID and pay per token, with no infrastructure. Custom training runs your code on accelerator clusters you size. Hosting puts a model you own behind an endpoint. Pipelines and registry tie data preparation, training, evaluation and deployment into a repeatable graph with versioned artefacts. Agent runtimes, the newest layer, host long-running agents with memory, tool access and tracing. Governance wraps all of it: identity, private networking, quotas, safety filters, evaluation and audit.

A retrieval chatbot uses the catalogue, an agent runtime and governance. A fraud model trained on tabular data uses training, hosting and pipelines and never touches the catalogue. Deciding which layers a workload needs is the first step, because it narrows the comparison to the products that matter.

Who ships what

LayerGoogle CloudAWSAzure
Catalogue and managed APIsModel Garden in Agent Platform: Gemini and third-party modelsBedrock: first- and third-party models, Converse APIFoundry Models, including Azure OpenAI deployments
Custom trainingcustom training jobs on GPUs or TPUsSageMaker AI training jobs; HyperPod for large clustersAzure Machine Learning jobs on compute clusters
Hosting your modelendpoints with traffic splitting; batch predictionSageMaker AI real-time, serverless, async, batch transformmanaged online endpoints with deployments; batch endpoints
Pipelines and registrypipelines and model registrySageMaker Pipelines and model registryAzure ML pipelines, registries
Agent runtimeAgent Runtime, Memory Bank, Agent Development KitBedrock AgentCore: runtime, memory, identity, gateway, tools, observabilityFoundry Agent Service
Governance hooksIAM, VPC Service Controls, Model ArmorIAM, PrivateLink, Bedrock GuardrailsEntra ID, private endpoints, content safety

The table is a routing aid, not a feature scorecard. Each cloud has something in every cell; the differences are in depth, regional availability and which models are offered, and those change monthly. Three differences stand out. Google owns its accelerator line, TPUs, alongside GPUs. Microsoft hosts OpenAI's proprietary models as Azure OpenAI deployments inside Foundry. AWS splits generative serving (Bedrock) from classic ML (SageMaker AI) most sharply, which is clearer to reason about and means two consoles, two quota systems and two sets of IAM actions.

Advertisement

The request path, end to end

Trace one call to a managed model and the same stages appear on every cloud. The application authenticates with a workload identity, never a long-lived key where it can be avoided. The request leaves through a private endpoint if you configured one, reaches a regional front door, passes quota and rate checks, optionally passes a safety filter on the input, is routed to capacity that may sit in another region if you opted into cross-region routing, generates, passes the output filter and returns, with token counts attached for billing.

Each stage is also a failure point you must design for. Quota checks return throttling errors, HTTP 429, long before the model is overloaded. Safety filters can block legitimate input and need a user-facing path. Cross-region routing trades availability for data-residency questions that your compliance team will ask. Private endpoints, compared in the private connectivity article, keep traffic off the internet but need DNS set up in every network that calls them. Identity design follows the cloud IAM article: one role per application, scoped to the model IDs it may call.

Calling each platform

The SDKs differ in shape but not in substance. Model IDs below are placeholders; copy real ones from each catalogue, because they are versioned and region-specific.

# AWS: a catalogue model through Bedrock's Converse API
import boto3, json
brt = boto3.client("bedrock-runtime", region_name="us-east-1")
resp = brt.converse(
    modelId="MODEL_ID_OR_INFERENCE_PROFILE_ID",
    messages=[{"role": "user", "content": [{"text": prompt}]}],
    inferenceConfig={"maxTokens": 512, "temperature": 0.2},
)
answer = resp["output"]["message"]["content"][0]["text"]

# AWS: your own model on a SageMaker AI real-time endpoint
smr = boto3.client("sagemaker-runtime")
r = smr.invoke_endpoint(EndpointName="churn-xgb-prod",
                        ContentType="application/json", Body=json.dumps(features))
score = json.loads(r["Body"].read())

# Google Cloud: a catalogue model through the Gen AI SDK in Vertex mode
from google import genai
client = genai.Client(vertexai=True, project="my-project", location="us-central1")
answer = client.models.generate_content(model="MODEL_ID", contents=prompt).text

# Google Cloud: your own model on a deployed endpoint
from google.cloud import aiplatform
aiplatform.init(project="my-project", location="us-central1")
score = aiplatform.Endpoint("ENDPOINT_ID").predict(instances=[features]).predictions

On Azure, hosting your own model shows the deployment model most clearly. An endpoint is a stable URL; deployments behind it are versions; traffic weights move load between them, which gives blue-green and canary releases without touching clients.

from azure.ai.ml import MLClient
from azure.ai.ml.entities import ManagedOnlineEndpoint, ManagedOnlineDeployment, Model
from azure.identity import DefaultAzureCredential

ml = MLClient(DefaultAzureCredential(), SUBSCRIPTION_ID, RESOURCE_GROUP, WORKSPACE)
ep = ManagedOnlineEndpoint(name="churn-prod", auth_mode="key")
ml.online_endpoints.begin_create_or_update(ep).result()

green = ManagedOnlineDeployment(
    name="green", endpoint_name="churn-prod",
    model=Model(path="./model"), environment=ENV, code_configuration=CODE,
    instance_type="Standard_DS3_v2", instance_count=2,
)
ml.online_deployments.begin_create_or_update(green).result()

ep.traffic = {"blue": 90, "green": 10}      # canary 10 percent to the new version
ml.online_endpoints.begin_create_or_update(ep).result()

Google's endpoints and SageMaker AI's production variants offer the same traffic-weight idea. Whichever you use, treat the endpoint name as configuration and the deployment as disposable.

Hosting options and their limits

Hosting is where the clouds look most alike and where the limits bite. AWS publishes the clearest set for SageMaker AI, and they illustrate the trade-offs every platform makes:

OptionPayloadProcessing timeFits
Real-time endpointup to 25 MB60 s, or 8 min when streamingsteady interactive traffic
Serverless inferenceup to 4 MBup to 60 sspiky, low-volume traffic that tolerates cold starts
Asynchronous inferenceup to 1 GBup to 1 hourlarge inputs, queued work, scale to zero
Batch transformdatasets of many GBdaysoffline scoring with no endpoint

The pattern generalises. Always-on endpoints give predictable latency and bill by the instance-hour whether or not traffic arrives. Serverless options bill per request and pay for it in cold starts and smaller limits. Asynchronous and batch paths accept a queue in exchange for large inputs and the lowest cost per item. Choose by traffic shape first and price second; an idle real-time GPU endpoint is the most common line item in a surprising AI bill.

How the money works

Three pricing shapes cover almost everything. Per token for managed models: you pay for input and output tokens, output usually costing more, and the bill scales with traffic and prompt length. Reserved or provisioned throughput: all three clouds sell committed model capacity for predictable latency and volume, at a price that only beats per-token billing when utilisation stays high. Instance-hours for training and self-hosted endpoints: you pay for accelerators whether they are busy or idle.

The per-token versus self-hosted decision is arithmetic. Estimate monthly tokens, multiply by the per-token rate, and compare with the instance-hours of the smallest self-hosted setup that meets your latency objective, plus the engineers who will run it. Low and spiky volume favours per-token; high, steady volume on an open-weight model can favour self-hosting. Batch APIs, offered on all three clouds, usually undercut interactive per-token prices for work that can wait. For training and self-hosted fleets, the spot and reservation tactics in the spot capacity article and the discipline in the FinOps article apply unchanged.

Worked example: choosing for one company

A retailer runs most systems on AWS, has a data warehouse on Google BigQuery, and uses Microsoft 365 with Entra ID for staff identity. It needs a customer support assistant grounded in product documents, a demand-forecast model retrained weekly on warehouse data, and an internal document assistant for employees.

The support assistant goes on Bedrock with AgentCore: the order and customer systems it calls live on AWS, so tools, private networking and IAM are already there. The forecast model trains where its data is, as a custom training job next to BigQuery, and writes forecasts back to the warehouse; moving terabytes weekly to another cloud would cost more than any platform difference. The employee assistant goes on Microsoft Foundry, because its users, documents and access controls already sit in Entra ID and Microsoft 365. Three platforms, each chosen by data gravity and identity, with one shared rule: application code talks to models through an internal interface, so a model can move without a rewrite. That is a deliberate multi-cloud choice, and the trade-offs in the multi-cloud strategy article apply: three sets of skills, three bills, and governance duplicated three times.

Keeping the seam portable

Lock-in is uneven. Calling a catalogue model is easy to move: the request is a prompt and parameters. Agent runtimes, pipelines and evaluation tooling are sticky, because their state, traces and configuration are platform-shaped. Keep the portable layer portable with a thin adapter and put platform-specific code behind it.

import boto3
from google import genai

class ChatModel:                      # the only interface application code sees
    def complete(self, prompt: str, max_tokens: int = 512) -> str: ...

class BedrockChat(ChatModel):
    def __init__(self, model_id): self.c, self.m = boto3.client("bedrock-runtime"), model_id
    def complete(self, prompt, max_tokens=512):
        r = self.c.converse(modelId=self.m,
                            messages=[{"role": "user", "content": [{"text": prompt}]}],
                            inferenceConfig={"maxTokens": max_tokens})
        return r["output"]["message"]["content"][0]["text"]

class VertexChat(ChatModel):
    def __init__(self, model_id, project, location):
        self.c = genai.Client(vertexai=True, project=project, location=location)
        self.m = model_id
    def complete(self, prompt, max_tokens=512):
        cfg = {"max_output_tokens": max_tokens}
        return self.c.models.generate_content(model=self.m, contents=prompt, config=cfg).text

The adapter will not make models interchangeable, because prompts tuned for one model degrade on another, but it turns a migration into a configuration change plus an evaluation run instead of a rewrite. Keep prompts, evaluation sets and traces in your own storage for the same reason. Your evaluation set is the most valuable portable asset you own: a few hundred real prompts with expected answers let you compare a new model on any cloud in an afternoon, and without it every migration becomes an argument about impressions rather than numbers.

Failure modes

  • Quota before capacity. Default tokens-per-minute and requests-per-minute quotas are low; launch day fails with 429 errors unless increases were requested weeks earlier.
  • Model not in your region. Catalogue availability differs by region; a residency requirement can rule out the model you prototyped with.
  • Silent model version changes. Aliases that float to the newest version change behaviour under you; pin versions and re-run evaluations before moving.
  • Idle endpoints. Real-time GPU endpoints left running after a pilot bill every hour.
  • Timeouts at the edge. A 60-second endpoint limit meets a long generation; stream, or move the job to an asynchronous path.
  • Safety filters without a fallback. Input or output filters block a legitimate request and the application shows a raw error; handle the blocked-content response explicitly and log it for review.
  • Egress surprises. Training in one cloud on data stored in another pays data transfer on every epoch; copy once, or train where the data lives.
  • Rename confusion. Documentation, IAM action names and SDK packages lag product names; search by the old and new name before concluding a feature is gone.

What to do next

  1. List each workload's layers: catalogue, training, hosting, pipelines, agents, governance.
  2. Place each workload where its data and identity already live, and write down the exception if you choose otherwise.
  3. Check that the models you need exist in your required regions, and request quota increases before launch.
  4. Pick hosting by traffic shape: real-time, serverless, async or batch; set an idle-endpoint alarm.
  5. Compare per-token and provisioned or self-hosted cost with your real token estimates.
  6. Put a model adapter between application code and platform SDKs, pin model versions, and keep prompts and evaluation sets in your own storage.
  7. Build an evaluation set of real prompts with expected answers before choosing a model, and re-run it on every model or version change.
Key takeaway: Gemini Enterprise Agent Platform (the evolution of Vertex AI), Bedrock with SageMaker AI, and Microsoft Foundry with Azure Machine Learning all provide the same six layers: model catalogue, custom training, hosting, pipelines, agent runtimes and governance. Choose by data gravity and identity rather than feature lists, pick hosting by traffic shape, compare per-token and provisioned cost with real volumes, plan quotas and regional availability early, and keep a thin adapter between your code and the platform so the most portable layer stays portable.