"Cloud-native ML" means running the whole lifecycle of your own models, from data to training to evaluation to a versioned artifact to a serving endpoint, on cloud infrastructure that scales on demand. Teams face one decision early and live with it for years: use a cloud's managed ML platform, or build the same lifecycle on Kubernetes with open-source components.
This page compares those two paths component by component, across the three large clouds. It is about models you train or fine-tune yourself. Calling hosted foundation models and building agents is a different problem, covered in Cloud AI Platforms Compared, which also lists the managed hosting limits. Here you will get the job contract that makes either path portable, a working checkpoint-and-resume entrypoint, Kubernetes manifests for queued training and canary serving, a cost model you can fill with your own prices, and the failure modes that decide the choice in practice.
The lifecycle and its six parts
Whatever you run it on, the lifecycle has the same parts. A data layer holds training data in object storage or a warehouse, sometimes with a feature store to keep training and serving features consistent. A training service runs a container on accelerators with inputs and outputs in object storage. An orchestrator runs multi-step pipelines and records lineage. A registry stores versioned model artifacts with metadata and an approval state. A serving layer turns an artifact into an endpoint with autoscaling and traffic splitting. Monitoring watches latency, errors and data drift and triggers retraining.
Each managed platform bundles all six behind one console and one IAM model. The Kubernetes-native approach assembles them from separate projects on a cluster you operate. The rest of this page compares each part.
The two architectures side by side
| Part | AWS | Google Cloud | Azure | Kubernetes-native |
|---|---|---|---|---|
| Training | SageMaker AI training jobs; HyperPod for large persistent clusters | custom training jobs on GPUs or TPUs (Gemini Enterprise Agent Platform, formerly Vertex AI) | Azure Machine Learning command jobs on compute clusters | Job or JobSet under Kueue; Ray; Kubeflow Trainer |
| Pipelines | SageMaker Pipelines | Pipelines, which run compiled Kubeflow Pipelines v2 definitions | Azure Machine Learning pipelines | Kubeflow Pipelines, Argo Workflows |
| Registry | SageMaker Model Registry (model package groups) | Model Registry | Azure Machine Learning models and registries, MLflow-compatible | MLflow Model Registry or equivalent |
| Serving | real-time, serverless, async endpoints; batch transform | online endpoints with traffic split; batch prediction | managed online and batch endpoints | KServe, Ray Serve, plain Deployments |
| Preemptible capacity | managed spot training | Spot VMs for custom training | low-priority VMs on compute clusters | spot node pools plus Kueue |
Two things stand out. First, the managed platforms have converged: each offers the same parts with different names, and each runs your container. Second, Google's pipelines execute the Kubeflow Pipelines v2 format and Azure Machine Learning uses MLflow for tracking and model formats, so the managed and open-source worlds already share some formats. The Google platform page goes deeper on one managed stack, and the managed Kubernetes comparison covers the clusters the open-source path runs on.
The portable job contract
Every training service, managed or not, runs a container and gives it three locations: where to read data, where to write checkpoints, and where to write the final model. The platforms advertise these in different ways: SageMaker AI sets SM_MODEL_DIR and one SM_CHANNEL_<NAME> per input channel as local directories; Google's custom training sets AIP_MODEL_DIR and AIP_CHECKPOINT_DIR as Cloud Storage URIs; Azure Machine Learning passes paths through job input and output bindings on the command line. If your entrypoint takes the three paths as arguments, with local-path variables only as fallbacks, the same image runs on any of them and on a plain Kubernetes Job.
The entrypoint must also resume. Preemptible capacity, node failures and maintenance all kill jobs, and a job that restarts from step zero turns a cheap spot discount into an expensive loop.
import argparse, glob, os, torch
def first(*vals):
return next((v for v in vals if v), None)
ap = argparse.ArgumentParser()
ap.add_argument("--data", default=first(os.environ.get("SM_CHANNEL_TRAIN")))
ap.add_argument("--checkpoint-dir", default="/opt/ml/checkpoints") # SageMaker AI syncs this to S3
ap.add_argument("--model-dir", default=first(os.environ.get("SM_MODEL_DIR")))
ap.add_argument("--steps", type=int, default=20_000)
ap.add_argument("--save-every", type=int, default=500)
args = ap.parse_args()
model, opt = build_model(), build_optimizer()
start = 0
ckpts = sorted(glob.glob(os.path.join(args.checkpoint_dir, "step_*.pt")))
if ckpts: # resume after preemption
state = torch.load(ckpts[-1], map_location="cpu")
model.load_state_dict(state["model"]); opt.load_state_dict(state["opt"])
start = state["step"] + 1
for step in range(start, args.steps):
train_step(model, opt, next_batch(args.data, step)) # data order keyed by step
if step % args.save_every == 0:
tmp = os.path.join(args.checkpoint_dir, f".tmp_{step}.pt")
torch.save({"model": model.state_dict(), "opt": opt.state_dict(), "step": step}, tmp)
os.replace(tmp, os.path.join(args.checkpoint_dir, f"step_{step:08d}.pt")) # atomic
save_final(model, args.model_dir) # the registry picks this upThree details matter. The write goes to a temporary name and is renamed, so a job killed mid-save never leaves a truncated checkpoint that the next attempt would load. The data order is a function of the step, so a resumed job does not re-see or skip examples. And the code assumes local paths: SageMaker AI sets SM_MODEL_DIR to a local directory and managed spot training syncs /opt/ml/checkpoints to the S3 location you configure. Google's AIP_MODEL_DIR and AIP_CHECKPOINT_DIR are Cloud Storage URIs, not local paths, so on Google pass a local or mounted directory with --checkpoint-dir and upload checkpoints and the final model to those URIs yourself; renames on a mounted bucket are not guaranteed to be atomic, so write locally first. On Kubernetes you mount a volume or sync to object storage in the same way.
Queued training on Kubernetes
Kubernetes on its own schedules pods, not jobs. A distributed training job needs all its workers at once; if half start and half wait for GPUs, the running half burn money doing nothing. Kueue adds queues, quotas per team and all-or-nothing admission: a Job is held suspended until the whole job fits within its queue's quota, then released. Submitting to it is a label on an ordinary Job.
apiVersion: batch/v1
kind: Job
metadata:
name: finetune-ranker-0930
labels:
kueue.x-k8s.io/queue-name: team-ranking # a LocalQueue in this namespace
spec:
parallelism: 4
completions: 4
completionMode: Indexed
backoffLimit: 20 # tolerate preemptions; resume makes retries cheap
template:
spec:
restartPolicy: OnFailure
containers:
- name: trainer
image: registry.example.com/ranker-train:1.8.2
args: ["--data", "/data/train", "--checkpoint-dir", "/ckpt", "--model-dir", "/out"]
resources:
limits:
nvidia.com/gpu: 1The managed platforms give you the equivalent without running anything: submit a job specification with an instance type and count, and the service provisions, runs and tears down the machines. What Kubernetes buys is sharing: one pool of GPUs across teams, with quotas and borrowing, which is where reserved capacity pays for itself. The spot capacity guide covers interruption handling and pool diversification in depth.
Pipelines, lineage and the registry
A pipeline is a DAG of steps: prepare data, train, evaluate, and conditionally register. Its value is less the automation than the record: which data, code version and parameters produced which artifact. Managed pipelines record lineage automatically. On Kubernetes, Kubeflow Pipelines records runs and artifacts, and Argo Workflows records only what your steps write.
The registry is the handoff between training and serving. Each entry should carry the artifact location, the training run and data snapshot that produced it, evaluation metrics against a fixed benchmark, and an approval state that deployment automation reads. SageMaker AI calls the container of versions a model package group and has an approval status per version; Azure Machine Learning and MLflow use registered models with versions and aliases or stages. The model registry guide covers the design in full. The rule that matters most: deployment should pull a model by registry version, never by a raw bucket path, so a rollback is a registry change.
Serving and safe rollout
Managed endpoints give autoscaling, traffic splitting between variants and several modes for different traffic shapes. On Kubernetes, KServe provides the same with an InferenceService resource. In its Serverless (Knative) deployment mode, setting canaryTrafficPercent on an update sends that share of traffic to the new revision while the previous one keeps the rest; promoting is removing the field. In raw deployment mode this field does not apply, and you split traffic with your ingress or service mesh instead.
apiVersion: serving.kserve.io/v1beta1
kind: InferenceService
metadata:
name: ranker
spec:
predictor:
canaryTrafficPercent: 10 # new revision gets 10 percent
minReplicas: 2
maxReplicas: 20
model:
modelFormat:
name: pytorch
storageUri: s3://models/ranker/v42 # resolved from the registry by CI, not typed by hand
resources:
limits:
nvidia.com/gpu: 1Whichever platform you choose, decide in advance which metrics end a canary: error rate, latency percentiles, and a business metric such as click-through. Automate the rollback on those thresholds.
Monitoring closes the loop. Operational signals, latency and errors, come from the endpoint on either path. Model signals need more work: log a sample of requests and predictions to object storage, compare input feature distributions with the training snapshot recorded in the registry, and join predictions with outcomes when labels arrive, often days later. The managed platforms offer model-monitoring features that compute drift against a baseline on a schedule; on Kubernetes you typically run the same comparison as a scheduled pipeline. Either way, a drift alert should open a ticket or trigger a retraining pipeline run with the same entrypoint, never deploy on its own: the new model still passes through evaluation, registration and a canary.
Worked example: choosing for one team
A team of six trains a ranking model weekly on about 64 GPU-hours, fine-tunes a small language model monthly on about 200 GPU-hours, and serves both at steady traffic. No one on the team runs Kubernetes today.
Build the cost model with your own prices. Let P be the on-demand price of the GPU instance, m the managed platform's markup over the raw instance for the same hardware, H the GPU-hours per month, and E the monthly cost of the engineering time needed to run the open-source stack. Managed costs roughly H * P * (1 + m); self-run costs roughly H * P + E + cluster overhead. Managed is cheaper whenever H * P * m is less than E. With about 460 GPU-hours a month, the markup on a few hundred GPU-hours is small next to even a fraction of an engineer, so this team should use its cloud's managed platform, with the portable entrypoint above so the decision stays reversible.
The answer flips when H grows by one or two orders of magnitude, when several teams can share reserved GPUs through quotas, or when a platform team already operates Kubernetes. It also flips for reasons that are not cost: data that must stay in a specific cluster, or a requirement to run the same stack on-premises. Choose the cloud by data gravity first, because moving training data between clouds every week costs more than any platform difference.
Failure modes
- Non-resumable training. Jobs on preemptible capacity without checkpoint-and-resume restart from zero; watch for jobs whose retries exceed their useful steps.
- Platform-shaped code. Training scripts that import a platform SDK for paths and logging cannot move. Keep SDK calls in the job specification, not in the container.
- Idle endpoints. Always-on GPU endpoints for models called a few times a day are the largest avoidable line item; use batch or asynchronous modes for them.
- Registry bypass. A model deployed from a bucket path has no lineage and no clean rollback.
- Training-serving skew. Features computed one way in the pipeline and another way online. A shared feature definition, as described in the feature store guide, removes the most common cause.
- Partial gang starts. Distributed jobs on plain Kubernetes that start some workers and wait for the rest; use Kueue or an operator with gang scheduling.
- Underestimating the open-source stack. Kubeflow, KServe and MLflow each need upgrades, databases, backups and security patches. Count that work in
E.
What to do next
- Write down your GPU-hours per month, number of teams and whether anyone already operates Kubernetes; compute the break-even with the cost model above.
- Make the training entrypoint take data, checkpoint and model paths as arguments with platform variables as fallbacks, and test a kill-and-resume.
- Put every deployable model in a registry with lineage, metrics and an approval state, and deploy only by registry version.
- Move rarely-called models to batch or asynchronous serving and remove idle endpoints.
- Define canary metrics and automatic rollback thresholds before the next model release.
- Revisit the managed versus Kubernetes decision when GPU-hours grow tenfold or a second team needs shared capacity.