Most teams meet MLOps as a shopping list: an experiment tracker, a feature store, a registry, a serving layer, a drift dashboard. Buying all of them does not produce a system that ships models safely, in the same way that owning a compiler, a test runner and a container registry does not produce continuous delivery. What makes the parts into an architecture is the set of contracts between them and the control loops that run across them.

This article treats MLOps as that architecture. It starts from the one property that makes machine learning different from other software, derives the planes a platform needs, and then spends most of its time on the parts that decide whether a team can release a model on a Tuesday afternoon without a meeting: what each kind of automation tests, how lineage is recorded, how promotion rules become code, when retraining fires and who owns which failure. The individual stages of an ML pipeline are covered in AI/ML pipeline architecture; this page is about the platform those stages run on.

Advertisement

Why ML systems change in three ways

A conventional service changes when someone commits code. A model-backed service changes in three independent ways. The code changes: feature logic, training scripts, the serving wrapper. The data changes: new rows arrive, an upstream team renames a column, user behaviour shifts after a product launch. And the model changes: retraining on the same code and a new snapshot produces different weights, and so different behaviour, without a single line of diff.

Two consequences follow. A green CI build proves little about what users experience, because most of what determines predictions never passed through CI. And failures are silent: a broken join that fills a feature with zeros raises no exception, and the only symptom is a business metric drifting down over a fortnight.

MLOps is the discipline of making all three kinds of change go through gates that are as explicit as code review, and of recording enough about each change that any prediction can be traced back to the code, data and configuration that produced it.

The planes of an MLOps architecture

It helps to split the platform into five planes with different owners and failure modes:

  • Data plane. Sources, validation and immutable snapshots. Its contract is that every training run can name the exact data it read, and that data failing validation never reaches training.
  • Training plane. Parameterised pipelines that turn a code commit and a data snapshot into a candidate model plus an evaluation report. Its contract is reproducibility: the same inputs produce an equivalent model.
  • Release plane. The registry, the promotion gate and deployment. Its contract is that only models that passed a written policy reach traffic, and that the previous model is one command away.
  • Serving plane. Online endpoints or batch scoring jobs, with feature retrieval that matches training. Covered in model serving architecture.
  • Observation plane. Monitoring of inputs, outputs, live quality and system SLOs, feeding signals back into the trigger policy.

Underneath all five sits the metadata and lineage store. It is the least glamorous component and the most important one: without it the planes are separate tools, with it they are one system you can reason about.

Code + configgit, CI testsData sourcestables, eventsValidate + versionschema, snapshotTraining pipelinetrain, evaluateModel registryversions, aliasesServingonline / batchMonitoringquality, drift, SLOsTrigger policyschedule, drift, dataMetadata and lineage storerun -> data snapshot -> code commit -> model version -> deploymentCI buildsCDsignalsCTsnapshot idlineagedeploy log
An MLOps platform as planes plus control loops. CI gates code, CT turns signals into training runs, CD gates promotion, and every step writes to one lineage store.
Advertisement

Maturity levels and what each one buys

Google Cloud's widely cited MLOps guidance describes three levels, and they are a useful way to decide what to build next rather than a ladder to climb for its own sake.

LevelWhat is automatedTypical symptom that you need the next level
0: manual processNothing end to end. A data scientist trains in a notebook and hands over a model file.Nobody can say which data produced the model in production; retraining takes weeks.
1: ML pipeline automationThe training pipeline itself, so the system retrains on new data (continuous training) and validates data and models automatically.Pipeline code changes are deployed by hand and break production runs.
2: CI/CD pipeline automationBuilding, testing and deploying the pipeline code, in addition to running it.At this level the bottleneck is usually evaluation quality, not automation.

The key insight at level 1 is that the artefact you deploy is the pipeline, not the model; models become its continuous outputs, judged without a human in every loop. Teams with only a few production models get most of the value from a disciplined level 1 plus lightweight CI for pipeline code.

CI, CD and CT test different things

These three acronyms are often used interchangeably. They should not be, because each one guards a different kind of change.

LoopTriggered byWhat it must testWhat passing proves
CIA code commitUnit tests on feature functions, schema contracts, a tiny end-to-end training run on a fixture dataset, serving wrapper testsThe pipeline code is not broken
CTSchedule, new data, drift or a quality alertData validation on the new snapshot, training, evaluation against a frozen holdout and against the championA new candidate exists and is measured
CDA candidate passing CTThe promotion policy, a load test, then shadow or canary trafficThe candidate is safe to serve

The small end-to-end training run in CI is worth singling out: train a few steps on a few hundred fixture rows, assert that the loss decreases and that the exported model loads in the serving container. It takes a minute and catches bugs that otherwise surface as a failed overnight run.

The metadata spine: lineage you can query

Every run should write a record that links a data snapshot id, a git commit, the resolved configuration, the environment image digest, the evaluation set version and the resulting model version. Every deployment should write which model version went where and when. With those records you can answer the questions incidents actually ask: which models read the table that was corrupted last Thursday, what changed between the model that worked and the one that does not, which customers were scored by version 41.

Registries, trackers and orchestrators each hold part of this graph. Make one, usually the registry, the source of truth and have the others write their ids into it as tags. The registry's side of the contract is described in ML model registry, and point-in-time correct features, which make snapshots meaningful, in feature store architecture.

Promotion policy as code

A promotion decision made in a meeting is slow and inconsistent; one made by a script nobody reviewed is dangerous. The middle path is to write the policy as code, review it like any other code, and run it as the last step of every training pipeline. The example below uses MLflow's model registry. Registry stages such as Staging and Production are deprecated in MLflow; the supported mechanism is aliases, mutable names that point at one version, set with set_registered_model_alias and resolved with get_model_version_by_alias or the URI models:/churn@champion.

# promote.py - runs as the last step of every training pipeline run.
# Policy is code: reviewed in git, versioned with the pipeline, same rules for every model.
from mlflow import MlflowClient

POLICY = {
    "primary_metric": "auc",
    "min_gain": 0.005,          # challenger must beat champion by this much on the frozen holdout
    "max_slice_drop": 0.010,    # no business slice may regress by more than this
    "max_p95_latency_ms": 40,   # measured by the pipeline's load-test step
    "required_tags": ["data_snapshot", "git_commit", "eval_set_version"],
}

def decide(client, name, candidate_version):
    cand = client.get_model_version(name, candidate_version)
    run = client.get_run(cand.run_id)
    missing = [t for t in POLICY["required_tags"] if t not in run.data.tags]
    if missing:
        return False, f"lineage incomplete: {missing}"

    try:
        champ = client.get_model_version_by_alias(name, "champion")
    except Exception:
        return True, "no champion yet"          # first model: human review still required
    champ_run = client.get_run(champ.run_id)
    if run.data.tags["eval_set_version"] != champ_run.data.tags["eval_set_version"]:
        return False, "eval sets differ: re-score champion first"

    m, cm = run.data.metrics, champ_run.data.metrics
    key = POLICY["primary_metric"]
    if m[key] - cm[key] < POLICY["min_gain"]:
        return False, f"{key} gain {m[key] - cm[key]:.4f} below threshold"
    for k, v in m.items():
        if k.startswith(f"slice.{key}.") and cm.get(k, v) - v > POLICY["max_slice_drop"]:
            return False, f"slice regression on {k}: {cm[k]:.3f} -> {v:.3f}"
    if m.get("p95_latency_ms", 1e9) > POLICY["max_p95_latency_ms"]:
        return False, "latency budget exceeded"
    return True, "passes policy"

def promote(name, version):
    client = MlflowClient()
    ok, reason = decide(client, name, version)
    client.set_model_version_tag(name, version, "gate_reason", reason)
    if ok:
        client.set_registered_model_alias(name, "challenger", version)   # canary picks this up
    return ok, reason

Three design choices matter more than the thresholds. The policy refuses to compare models evaluated on different evaluation sets, because a better number on an easier set is not an improvement. It checks slices, not only the headline metric. And it promotes to a challenger alias, not straight to champion: the canary or shadow stage moves champion only after live evidence, and rollback is a single alias change back to the previous version. Serving code loads by alias, never by version number.

A worked example: a weekly churn model

A subscription business retrains a churn classifier every Monday on the previous 12 months of labelled accounts. The champion scores an AUC of 0.842 on the frozen holdout. This week the pipeline produces a challenger with 0.851, a gain of 0.009, which clears the 0.005 threshold.

The slice check then compares per-segment AUC. Established customers improve from 0.861 to 0.872, but customers in their first 90 days drop from 0.781 to 0.766, a regression of 0.015, above the 0.010 limit. The gate blocks promotion and tags the version with the reason. Investigation shows a new onboarding flow launched three weeks ago changed the meaning of an early-engagement feature, and the challenger learned the old meaning from 11 months of history.

Without the slice gate, the headline number would have shipped a model that got worse for exactly the customers the retention team cares most about. With it, the fix is an ordinary feature change through CI.

Retrain triggers and the control loop

Continuous training needs a policy for when to run. Schedules are simple and predictable; data-volume triggers use fresh labels as soon as they are worth it; drift and quality triggers react to change. Combine them, and make the policy refuse to train while an upstream data incident is open, so a broken feed cannot produce a model trained on bad data.

# Continuous-training trigger, evaluated hourly by the orchestrator.
def should_retrain(state, now):
    reasons = []
    if now - state.last_train_at > state.max_model_age:             # e.g. 7 days
        reasons.append("schedule")
    if state.new_labelled_rows >= state.min_new_rows:              # e.g. 50,000
        reasons.append("new data")
    if state.feature_psi_max > 0.2 and state.drift_hours >= 6:     # sustained, not a blip
        reasons.append("input drift")
    if state.live_metric is not None and state.live_metric < state.alert_floor:
        reasons.append("live quality below floor")
    if state.upstream_incident_open:                               # never train on broken data
        return []
    return reasons

The thresholds here, a population stability index above 0.2 sustained for six hours for example, are starting points to tune against your own false-alarm rate. Drift detection covers how to compute the signals. Remember that drift is a reason to evaluate, not proof that a new model will be better; the promotion gate still decides.

Failure modes that platforms must catch

  • Training-serving skew. Features computed one way offline and another way online. Share feature code or definitions between both paths, and log served features to compare against training distributions.
  • Evaluation leakage. The holdout slowly leaks into training through re-splits or duplicated entities. Freeze the evaluation set, version it, and split by entity and time rather than by row.
  • Feedback loops. The model's own decisions shape its future labels, so a fraud model that blocks transactions never sees whether they were fraud. Keep a small randomised holdout of traffic where the model does not act.
  • Orphaned models. A model keeps serving after its owner leaves and its pipeline stops. Alert on model age and on pipelines that have not produced a candidate within their schedule.
  • Silent upstream changes. A column changes units or encoding. Data validation with explicit expectations on ranges, null rates and categories catches most of these before training.

Ownership, SLOs and build versus buy

The most common organisational failure is that nobody owns the model once it ships. Split ownership explicitly: the platform team owns the planes, the metadata store and the gates as mechanisms; the model team owns its pipeline, its evaluation set, its thresholds and its on-call for quality alerts; the product owner signs off the slice definitions, because they encode what the business cares about.

Give each model two kinds of SLO: system SLOs for latency, availability and feature freshness, and model SLOs for a live quality proxy, maximum model age and time to roll back. Without a quality proxy, at least alert on prediction distributions.

On build versus buy: managed platforms remove the undifferentiated work of running trackers, registries and serving clusters, and suit small teams. What you cannot buy is your evaluation sets, slice definitions, promotion policy and lineage discipline; keep those in your own repository.

What to do next

  1. Pick one production model and write down its code commit, data snapshot and evaluation set. If you cannot, lineage is your first project.
  2. Freeze and version an evaluation set with at least three business slices, and re-score the current champion on it.
  3. Add a one-minute end-to-end training test to CI on a fixture dataset.
  4. Write the promotion policy as a reviewed script and run it at the end of every training run, promoting to a challenger alias only.
  5. Make serving load by alias and rehearse a rollback by moving the alias back.
  6. Add a retrain trigger that combines schedule, new data and sustained drift, and blocks during upstream incidents.
  7. Assign an owner, a model-age limit and a quality alert to every model in production.
Key takeaway: MLOps is not a set of tools but an architecture built around one fact: ML systems change through code, data and retrained models, and only code passes through ordinary CI. Split the platform into data, training, release, serving and observation planes joined by one lineage store. Automate the pipeline before you automate its deployment, and test code in CI, new candidates in CT and safety in CD. Encode promotion rules, including slice checks and matching evaluation sets, as reviewed code that moves an alias, and give every model an owner and SLOs. The expensive mistakes are silent ones, and the platform exists to make them loud.