Azure Machine Learning studio is the web interface at ml.azure.com for an Azure Machine Learning workspace. It is easy to treat the studio as the product, clicking through notebooks, the designer and AutoML wizards, and then discover that nothing you built can be reproduced, reviewed or promoted to production. The studio is better understood as a window onto a set of versioned Azure resources: data assets, environments, compute, jobs, models and endpoints. Everything you can click, you can also create from the v2 CLI or the Python SDK, azure-ai-ml, against the same REST API.
This page explains that resource model from first principles, then walks one model, a customer churn classifier, from versioned data to a managed online endpoint with staged traffic, with real SDK v2 code at each step. It covers the managed virtual network and its one-way settings, the failures teams hit most often, and cost controls. For how Azure Machine Learning compares with SageMaker and Google's platform, see cloud-native ML compared. Microsoft steers new generative AI and agent projects towards its Foundry portal; this page is about the workspace you use to train, track and deploy your own models.
The workspace and its dependent resources
A workspace is an Azure resource in a resource group. On its own it stores metadata: the history of jobs, the versions of assets, the definitions of endpoints and who may do what. Data and artefacts live in a set of dependent resources linked to it. A storage account holds the default datastores where uploaded data, job outputs and logs land. A Key Vault holds secrets such as datastore credentials. Application Insights receives telemetry from online endpoints. A container registry stores the Docker images built from environments, and can be created on demand the first time an image is built if you do not supply one.
Those dependencies matter operationally. Deleting the workspace does not delete them, network rules on the storage account can break the studio's file browser, and the identity the workspace uses must hold roles on each one. Access to the workspace itself is Azure role-based access control through Microsoft Entra ID; the built-in AzureML Data Scientist role lets people run jobs and manage assets without changing the workspace's configuration or compute. Identity design for the whole estate is covered in Microsoft Entra ID in depth.
Assets: data, environments, models and components
Assets are versioned, named objects. Referencing churn-raw:2026-10-01 in a job means the same input every time, which is what makes a run reproducible.
- Datastores are connections to storage, such as a blob container or an ADLS Gen2 filesystem, preferably using identity-based access instead of stored keys. Paths inside them are addressed as
azureml://datastores/<name>/paths/<path>. - Data assets are named, versioned pointers of type
uri_file,uri_folderormltable. Registering a data asset does not copy the data, so the underlying folder must be treated as immutable; write each version to a new path. - Environments describe the software a job runs in: a base Docker image plus a conda or pip specification, or a complete Dockerfile. Microsoft publishes curated environments in the shared
azuremlregistry. Pin every package version, or the same environment version can build differently next month. - Models are registered artefacts with a type. MLflow models carry their own signature and dependencies, which enables deployment without a scoring script.
- Components are reusable, versioned job steps with typed inputs and outputs, the building blocks of pipelines.
Registries extend this across workspaces: an environment, component or model published to a registry can be used by a development workspace and a production workspace in different subscriptions, which is the clean way to promote a model rather than copying files between workspaces.
Compute: instances, clusters and serverless
A compute instance is a single managed VM for one person's notebooks and VS Code sessions. It bills while running, so configure idle shutdown and a schedule; forgotten instances are the most common source of surprise cost. A compute cluster is a managed, autoscaling pool of VMs for jobs. Set the minimum node count to zero so it scales to nothing when idle, and choose an idle time before scale-down that balances queue latency against paying for idle nodes. Clusters can use dedicated or low-priority VMs.
Serverless compute removes the cluster from the picture. When a command, sweep, AutoML or pipeline job omits its compute target, Azure Machine Learning provisions capacity for that job and releases it afterwards, choosing a VM size from your quota unless you specify instance_type and instance_count in the job's resources. The queue_settings job tier is Standard (dedicated, the default) or Spot (low priority and preemptible). Serverless jobs consume the same Azure Machine Learning quota as clusters, so quota, not cluster size, becomes the thing to manage. Attached Kubernetes clusters, including AKS, are the option when you must run on infrastructure you already operate.
Worked example: from versioned data to a registered model
The churn team keeps daily extracts in a blob container. Each extract is registered as a version of the churn-raw data asset. A versioned environment pins scikit-learn and MLflow. The training script reads --data, trains a gradient-boosted classifier, logs metrics with MLflow and saves an MLflow model to --model-out. The code below registers the inputs, submits the job to serverless Spot capacity, streams its logs and registers the output as a model version.
from azure.ai.ml import MLClient, command, Input, Output
from azure.ai.ml.constants import AssetTypes
from azure.ai.ml.entities import Data, Environment, Model
from azure.identity import DefaultAzureCredential
ml = MLClient(DefaultAzureCredential(), "<subscription-id>", "rg-ml-prod", "ws-churn")
# 1. Version the input data. The path points at a datastore; nothing is copied.
data = ml.data.create_or_update(Data(
name="churn-raw", version="2026-10-01", type=AssetTypes.URI_FOLDER,
path="azureml://datastores/workspaceblobstore/paths/churn/2026-10-01/"))
# 2. Version the environment. The image is built once and cached in the workspace registry.
env = ml.environments.create_or_update(Environment(
name="churn-train", version="3",
image="mcr.microsoft.com/azureml/openmpi4.1.0-ubuntu22.04:latest",
conda_file="env/conda.yml")) # pinned versions only
# 3. A command job. No compute= argument, so it runs on serverless compute.
job = command(
code="./src",
command="python train.py --data ${{inputs.data}} --model-out ${{outputs.model}}",
inputs={"data": Input(type=AssetTypes.URI_FOLDER, path="azureml:churn-raw:2026-10-01")},
outputs={"model": Output(type=AssetTypes.MLFLOW_MODEL)},
environment="churn-train:3",
experiment_name="churn",
queue_settings={"job_tier": "Spot"}, # Standard (default) or Spot; Spot can be preempted
)
run = ml.jobs.create_or_update(job)
ml.jobs.stream(run.name) # blocks and prints logs until the job ends
# 4. Register the job's output as a versioned model.
model = ml.models.create_or_update(Model(
name="churn", type=AssetTypes.MLFLOW_MODEL,
path=f"azureml://jobs/{run.name}/outputs/model"))
print(model.name, model.version)The ${{inputs.data}} placeholders are resolved at run time to a mounted or downloaded path on the compute, so the script never contains a storage URL or credential. Because the job ran on Spot capacity it can be preempted, so the training script should checkpoint into its outputs and resume from them. Every run records its code snapshot, environment version, input versions and metrics, which the studio displays under the experiment and is what lets you answer, months later, exactly which data and code produced the model in production.
Deploying: managed online endpoints and staged traffic
A managed online endpoint is a stable HTTPS address with authentication. Behind it sit one or more deployments, each a model plus environment plus VM size and instance count, and the endpoint splits traffic between them by percentage. That split is the mechanism for safe releases: deploy the new model as a second deployment, test it directly by name, mirror a share of live traffic to it if you want to compare without affecting users, then shift live traffic in steps. The release pattern itself is described in canary deployments.
from azure.ai.ml.entities import ManagedOnlineEndpoint, ManagedOnlineDeployment
ENDPOINT = "churn-scorer-weu" # must be unique in the Azure region
ep = ManagedOnlineEndpoint(name=ENDPOINT, auth_mode="key")
ml.online_endpoints.begin_create_or_update(ep).result()
# An MLflow model needs no scoring script or environment: the platform supplies both.
green = ManagedOnlineDeployment(
name="green", endpoint_name=ENDPOINT,
model=f"azureml:churn:{model.version}",
instance_type="Standard_DS3_v2", # pick from the managed online endpoint SKU list
instance_count=2)
ml.online_deployments.begin_create_or_update(green).result()
# Smoke test the new deployment directly before it gets any user traffic.
print(ml.online_endpoints.invoke(endpoint_name=ENDPOINT, deployment_name="green",
request_file="sample-request.json"))
# Shift traffic in steps; "blue" is the deployment already serving.
ep = ml.online_endpoints.get(ENDPOINT)
ep.traffic = {"blue": 90, "green": 10}
ml.online_endpoints.begin_create_or_update(ep).result()Plan instance counts for production at more than one instance per deployment, and keep spare quota for the platform to perform upgrades. For custom models that are not MLflow, you supply a scoring script with init(), which loads the model once per worker, and run(), which handles each request. For scoring large datasets on a schedule rather than per request, a batch endpoint runs a job over a data asset instead of keeping VMs warm.
The managed virtual network
By default the workspace's compute has outbound internet access and the workspace is reachable over public endpoints. A managed virtual network moves compute and endpoints into a Microsoft-managed network with private endpoints to the dependent resources. The isolation mode is set per workspace: allow_internet_outbound permits outbound internet, allow_only_approved_outbound permits only private endpoints, service tags and fully qualified domain names you list. FQDN rules are implemented with Azure Firewall, which adds its cost to your bill.
Three properties catch teams out. The settings are one-way: a workspace moved to allow internet outbound cannot return to disabled, and one moved to approved outbound only cannot return to allow internet outbound. The managed network is not created when you set the mode but when the first compute is created, which can then take around 30 minutes, or when you provision it manually. And an online deployment needs the network provisioned first, so provision explicitly as part of workspace setup. Private connectivity from your own networks to the workspace uses private endpoints, described in private connectivity.
from azure.ai.ml.entities import ManagedNetwork, IsolationMode
ws = ml.workspaces.get("ws-churn")
ws.managed_network = ManagedNetwork(IsolationMode.ALLOW_ONLY_APPROVED_OUTBOUND)
ml.workspaces.begin_update(ws).result()
# Provision now rather than on first compute creation; required before an online deployment.
ml.workspaces.begin_provision_network(workspace_name="ws-churn").result()
Failure modes
| Symptom | Usual cause | Fix |
|---|---|---|
| Environment image build fails after enabling approved outbound only | pip or conda cannot reach package indexes | add FQDN rules for your package sources, mirror them privately, or build from a prebuilt image |
| Job waits in queue, then errors on quota | no quota for the VM family in the region | request quota for that family, choose another size, or let serverless pick from available quota |
| Online deployment fails readiness or liveness probes | slow init() loading a large model | raise the probe initial delay, load lazily, or use a larger instance |
| Job cannot read a datastore | identity-based access but the user or managed identity lacks a storage data role | grant Storage Blob Data Reader on the container to the identity the job runs as |
| First compute takes half an hour or times out | managed network provisioned implicitly | provision the network explicitly before creating compute |
| Bill dominated by idle compute | instances left running, cluster minimum above zero | idle shutdown, schedules, minimum nodes of zero, serverless for batch work |
| Spot job restarts from scratch | preemption without checkpoints | checkpoint to outputs and resume |
If you still have code on the v1 SDK, azureml-core, treat migration as overdue: Microsoft ended support for SDK v1 on 30 June 2026, and support for the v1 CLI extension ended on 30 September 2025. The v2 concepts map closely: datasets become data assets, run configurations become command jobs, and ScriptRunConfig becomes command(). Teams that run Spark on the same data often pair the workspace with Azure Databricks, covered in Azure Databricks in depth.
What to do next
- List what lives in the workspace today and which of it exists only as studio clicks; recreate it as SDK v2 or CLI v2 definitions in source control.
- Switch datastores to identity-based access and grant storage data roles to the identities that run jobs.
- Version every data asset to a new, immutable path and pin every environment package.
- Move batch training to serverless or to clusters with a minimum of zero nodes; set idle shutdown on every compute instance.
- Register models from job outputs, not uploaded files, so lineage is recorded.
- Deploy through a second deployment on the same endpoint and shift traffic in steps.
- Decide on the managed network mode deliberately, knowing it is one-way, and provision it before creating compute.
- Remove any remaining
azureml-corecode.