Azure Databricks is sold as a first-party Azure service: you create it from the Azure portal, it bills through your Azure subscription, and it signs users in with Microsoft Entra ID. Underneath, it is the Databricks platform split across two administrative worlds. Databricks operates a control plane that holds the web application, the job scheduler and the catalog services; the machines that run your Spark code live either in your own subscription (classic compute) or in a Databricks-operated serverless plane. Almost every networking ticket, permission error and surprise bill comes from forgetting which side of that split something lives on.
This article builds the architecture from first principles, then walks a worked deployment: a Premium workspace injected into a corporate virtual network, Unity Catalog reading Azure Data Lake Storage Gen2 through a managed identity, secrets held in Key Vault, and jobs promoted by CI. It ends with recurring failure modes, cost levers and a checklist.
The two planes, and what lives where
The control plane is Databricks' own service in each Azure region. It serves the workspace UI and REST API, stores notebook source and job definitions, runs the scheduler, and hosts the Unity Catalog metastore service that records which tables exist, who may read them and where their files live.
The classic compute plane is in your subscription. When you create a workspace, Azure also creates a managed resource group that Databricks controls through a deny assignment: you can see the virtual machines, disks and (by default) a virtual network inside it, but you cannot change them. When a cluster starts, the control plane asks Azure to create VMs there, the VMs boot a Databricks runtime image, and they connect back to the control plane for instructions. You pay Azure for the VMs and Databricks for the Databricks Units (DBUs) the cluster consumes.
The serverless compute plane is run by Databricks inside its own Azure account in the same region. Serverless SQL warehouses, serverless jobs and notebooks start in seconds because a pool of capacity is already warm, and you pay only DBUs, which include the infrastructure. The trade is control: serverless machines are not in your VNet, so the network rules that protect your storage must explicitly admit them.
One commercial fact shapes every new design. Microsoft Learn's end-of-life notice for the Standard tier states that from 1 April 2026 all new workspaces must be Premium, and that Standard workspaces still remaining on 1 October 2026 are upgraded to Premium automatically. Unity Catalog, Databricks SQL, serverless compute and Lakeflow pipelines are listed as Premium-only, so treat Premium plus Unity Catalog as the baseline, and check your own workspaces' tier in the account console because the upgrade changes the DBU rate on your bill.
Networking: VNet injection and secure cluster connectivity
By default the managed resource group contains a VNet you cannot touch. Enterprises almost always use VNet injection instead: you supply your own VNet with two dedicated, delegated subnets, conventionally called host (public) and container (private), and every cluster node takes one address in each. That lets you peer to a hub and route egress through a firewall.
Size the subnets for peak concurrent nodes across all clusters, because addresses, not quotas, become the hard ceiling. Azure reserves five addresses in every subnet, so a /24 yields 251 usable addresses and therefore at most 251 concurrently running nodes in that workspace, while a /26 yields only 59. Undersized subnets show up as clusters stuck pending with an address-exhaustion error. Subnets cannot be resized casually under a running workspace, so err large.
Secure cluster connectivity (also called No Public IP) should be on for every workspace. Nodes get no public addresses and no inbound ports are opened; instead each cluster opens an outbound connection over HTTPS on port 443 to a relay in the control plane, and all administration flows through that tunnel. Because the connection is outbound, your egress path must allow it: with a firewall or NAT gateway in front of the subnets, a missing rule to the regional relay endpoint shows up as clusters that launch VMs and then fail to become ready.
Private Link comes in two directions. Front-end Private Link puts a private endpoint for the workspace UI and REST API into your network, so browsers, BI tools and CI reach it without traversing the internet. Back-end Private Link makes the classic cluster-to-relay traffic use a private endpoint too. Regulated deployments use both and then disable public network access on the workspace. Serverless compute does not use secure cluster connectivity at all; to reach firewalled storage it needs its own allowance, configured at the account level (Databricks calls this a network connectivity configuration), which can add private endpoints from the serverless plane to your storage account.
Identity and governance with Unity Catalog
Users and groups should come from Entra ID, provisioned into the Databricks account rather than individual workspaces, so a group defined once can be granted rights everywhere. Automation runs as service principals or managed identities, not as a person.
Unity Catalog reaches storage without anyone holding a key. The chain has three links, and every storage permission problem is a broken link in it:
- An Access Connector for Azure Databricks, an Azure resource that carries a system-assigned or user-assigned managed identity. You grant that identity an Azure RBAC role such as Storage Blob Data Contributor on the ADLS Gen2 account or container.
- A storage credential in Unity Catalog that references the connector's identity.
- An external location that pairs an
abfss://path with that credential, and on which you grant fine-grained rights to groups.
Using a managed identity rather than a service principal secret removes secret rotation and, according to Microsoft's documentation, lets Unity Catalog reach storage accounts protected by network rules, which a service principal cannot. Creating the connector needs Contributor or Owner on a resource group; granting its identity access to storage needs Owner or User Access Administrator on the storage account, which is why this step usually waits on a platform team.
-- Run as a metastore or catalog admin, after the connector identity has its RBAC role
-- and the storage credential lake_mi has been created from the connector's resource ID.
CREATE EXTERNAL LOCATION IF NOT EXISTS sales_raw
URL 'abfss://raw@contosolake.dfs.core.windows.net/sales'
WITH (STORAGE CREDENTIAL lake_mi);
GRANT READ FILES ON EXTERNAL LOCATION sales_raw TO `data-engineers`;
CREATE CATALOG IF NOT EXISTS sales
MANAGED LOCATION 'abfss://managed@contosolake.dfs.core.windows.net/sales';
GRANT USE CATALOG, CREATE SCHEMA ON CATALOG sales TO `data-engineers`;
GRANT USE CATALOG ON CATALOG sales TO `analysts`;The storage credential itself is normally created in Catalog Explorer or through the API or Terraform by pasting the connector's resource ID, which is why the script starts after it. After this, analysts read sales.gold.orders as a governed table: Unity Catalog checks their grants, then the cluster accesses the files using the connector identity on their behalf. Retire the legacy pattern of dbutils.fs.mount with a stored secret: every user of a mount inherits its access.
Worked example: a governed workspace, deployed as code
Contoso wants one production workspace in West Europe, injected into a spoke VNet, with no public IPs, front-end Private Link for analysts, and a lake in ADLS Gen2 that only Unity Catalog can reach. The Azure half is Terraform with the azurerm provider; the Databricks half uses the databricks provider or CI-driven SQL.
resource "azurerm_databricks_workspace" "prod" {
name = "dbw-contoso-prod-weu"
resource_group_name = azurerm_resource_group.data.name
location = "westeurope"
sku = "premium"
managed_resource_group_name = "rg-dbw-contoso-prod-managed"
public_network_access_enabled = false
network_security_group_rules_required = "NoAzureDatabricksRules"
custom_parameters {
no_public_ip = true # secure cluster connectivity
virtual_network_id = azurerm_virtual_network.spoke.id
public_subnet_name = azurerm_subnet.host.name # /24 each
private_subnet_name = azurerm_subnet.container.name
public_subnet_network_security_group_association_id = azurerm_subnet_network_security_group_association.host.id
private_subnet_network_security_group_association_id = azurerm_subnet_network_security_group_association.container.id
}
}
resource "azurerm_databricks_access_connector" "uc" {
name = "dbac-contoso-prod"
resource_group_name = azurerm_resource_group.data.name
location = "westeurope"
identity { type = "SystemAssigned" }
}
resource "azurerm_role_assignment" "uc_lake" {
scope = azurerm_storage_account.lake.id
role_definition_name = "Storage Blob Data Contributor"
principal_id = azurerm_databricks_access_connector.uc.identity[0].principal_id
}Disabling public network access while requiring no Databricks-managed NSG rules is the setting that makes front-end and back-end Private Link mandatory, so the private endpoints and private DNS zones must exist before anyone can log in. Check the provider documentation for the combinations your version accepts.
Jobs and pipelines are then defined in a bundle in the same repository and promoted by CI with databricks bundle validate followed by databricks bundle deploy -t prod, running as a service principal. Secrets the code needs (an API token for a SaaS source, say) live in Azure Key Vault and are exposed through a Key Vault-backed secret scope, read at runtime:
token = dbutils.secrets.get(scope="kv-contoso-prod", key="crm-api-token")
df = (spark.read.format("json")
.load("abfss://raw@contosolake.dfs.core.windows.net/crm/2026/10/01/"))
df.write.mode("append").saveAsTable("sales.bronze.crm_events") # governed, UC-managed
Data flow of a single query
An analyst in Power BI runs a query against a SQL warehouse. The request reaches the workspace through the front-end private endpoint; Entra ID authenticates the user. The warehouse asks Unity Catalog whether the user may SELECT from sales.gold.orders, applying any row filters or column masks defined on the table. Unity Catalog returns the table's storage path and short-lived, path-scoped credentials derived from the Access Connector identity. The warehouse reads the Delta transaction log, prunes files using statistics, reads only the needed Parquet files from ADLS, executes (with Photon on warehouses), and streams results back.
Three consequences follow. Storage firewall rules must admit whichever compute ran the query, classic subnets or serverless. Audit logs record the user identity even though storage sees the connector identity. And a user who bypasses Unity Catalog, for instance with a personal access key to the storage account, bypasses every grant, so storage keys should be disabled where possible.
Cost: what you are really paying for
A classic cluster's hourly cost is the VM price times nodes plus the DBU rate times DBUs consumed; serverless rolls both into a single DBU price. Bursty or idle-heavy work usually costs less on serverless; long jobs that run flat out can be cheaper on classic with reserved or spot workers.
| Lever | Effect | Watch out for |
|---|---|---|
| Auto-termination on interactive clusters | Stops idle clusters; the commonest single saving | Too short annoys users and they disable it |
| Job clusters instead of all-purpose | Job compute carries a lower DBU rate | Startup time per run |
| Spot workers, on-demand driver | Large VM savings for fault-tolerant batch | Evictions lengthen runs; never spot the driver |
| Cluster policies | Cap node types, sizes and runtimes per team | Policies nobody updates block new runtimes |
| Serverless for SQL and short jobs | No idle billing, fast start | Needs network connectivity config for private storage |
| Tags and system billing tables | Attribute spend to teams and jobs | Untagged clusters make chargeback impossible |
Make budgets observable before optimising them: tag every cluster and job with a cost centre through a policy, then query the billing system tables weekly. For the engine-level settings that decide runtime and therefore cost, see Spark adaptive query execution and the Photon engine.
Failure modes that recur
- Clusters pending forever. Subnet addresses exhausted or a vCPU quota hit in the region. Check the cluster event log for the Azure error, then raise the quota or widen the subnets.
- Clusters launch then fail to start. Egress to the secure cluster connectivity relay, the artifact storage or the metastore blocked by a firewall or user-defined route. Compare your egress allow-list with the regional endpoints Microsoft documents.
- 403 on
abfss://paths. One link in the connector-credential-location chain is missing: the RBAC role, the storage credential, the external location, or the user's grant. Or the storage firewall does not admit the compute that ran the query, which is the classic cause when something works on a classic cluster and fails on serverless. - Login works from the office but not from CI. Front-end Private Link with public access disabled; the CI runner is outside the network. Run CI on self-hosted agents inside it.
- Legacy mounts leak access. A mount created with one principal's credentials gives every workspace user that principal's access. Migrate to external locations and remove mounts.
Operational guidance and trade-offs
Run one metastore per region (Unity Catalog's model) and attach every workspace in that region to it; separate environments by workspace and catalog, with production catalogs bound only to production workspaces. Keep the Delta tables in ADLS Gen2 with hierarchical namespace enabled; the ADLS Gen2 guide covers the storage side, and the Delta Lake article covers the table format that makes the transaction log and time travel work.
The central trade-off is control versus convenience. Classic compute in an injected VNet gives packet-level control and your own reserved VMs, at the price of subnet planning, firewall rules and slower startup; serverless removes that load and idle cost but runs on Databricks' network. Most estates run both. Either way, put the estate in code and ship it through a pipeline.
What to do next
- Confirm every workspace you own is on Premium, and model the DBU-rate change on your bill.
- Enable secure cluster connectivity, inject new workspaces into your own VNet, and size host and container subnets for peak concurrent nodes plus headroom.
- Create an Access Connector, grant its identity a storage role, and replace every mount and stored storage key with storage credentials and external locations.
- Provision users and groups from Entra ID at the account level; run all jobs as service principals.
- Decide which workloads go serverless and configure network connectivity so serverless can reach your firewalled storage.
- Put workspace, connector, role assignments, grants and jobs in Terraform and bundles, deployed by CI.
- Enforce cluster policies with auto-termination and cost tags, and review billing system tables weekly.