Microsoft Fabric is a software-as-a-service analytics platform that puts data engineering with Spark, a T-SQL warehouse, real-time analytics with KQL, data pipelines and Power BI behind one portal, one storage layer and one billing unit. The pitch is unification. The engineering reality is two shared resources that every team in a tenant depends on: OneLake, where all data lives as Delta tables, and the capacity, a pool of compute that every engine in its workspaces draws from.
Most Fabric incidents trace back to misunderstanding one of those two. This article explains the architecture briefly, then spends its depth where production teams struggle: how capacity units are consumed, smoothed and throttled, with the arithmetic worked through; how table layout decides whether Direct Lake reports are fast or fail; and how to isolate, deploy and operate workloads so one team's job does not take down another team's dashboards. For how Fabric compares with BigQuery, Snowflake and others, see cloud-native analytics engines.
The architecture in one picture
A Fabric deployment lives in a Microsoft Entra tenant. Inside it are one or more capacities, each a distinct pool of compute, and workspaces, containers for items such as lakehouses, warehouses, notebooks, pipelines, eventhouses and semantic models. Each workspace is assigned to one capacity, and every operation in it is billed against that capacity.
All of these items store tabular data in OneLake, a single logical lake per tenant built on Azure Data Lake Storage, with tables in the open Delta Lake format over Parquet. That is the key design choice: the Spark engine, the warehouse engine and Power BI read the same files, so data written once can be queried by any engine without copying. The Delta format itself is explained in Delta Lake on Spark.
Getting data in: copy, replicate or reference
There are three ways to make data available in OneLake, and they have different cost and freshness profiles. Pipelines and dataflows copy data on a schedule, consuming capacity on every run. Mirroring continuously replicates an operational database into OneLake as Delta tables, without you building a change-data-capture pipeline. Shortcuts copy nothing: a shortcut is a pointer from a lakehouse folder to data elsewhere, in another OneLake location, ADLS Gen2, Amazon S3 or Google Cloud Storage, and engines read through it at query time.
Shortcuts are the most attractive and the most misunderstood. They remove duplication, but every read pays the latency and, for other clouds, the egress cost of the source, and the data's layout is whatever the owner wrote. Use shortcuts for data that is already well-formed Delta and read moderately; copy or mirror data that is hot, needs reshaping, or lives far away. Teams that also run Databricks on Azure can share tables both ways through shortcuts; Databricks on Azure integration covers that side.
Capacity: SKUs and capacity units
Compute is purchased as F SKUs through Azure, billed per second with a one-minute minimum, with optional reservations. The number in the SKU is the number of capacity units (CUs): an F2 has 2 CUs, an F64 has 64, and the range runs to F8192. Capacities can be scaled up and down and paused; pausing stops the per-second charge.
| SKU | CUs | Note |
|---|---|---|
| F2 - F32 | 2 - 32 | Viewers of Power BI content need a Pro or PPU licence |
| F64 | 64 | Users with a free licence and a viewer role can view Power BI content |
| F128 - F8192 | 128 - 8192 | Larger Direct Lake guardrails; see below |
The F64 line is a licensing threshold as much as a compute one: below it, every person who views a report needs a paid per-user licence. For an organisation with many report consumers, F64 can be cheaper than a smaller capacity plus hundreds of licences, even if the compute is not needed. Price the two together.
Bursting, smoothing and throttling
Fabric lets an operation use more compute than the SKU provides, called bursting, and then charges that usage against the future, called smoothing. Time is divided into 30-second timepoints, 2,880 per day. Interactive operations, such as report queries, are smoothed over at least five minutes and up to 64 minutes depending on their size; background operations, such as Spark jobs, pipeline runs and most warehouse work, are smoothed over 24 hours.
When smoothed usage exceeds what the capacity provides, the excess accumulates as carried-forward debt that idle capacity later burns down. Throttling depends on how far into the future the capacity is already spent:
| Future capacity already used | Effect |
|---|---|
| Up to 10 minutes | Overage protection: no throttling |
| 10 to 60 minutes | New interactive operations delayed 20 seconds |
| 60 minutes to 24 hours | New interactive operations rejected; background still runs |
| Over 24 hours | All new operations rejected |
Operations already running are never throttled; only new ones are. Some workloads differ: real-time intelligence skips the 20-second delay stage, and eventstreams reduce allocated compute rather than rejecting. Throttling is per capacity, which is the most important operational fact in Fabric: everything assigned to that capacity shares its fate.
Worked example: why the 9 a.m. dashboards slow down
A team runs an F64. Its nightly ETL consumes 400 CU-hours. At 9 a.m., two hundred people open the same dashboards and generate 60,000 CU-seconds of queries in five minutes. The arithmetic:
SKU_CU = 64 # F64
TIMEPOINT_S = 30
CU_S_PER_TP = SKU_CU * TIMEPOINT_S # 1,920 CU-seconds per 30 s timepoint
def background_share(cu_hours):
# background work is smoothed over 24 h = 2,880 timepoints
per_tp = cu_hours * 3600 / 2880
return per_tp / CU_S_PER_TP
def interactive_share(cu_seconds, smooth_minutes=5):
# interactive work is smoothed over at least 5 minutes (10 timepoints)
per_tp = cu_seconds / (smooth_minutes * 2)
return per_tp / CU_S_PER_TP
print(f"{background_share(400):.0%}") # nightly ETL, 400 CU-hours -> 26% of every timepoint
print(f"{interactive_share(60_000):.0%}") # 9 a.m. report burst, 60k CU-s -> 312% for 5 minutesThe ETL, smoothed over 24 hours, takes about a quarter of every timepoint all day, which is affordable. The morning burst, smoothed over only five minutes, needs about three times what the capacity provides during that window. Overage protection absorbs some of it, but a burst of this size, on top of the ETL's background share, pushes the capacity into the 20-second delay stage, so every report interaction feels slow precisely at peak.
The remedies, in order of preference: make the queries cheaper by fixing the model and the Delta layout; move the ETL to a separate capacity so its background share does not eat into interactive headroom; scale up for the morning window; or, as a last resort, accept overage billing, which costs three times the normal rate. Pausing and resuming clears accumulated debt immediately, but it bills the accumulated usage at once and makes all content on the capacity unavailable while paused.
Table layout decides Direct Lake
Direct Lake is a Power BI storage mode in which the semantic model reads Delta columns from OneLake directly into the VertiPaq engine instead of importing a copy. A refresh is just framing, which updates the model to point at the latest Delta table version, and takes seconds. There are two variants. Direct Lake on OneLake can combine tables from several Fabric sources and never falls back to DirectQuery. Direct Lake on SQL goes through a single source's SQL analytics endpoint and falls back to DirectQuery when it cannot read Delta directly, such as for SQL views or when the warehouse enforces SQL-level row security; the model's Direct Lake behavior property controls whether fallback is allowed.
Each SKU has guardrails per table. On F2 to F32, a table may have at most 1,000 Parquet files, 1,000 row groups and 300 million rows; on F64, 5,000 files, 5,000 row groups and 1,500 million rows. Exceed one and Direct Lake on OneLake fails refresh until the table is fixed, while Direct Lake on SQL, if fallback is enabled, refreshes with only a warning, falls back to DirectQuery and gets slower. Thousands of small files from frequent appends are the usual cause, so table maintenance is not optional:
# Fabric notebook (PySpark), attached to the silver lakehouse
spark.conf.set("spark.sql.parquet.vorder.default", "true") # this table feeds Direct Lake
src = "abfss://Sales@onelake.dfs.fabric.microsoft.com/Bronze.Lakehouse/Tables/orders_raw"
orders = (spark.read.format("delta").load(src)
.where("order_ts >= current_date() - INTERVAL 1 DAY")
.dropDuplicates(["order_id"]))
from delta.tables import DeltaTable
target = DeltaTable.forName(spark, "orders") # silver lakehouse table
(target.alias("t")
.merge(orders.alias("s"), "t.order_id = s.order_id")
.whenMatchedUpdateAll()
.whenNotMatchedInsertAll()
.execute())
# keep file and row-group counts inside Direct Lake guardrails
spark.sql("OPTIMIZE orders VORDER")
spark.sql("VACUUM orders RETAIN 168 HOURS")V-Order is Fabric's write-time Parquet optimisation that speeds up reads, especially by Direct Lake. New workspaces default to the writeHeavy Spark resource profile, which sets spark.sql.parquet.vorder.default to false to make ingestion cheaper. Enable V-Order deliberately on tables that feed reports, and leave it off on raw ingestion tables.
Isolating workloads and deploying changes
Because throttling is per capacity, capacity assignment is your blast-radius tool. A common layout separates capacities by workload class: one for scheduled engineering, one for business-facing reports, and a small one for development. Organise workspaces by layer and domain, for example bronze, silver and gold per domain, so permissions and capacity assignment follow data ownership. Workspace roles (Admin, Member, Contributor and Viewer) are coarse; keep row-level rules in the semantic model and grant consumers access to gold workspaces only.
Treat Fabric items as code. Workspaces can be connected to a Git repository so notebooks, pipelines and semantic model definitions are versioned and reviewed, and deployment pipelines promote content from development through test to production. Promote by pipeline, never by editing production directly, and test semantic model changes against production-sized data in test, since guardrails depend on table sizes.
Monitoring
Install the Microsoft Fabric Capacity Metrics app on day one. Its throttling charts show when smoothed usage crosses the 10-minute, 60-minute and 24-hour limits, and the timepoint drill-through shows which operations, users and items consumed the CUs. Set capacity threshold alerts, and review the top consumers weekly: a small number of expensive items usually dominate. When users report the error code CapacityLimitExceeded, the capacity is rejecting work; slowness without that error is more often item design than throttling.
Failure modes and trade-offs
| Symptom | Likely cause | Fix |
|---|---|---|
| Reports slow every morning | Interactive burst plus background share on one capacity | Optimise models; split capacities; scale for the window |
| Direct Lake refresh fails | Table over guardrails, usually small files | OPTIMIZE regularly; reduce partition count |
| Direct Lake report suddenly slow | Fallback to DirectQuery on a SQL view or SQL row security | Materialise the view; disable fallback to surface the problem |
| One team's job breaks everyone's reports | Shared capacity | Separate capacities by workload class |
| Unexpected bill spike | Overage billing or a forgotten large capacity | Alert on usage; pause dev capacities outside hours |
| Slow queries over shortcut data | Remote, poorly laid-out source | Copy or mirror hot data into OneLake |
The central trade-off is simplicity versus control. Fabric removes cluster sizing and storage plumbing, and its open Delta storage means data is not locked into one engine; in exchange, compute is a shared pool with rules you must learn, and Spark autoscale billing, which bills pay-as-you-go, gives up smoothing entirely. Teams that already run open tables elsewhere, such as an Iceberg and Trino stack, should weigh that control against Fabric's integrated BI.
What to do next
- Price F SKUs together with per-user licences, using the F64 viewer threshold.
- Install the Capacity Metrics app and set threshold alerts before onboarding users.
- Separate capacities for scheduled engineering, business reporting and development.
- Choose per source between copying, mirroring and shortcuts based on freshness, layout and egress.
- Schedule OPTIMIZE and VACUUM, and enable V-Order on tables that feed Direct Lake.
- Check each Direct Lake table against your SKU's guardrails, and decide whether fallback should be allowed.
- Connect workspaces to Git and promote changes only through deployment pipelines.