BigQuery runs every query on slots, units of compute that its scheduler hands out to the stages of a query plan. How you pay for those slots decides both your bill and how predictable your performance is. On-demand pricing charges for bytes scanned and draws slots from a shared pool. Capacity pricing, through BigQuery editions, charges for slots over time, and reservations let you divide that capacity between teams and workloads.

The overview in BigQuery architecture explains what a slot is and where it sits in the engine. This page is the operational guide: the object model of commitments, reservations and assignments, how baseline, autoscaling and idle sharing interact, how to size from your own job history, the SQL to set it up, a worked example with three teams, and the failure modes that make capacity bills surprising.

What a slot does and how slots are shared

A query is compiled into stages, and each stage is split into many units of work that run in parallel. A slot runs one unit of work at a time. A stage that can be split into 2,000 pieces can use 2,000 slots at once; a final aggregation that cannot be split uses few. So slot demand is spiky within a single query, and doubling the slots available rarely halves the run time, because some stages are serial and shuffles between stages cost time regardless.

The scheduler shares slots fairly. Within a reservation, available slots are divided between the projects running jobs, and then between the jobs in each project, and a running query's share shrinks when new queries arrive. This is why a reservation does not reject work when full: queries slow down rather than fail, and if demand persists they queue. Reservations also support a reservation-based fairness mode that divides idle slots equally across reservations rather than across projects.

The unit of measurement is slot milliseconds. A job's total_slot_ms divided by its duration in milliseconds gives its average slot usage. That figure, aggregated over time, is the foundation of every sizing decision below.

On-demand versus capacity editions

There are two ways to pay. On-demand bills per TiB of data processed, with a per-project concurrency limit on slots drawn from shared capacity; check the current price list for your region. It suits spiky, small or unpredictable workloads and needs no capacity planning, but a careless full-table scan is expensive and, unlike capacity pricing, the cost scales with data scanned, not time. Capacity pricing bills slot-hours through an edition, and queries cost nothing extra per byte. Most mature platforms use both.

StandardEnterpriseEnterprise Plus
Baseline slotsNo (autoscale only)YesYes
CommitmentsNo1-year or 3-year1-year or 3-year
Idle slot sharingNoYesYes
Max reservation size1,600 slotsQuota-boundQuota-bound
Assignment job typesQUERY, PIPELINEAlso CONTINUOUS, ML_EXTERNAL, BACKGROUNDSame as Enterprise
CMEK, BI EngineNoYesYes

Google lists commitment discounts of roughly 20 percent for one year and 40 percent for three years against pay-as-you-go slot prices; confirm the current figures before buying. Enterprise Plus adds managed disaster recovery and compliance controls. For dashboards, BI Engine can serve repeated small queries from memory and reduce the slots a BI reservation needs, so evaluate it before sizing that pool. An edition is fixed when a reservation is created: to change it you delete the reservation and create a new one, so plan migrations rather than flipping a setting.

Commitments, reservations and assignments

Capacity lives in an administration project, usually a dedicated one per region, so billing and permissions for capacity sit apart from the data projects. Inside it you create reservations, named pools with an edition, a baseline and an autoscaling ceiling. Commitments are optional purchases in the same project and region that discount a number of slots for a term. Assignments connect a project, folder or organisation to a reservation for a job type.

Capacity model: an admin project owns reservations; assignments map projects, folders or orgs to themCommitment (optional)1 or 3 years, discountedAdmin project, region-usholds reservations, commitmentsbacks baselineReservation: etlbaseline 500, max 1000Reservation: bibaseline 300, max 600Reservation: adhocbaseline 0, max 400projects/etl-prodjob_type PIPELINE, QUERYprojects/looker-prodjob_type QUERYfolders/analyticsjob_type QUERYnoneexplicitly on-demandprojects/sandboxIdle baseline slots flow between reservations of the same edition and region; autoscaled slots never do.
An admin project per region holds reservations. Assignments route projects or folders to them; a project assigned to none runs on-demand.

Assignment resolution follows the resource hierarchy: a project uses the single most specific reservation it is assigned to, so a project assignment beats a folder assignment, which beats an organisation assignment. A project assigned to the special none reservation runs on-demand, which is how you exempt a sandbox from a folder-wide assignment. Job types matter: a QUERY assignment does not cover load jobs, which need PIPELINE. Without a PIPELINE assignment, load jobs run in BigQuery's shared pool, which costs nothing extra but comes with no capacity guarantee, so a nightly load can take much longer on a busy night.

Keep the number of reservations small. Each one is a pool you have to size, monitor and explain on a chargeback report. A common starting shape is one reservation per workload class, such as batch ETL, dashboards and exploration, rather than one per team, with folders mapping teams onto classes.

Baseline, idle slots and autoscaling

A reservation gets slots from three places, in a fixed order. First its baseline, which is always allocated and always billed, used or not. Second, idle slots: unused baseline and committed slots from other reservations in the same admin project, edition and region. Third, autoscaling, up to the reservation's maximum. Autoscaled slots are added in multiples of 50, are billed for the slots scaled rather than the slots used, and by default carry a one-minute minimum once added. They are also never lent to other reservations as idle capacity.

Two consequences follow. Autoscaling is billed in coarse steps, so a workload of many short bursts can pay for considerably more slot time than it used; compare billed autoscale slot-hours against consumed slot-hours to see that waste. And idle sharing makes performance depend on neighbours: a reservation that runs fast because it borrows another team's idle baseline will slow down the moment that team gets busy. Set ignore_idle_slots to true on a reservation that must behave the same every day; it can still lend its own idle slots to others.

Predictable reservations add a different control: instead of a baseline plus autoscale_max_slots, you set max_slots as a cap on the reservation's total consumption, and a scaling_mode of ALL_SLOTS, IDLE_SLOTS_ONLY or AUTOSCALE_ONLY to say where slots above baseline may come from.

Sizing from job history

Size from history, not intuition. INFORMATION_SCHEMA.JOBS_TIMELINE records slot milliseconds per job per second in period_slot_ms. Summing it per second and dividing by 1,000 gives slots in use each second; percentiles over a few weeks give you the demand curve.

-- Slot demand per second over 28 days, then percentiles by hour of day
WITH per_second AS (
  SELECT
    period_start,
    SUM(period_slot_ms) / 1000 AS slots
  FROM `etl-prod.region-us`.INFORMATION_SCHEMA.JOBS_TIMELINE
  WHERE period_start >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 28 DAY)
  GROUP BY period_start
)
SELECT
  EXTRACT(HOUR FROM period_start) AS hour_utc,
  APPROX_QUANTILES(slots, 100)[OFFSET(50)] AS p50,
  APPROX_QUANTILES(slots, 100)[OFFSET(90)] AS p90,
  APPROX_QUANTILES(slots, 100)[OFFSET(99)] AS p99
FROM per_second
GROUP BY hour_utc
ORDER BY hour_utc;

Seconds with no running jobs do not appear in the view, so the percentiles describe busy seconds only; keep that in mind when reading the median. A practical rule: set baseline near the level demand sits above most of the working day, cover the p90 to p99 band with autoscaling, and accept that the spikes above p99 will queue or slow. For an on-demand project, the same query tells you what capacity would have cost: multiply slot-hours by the edition price and compare with the bytes-billed charges from INFORMATION_SCHEMA.JOBS. Cutting bytes scanned with partitioning and clustering lowers slot demand too, so do that before sizing.

Setting it up as code

Reservations and assignments can be created with DDL run in the admin project, which makes them easy to keep in version control alongside other infrastructure:

CREATE RESERVATION `capacity-admin.region-us.etl`
OPTIONS (
  edition = 'ENTERPRISE',
  slot_capacity = 500,            -- baseline
  autoscale_max_slots = 500       -- extra on top of baseline, so max 1000
);

CREATE RESERVATION `capacity-admin.region-us.bi`
OPTIONS (edition = 'ENTERPRISE', slot_capacity = 300, autoscale_max_slots = 300);

CREATE ASSIGNMENT `capacity-admin.region-us.etl.etl_prod_pipeline`
OPTIONS (assignee = 'projects/etl-prod', job_type = 'PIPELINE');

CREATE ASSIGNMENT `capacity-admin.region-us.etl.etl_prod_query`
OPTIONS (assignee = 'projects/etl-prod', job_type = 'QUERY');

CREATE ASSIGNMENT `capacity-admin.region-us.bi.looker`
OPTIONS (assignee = 'projects/looker-prod', job_type = 'QUERY');

The bq tool offers the same operations, for example bq mk --reservation with --slots, --edition and --autoscale_max_slots, and bq mk --reservation_assignment with --assignee_type, --assignee_id and --job_type. Restrict who can create assignments: anyone who can assign a project to a reservation can spend that capacity. Cloud IAM covers role scoping; the reservation admin roles belong on the admin project only.

Worked example: three workloads, one region

A data platform team runs three workloads in US: nightly ETL, a Looker dashboard estate and analysts' ad hoc queries. All three are on-demand, and the monthly bill is dominated by a few ETL jobs that rescan large tables. The sizing query shows ETL demand around 450 slots from 01:00 to 05:00 UTC, p99 near 1,000; BI around 200 slots during business hours with p99 near 550; ad hoc demand is low most of the day with short spikes to 400.

They create an Enterprise admin project with three reservations: etl with a baseline of 500 and autoscaling to 1,000, bi with a baseline of 300 and autoscaling to 600, and adhoc with a baseline of 0 and autoscaling to 400. Overnight, etl borrows the idle BI baseline, so it rarely autoscales. During the day, ad hoc spikes borrow idle ETL baseline. The dashboards had a latency SLO, so bi is set to ignore idle slots, keeping its behaviour independent of ETL. A one-year commitment covers the 800 baseline slots that are always billed anyway.

After a month they review billed against used slot-hours. Ad hoc autoscaling shows heavy waste: many spikes last seconds but are billed for a minute in 50-slot steps. They cap adhoc at 200, accept slower spikes, and move the sandbox project to none so experiments run on-demand with a per-user custom quota on bytes billed.

Failure modes

  • Performance changes when neighbours change. A reservation relied on idle slots that disappeared when another team's workload grew. Check how much of each reservation's usage is borrowed.
  • Wrong job type. Load or export jobs fall outside the reservation because only QUERY was assigned. Assign PIPELINE explicitly and check reservation_id in job metadata.
  • Region mismatch. Reservations are regional; a dataset in EU cannot use slots in US. Jobs run where their data lives and use whatever assignment, or on-demand pricing, applies in that region.
  • Autoscale ceiling as a budget. Teams set a high maximum "just in case" and pay for it every busy minute. Treat the maximum as a spending limit.
  • Edition locked in. Choosing Standard then needing idle sharing or CMEK means recreating reservations and assignments.
  • Uncontrolled assignments. A folder-level assignment silently moves a new project onto capacity, or onto none. Audit assignments in code review.

Trade-offs

On-demand rewards efficient queries and punishes scans; capacity rewards steady utilisation and punishes idle baseline. Baseline gives predictability at the cost of paying for idle time; autoscaling gives elasticity at the cost of coarse billing. Idle sharing raises utilisation and lowers isolation. One large reservation maximises sharing but makes chargeback hard and lets one team starve another; many small ones are fair and wasteful. Most teams settle on a few reservations by workload class, with isolation only where an SLO demands it.

What to do next

  • Run the slot-demand query for each major project and plot p50, p90 and p99 by hour.
  • Compare four weeks of on-demand bytes-billed cost against the capacity cost of the same slot-hours.
  • Pick an edition from the features you need, remembering that changing it means recreating reservations.
  • Create a dedicated admin project per region and define reservations and assignments as code.
  • Assign PIPELINE as well as QUERY where batch loads should use reserved capacity.
  • Turn on ignore_idle_slots only for reservations with a latency SLO.
  • Review billed versus used autoscale slot-hours monthly and adjust baselines and caps.
  • Buy commitments only for slots you have kept as baseline for several months.
Key takeaway: Slots are BigQuery's compute, and reservations decide who gets them. Size baselines from JOBS_TIMELINE history, cover peaks with autoscaling while watching its per-minute, 50-slot billing, use idle sharing for utilisation and ignore_idle_slots where an SLO needs isolation, route projects with explicit assignments per job type, and commit only to capacity you have proven you use.