Azure Functions is Microsoft's event-driven compute service: you write a function, attach it to a trigger such as an HTTP route, a queue or a timer, and the platform decides how many machines run it. That description fits every serverless product, and the general model of cold starts, concurrency and retries is covered in serverless in the cloud. What makes Azure Functions its own thing is the details: a two-process runtime, a hosting plan that is hard to change later, a binding system that removes plumbing code, and Durable Functions for work that outlives a single invocation.
This article explains those details from first principles, using facts checked against Microsoft's documentation on 2026-10-01, and closes with a worked order-processing app, the failures that show up in production, and a checklist.
What runs when a function runs
Every function app instance runs two processes. The Functions host is a .NET process owned by Microsoft. It listens to triggers, manages bindings, reads host.json, enforces timeouts and reports to the platform. The language worker is a separate process that runs your code: the .NET isolated worker, the Java worker, Node.js, Python, PowerShell or a custom handler. Host and worker talk over gRPC on the same machine. When a message arrives, the host claims it from the source, builds the input, sends an invocation request to the worker, waits for the result, writes any output bindings and only then completes the message.
Outside the instance, a scale controller watches the event sources, for example the length of a Service Bus queue or the rate of HTTP requests, and decides how many instances should exist. A storage account referenced by the AzureWebJobsStorage setting holds state the host needs: blob leases that make sure only one instance runs a timer, function keys, and on some plans the deployment package itself.
This split explains two surprises. A function timeout restarts the worker process, so every in-flight execution on that instance fails with it. And startup has budgets: 60 seconds for the worker, and on Flex Consumption 30 seconds for the whole app, so heavy dependency injection setup shows up as failed starts, not just slow first requests.
Choosing a hosting plan
The plan decides how you scale, what you pay for, whether you can reach a private network, and which operating system you run on. Microsoft now treats the original Consumption plan as legacy and recommends Flex Consumption for new serverless apps. Linux Consumption is retiring on 30 September 2028 and receives no new language versions; Windows Consumption is not affected yet. You cannot migrate an existing app into or out of Flex Consumption in place, you create a new app and redeploy, so this choice is worth getting right first.
| Plan | Scale and limits | Pick it when |
|---|---|---|
| Flex Consumption | Linux only, 512, 2,048 or 4,096 MB instances, up to 1,000 on-demand instances per scale group, optional always-ready instances, VNet integration, one app per plan | Default for new event-driven work, including private network access |
| Premium | Always at least one warm instance, prewarmed workers, Windows or Linux, containers on Linux, VNet, 100 instances on Windows and 20 to 100 on Linux depending on region | Steady load, larger instances, many apps sharing one plan, custom Linux images |
| Dedicated (App Service) | You choose instance counts and size, manual or rule-based autoscale, slots | Spare App Service capacity, predictable billing, App Service Environment isolation |
| Container Apps | Container-only, scale to zero optional, up to 1,000 replicas subject to core quota, GPU options | You want to own the image and run functions beside other containers |
| Consumption (legacy) | 5-minute default and 10-minute maximum timeout, 1.5 GB memory, no VNet | Existing Windows apps that depend on Windows-only features |
Two Flex limits are easy to miss. First, all Flex apps in one subscription and region share a default quota of 250 cores; a 2,048 MB instance counts as one core, so one app at 250 instances can starve every other Flex app in that region. Always-ready instances count against the quota. Second, Flex has no deployment slots: zero-downtime releases use a rolling-update site strategy, which was still in public preview when this was checked.
The isolated worker model and the November 2026 deadline
For years .NET functions ran in-process: your assembly loaded into the host and had to match its .NET version and dependencies. The isolated worker model runs your code in its own process like every other language, so you control the .NET version and dependency tree. Support for the in-process model ends on 10 November 2026; apps keep running afterwards without security or feature updates, and Flex Consumption never supported it. If you still have in-process apps, migration is the most urgent item on this page.
// Program.cs - .NET isolated worker
using Microsoft.Azure.Functions.Worker;
using Microsoft.Extensions.DependencyInjection;
using Microsoft.Extensions.Hosting;
var host = new HostBuilder()
.ConfigureFunctionsWebApplication() // ASP.NET Core integration for HTTP triggers
.ConfigureServices(services =>
{
services.AddApplicationInsightsTelemetryWorkerService();
services.AddSingleton<IOrderStore, CosmosOrderStore>(); // one client per process
})
.Build();
host.Run();
// IngestOrder.cs
public class IngestOrder(IOrderStore store, ILogger<IngestOrder> log)
{
[Function("IngestOrder")]
public async Task Run(
[ServiceBusTrigger("orders", Connection = "OrdersBus")] OrderMessage msg,
CancellationToken ct)
{
// Idempotent write keyed by the business id: redelivery must not double-insert.
await store.UpsertAsync(msg.OrderId, msg, ct);
log.LogInformation("Stored order {OrderId}", msg.OrderId);
}
}Note the singleton client: a new client per invocation exhausts outbound connections (600 active per instance on the legacy Consumption plan). Pass the CancellationToken through so shutdown stops work cleanly.
Triggers, bindings and host.json
A trigger starts an execution and supplies its input; each function has exactly one. Bindings are declarative inputs and outputs: instead of writing code to open a queue client and send a message, you declare an output binding and return a value. Bindings are implemented by extensions inside the host, which is why non-.NET apps pull them in through an extension bundle version range in host.json. The Python v2 programming model shows the idea compactly:
import json
import azure.functions as func
app = func.FunctionApp()
@app.route(route="orders", methods=["POST"], auth_level=func.AuthLevel.FUNCTION)
@app.queue_output(arg_name="out", queue_name="orders-in", connection="AzureWebJobsStorage")
def submit_order(req: func.HttpRequest, out: func.Out[str]) -> func.HttpResponse:
order = req.get_json()
if "orderId" not in order:
return func.HttpResponse("orderId required", status_code=400)
out.set(json.dumps(order)) # written by the host after the function returns
return func.HttpResponse(status_code=202)Host-wide behaviour lives in host.json: timeout, logging, and per-extension settings such as Service Bus concurrency per instance.
{
"version": "2.0",
"functionTimeout": "00:10:00",
"extensions": {
"serviceBus": { "maxConcurrentCalls": 16, "autoCompleteMessages": true }
},
"logging": {
"applicationInsights": {
"samplingSettings": { "isEnabled": true, "excludedTypes": "Request" }
}
}
}Connections should use managed identity rather than connection strings. For identity-based connections, a setting such as OrdersBus__fullyQualifiedNamespace replaces the secret, and the app's identity gets a data-plane role on the namespace; see Microsoft Entra ID in depth for how those identities and roles work. Anything that must stay a secret belongs in Key Vault, referenced from app settings, as discussed in cloud secrets management.
How scaling decisions are made
Scaling is driven by concurrency: how many executions one instance should handle at once. For queue-like sources the desired instance count is roughly the backlog divided by the target executions per instance; for HTTP it is the request rate against the configured HTTP concurrency. Higher concurrency means fewer, busier instances.
Flex Consumption adds per-function scaling. All HTTP triggers in an app scale together as one group, all Event Grid-based blob triggers form another, and all Durable Functions triggers a third. Every other function, for example each Service Bus or Event Hubs trigger, scales on its own instances. The maximum instance count you set applies to each group separately, and always-ready instances sit outside it. The scale-out rate itself is not configurable: Microsoft describes it as fast at small sizes and progressively more measured at high instance counts, and advises against designing to specific per-interval numbers.
So a noisy queue function no longer steals instances from your HTTP API, though the regional core quota is still shared. How scaling interacts with downstream limits in general is covered in cloud autoscaling.
Long work: the 230-second ceiling and Durable Functions
An HTTP-triggered function has at most 230 seconds to respond, whatever functionTimeout says, because of the idle timeout of the load balancer in front of it. Flex and Premium allow unbounded execution for non-HTTP triggers, but during scale-in an execution gets a 60-minute grace period, and platform updates allow 10 minutes. Long work must therefore be resumable whatever the plan.
Durable Functions is the built-in answer. An orchestrator function describes a workflow in ordinary code; each await on an activity checkpoints progress to a storage provider. When the orchestrator wakes up it replays from the start, skipping completed steps by reading their recorded results. That is why orchestrator code must be deterministic: no DateTime.UtcNow, no random numbers, no direct I/O. Use the context's clock and put side effects in activities.
[Function(nameof(FulfilOrder))]
public static async Task<string> FulfilOrder([OrchestrationTrigger] TaskOrchestrationContext ctx)
{
var order = ctx.GetInput<OrderMessage>()!;
await ctx.CallActivityAsync("ReserveStock", order);
try
{
await ctx.CallActivityAsync("ChargeCard", order);
}
catch (TaskFailedException)
{
await ctx.CallActivityAsync("ReleaseStock", order); // compensation
return "payment-failed";
}
// Durable timer: costs nothing while waiting, survives restarts.
await ctx.CreateTimer(ctx.CurrentUtcDateTime.AddHours(1), CancellationToken.None);
await ctx.CallActivityAsync("SendShippingReminder", order);
return "done";
}
[Function("StartFulfilment")]
public static async Task<HttpResponseData> Start(
[HttpTrigger(AuthorizationLevel.Function, "post")] HttpRequestData req,
[DurableClient] DurableTaskClient client)
{
var order = await req.ReadFromJsonAsync<OrderMessage>();
string id = await client.ScheduleNewOrchestrationInstanceAsync(nameof(FulfilOrder), order);
return await client.CreateCheckStatusResponseAsync(req, id); // 202 + status URLs
}The HTTP starter returns 202 with status URLs at once, so the request stays inside 230 seconds while the workflow runs for hours. On Flex Consumption, Azure Storage and the Durable Task Scheduler are the supported providers.
Worked example: an order pipeline on Flex Consumption
Suppose a shop receives orders through an API that peaks at a few hundred requests per second during sales. The design: an HTTP function behind an API gateway validates the body and writes it to a Service Bus queue, returning 202. A Service Bus-triggered function upserts each order into Cosmos DB keyed by order id. For orders that need payment, it starts a Durable orchestration like the one above.
Sizing it: choose 2,048 MB instances, set always-ready to 2 for the HTTP group so the first requests after a quiet night skip the cold start, and deploy the queue consumer as its own app with a maximum instance count of 40 (the ceiling is set per app and applied to each scale group), because the database is provisioned for roughly 40 times 16 concurrent writes. Without that cap, a 200,000-message backlog after an outage would scale the consumer far past what Cosmos DB can absorb, turning backlog into throttling and retries. With it, the backlog drains at a known rate. All of these instances must also fit in the region's 250-core quota alongside other Flex apps.
Failure handling: Service Bus redelivers a failed message and, after the queue's maximum delivery count (10 by default), dead-letters it. The upsert makes redelivery harmless, and an alert fires on any dead-letter message.
Failure modes seen in production
- Timeout kills the neighbours. One slow execution hits
functionTimeout, the worker restarts and every other execution on that instance fails. Fix the slow dependency, set client timeouts shorter than the function timeout, and move long work to Durable Functions. - Downstream stampede. Fast scale-out multiplies connections to a database or partner API. Set the app's maximum instance count, split consumers into their own apps when they need a different ceiling, and lower per-instance concurrency.
- Startup over budget. Eager cache loads or large packages push start past the 30-second Flex limit. Make startup lazy; mount large binaries from Azure Files.
- Poison messages loop. A message that always throws is retried up to the delivery limit. Dead-letter non-retryable errors explicitly.
- Non-deterministic orchestrators. Reading the clock or calling an API inside an orchestrator breaks replay.
- Quota surprise. A load test in one app consumes the regional Flex core quota and production apps in the same subscription stop scaling.
Operating it: identity, network, deploys and telemetry
Give each app a managed identity with only the data roles it needs. On Flex and Premium, use virtual network integration outbound and private endpoints inbound where the API must not be public. Deploy a zip package built in CI; on Flex it goes to a blob container that every instance loads at start. Keep the plan type in infrastructure code.
For telemetry, keep Application Insights sampling on but exclude requests so failure counts stay exact. Track age of the oldest message, failures by type, instances per scale group and dead-letter counts. Oldest-message age is the best single alert, because it captures every way the pipeline can fall behind. Give busy production apps their own AzureWebJobsStorage account so leases and timers do not contend with other apps.
When Azure Functions is the wrong shape
Functions is a poor fit for services busy around the clock, where fixed container capacity is cheaper, for long-lived connections held by your own process, and for very heavy startup. It shines for bursty, event-shaped work: webhooks, queue consumers, file processing and scheduled jobs, where bindings remove the plumbing.
What to do next
- List every function app with its plan, OS and worker model; migrate any .NET in-process app to the isolated worker before 10 November 2026.
- For new apps choose Flex Consumption unless you need Windows, slots, or a fixed fleet; for Linux Consumption apps, plan the move before 30 September 2028.
- Set a maximum instance count on every app with queue or event consumers, derived from what downstream can absorb; split consumers into separate apps where ceilings differ.
- Use managed identity connections and Key Vault references; remove connection strings from app settings.
- Move anything that can exceed 230 seconds behind an HTTP call into Durable Functions or a queue with a 202 response.
- Alert on oldest-message age, dead-letter count, failure rate and regional core quota usage, then run a load test that kills a dependency and watch the alerts fire.