Async/await lets one thread juggle thousands of waiting operations while the code still reads top to bottom. It is the default model for network services in Python, JavaScript, C#, Rust and Kotlin, and it is easy to write code that compiles, passes tests and then behaves badly under load: requests that run one after another when they could overlap, fan-outs that open ten thousand sockets at once, timeouts that leave work running in the background, and one blocking call that freezes every request on the server.
This article is a catalogue of the patterns that avoid those problems, with the reasoning behind each. It starts with what await actually does, because every pattern follows from that, then covers concurrent fan-out, structured task groups, bounded concurrency, timeouts and cancellation, bridging blocking code, and async streams with backpressure. Examples use Python's asyncio because its primitives are explicit, with notes on how the same ideas look in JavaScript and C#. A worked example combines them into a small service component. For the underlying coroutine mechanics see coroutines.
What await actually does
An async function does not run when you call it. It returns a coroutine (Python), a Promise that has started (JavaScript) or a Task that has started (C#). The compiler turns the function body into a state machine whose states are the points between awaits. When the code reaches await on something that is not finished, the function saves its locals, returns control to the event loop, and registers to be resumed when the awaited thing completes.
The event loop is a single thread that repeatedly takes a ready task, runs it until its next suspension, and asks the operating system's I/O selector which sockets and timers are ready. That gives three rules that every pattern below relies on. First, tasks only switch at awaits, so code between two awaits runs without interruption and needs no lock for in-memory state. Second, anything that blocks the thread without awaiting blocks every task on the loop. Third, concurrency appears only when several operations are started before any of them is awaited; writing await on each one in turn is sequential, however async the syntax looks.
Concurrent fan-out instead of accidental sequencing
The most common performance bug in async code is accidental sequencing. Each await below waits for one request to finish before the next one starts, so three 200 ms calls take 600 ms.
# Sequential: total latency is the sum.
user = await fetch_user(uid)
orders = await fetch_orders(uid)
recs = await fetch_recommendations(uid)When calls are independent, start them all and then wait for all of them. Total latency becomes the slowest call rather than the sum. In Python 3.11 and later the right tool is a TaskGroup, because it gives structured concurrency: if one task fails, the group cancels the others and the async with block raises an ExceptionGroup carrying the failure(s), so no task is left running unobserved.
import asyncio
async def load_profile(uid: str) -> dict:
async with asyncio.TaskGroup() as tg:
user_t = tg.create_task(fetch_user(uid))
orders_t = tg.create_task(fetch_orders(uid))
recs_t = tg.create_task(fetch_recommendations(uid))
# Leaving the block means every task finished successfully.
return {"user": user_t.result(), "orders": orders_t.result(), "recs": recs_t.result()}JavaScript's Promise.all([...]) and C#'s Task.WhenAll(...) give the same latency benefit, but neither cancels the siblings when one fails: the others keep running and their results are thrown away. In JavaScript, pass an AbortController signal to each call and abort it on failure; in C#, pass a CancellationToken from a linked CancellationTokenSource and cancel it. When partial results are acceptable, use Promise.allSettled or, in Python, asyncio.gather(..., return_exceptions=True) and inspect each outcome explicitly. The structured concurrency article explains why scoping task lifetimes to a block matters beyond error handling.
Bounded concurrency
Fan-out without a limit is a denial-of-service attack on your own dependencies. Starting one task per item for a list of 50,000 URLs opens 50,000 sockets, exhausts file descriptors, trips rate limits and makes every request time out together. The fix is a semaphore that caps how many tasks are inside the critical region at once.
async def fetch_all(urls: list[str], limit: int = 50) -> list[bytes | Exception]:
sem = asyncio.Semaphore(limit)
results: list[bytes | Exception] = [b""] * len(urls)
async def one(i: int, url: str) -> None:
async with sem: # at most `limit` requests in flight
try:
results[i] = await http_get(url)
except Exception as exc: # record, do not cancel the whole batch
results[i] = exc
async with asyncio.TaskGroup() as tg:
for i, url in enumerate(urls):
tg.create_task(one(i, url))
return resultsThis version still creates 50,000 task objects, which costs memory but is usually fine. For millions of items, or an unbounded input, use a fixed pool of worker tasks reading from a queue instead, which is shown in the streams section. Choose the limit from the dependency's capacity (its connection pool, its rate limit) rather than from your own CPU count: async concurrency is about how many things you are waiting on, not how many you can compute.
Timeouts and cancellation
Every await on the network needs a deadline, or one slow dependency will pin tasks forever. In Python 3.11 and later, asyncio.timeout() applies a deadline to a whole block, which is better than per-call timeouts because retries inside the block share one budget.
async def get_price(sku: str) -> Decimal:
try:
async with asyncio.timeout(0.8): # whole budget, including retries
for attempt in range(3):
try:
return await pricing_client.get(sku)
except TransientError:
await asyncio.sleep(0.05 * 2 ** attempt)
raise PricingUnavailable(sku)
except TimeoutError:
return cached_price(sku) # degrade instead of failing the pageA timeout works by cancelling the task: the awaited operation receives a CancelledError at its current await point. That has two consequences. Cleanup code in finally blocks and async with exits runs, which is how connections get returned to pools. And code must not swallow cancellation: a bare except Exception does not catch CancelledError in modern Python because it derives from BaseException, but except BaseException or a bare except: will, and then the task keeps running after its caller has given up.
Propagate cancellation through every layer. In C# that means accepting a CancellationToken in every async method and passing it down; in JavaScript, an AbortSignal. A timeout that only stops the caller from waiting, while the work continues underneath, still consumes the dependency's capacity and is how retry storms start.
Bridging blocking code
Blocking calls on the event loop are the second most common bug: a synchronous database driver, time.sleep, a large JSON parse, password hashing, or file I/O. Each one stops every task on the loop for its duration, so p99 latency for unrelated requests jumps to match the slowest blocking call.
import asyncio, hashlib
async def handle_upload(data: bytes) -> str:
# Blocking CPU work moved to a thread so the loop keeps serving other tasks.
digest = await asyncio.to_thread(lambda: hashlib.sha256(data).hexdigest())
await storage.put(digest, data) # genuinely async I/O stays on the loop
return digest
# Debug mode logs callbacks that hold the loop longer than slow_callback_duration.
loop = asyncio.get_running_loop()
loop.set_debug(True)
loop.slow_callback_duration = 0.05 # secondsasyncio.to_thread runs the function in the default thread pool. In Python, threads help with blocking I/O and with C code that releases the GIL, such as hashlib on large inputs, but pure-Python CPU work needs a process pool through loop.run_in_executor. C# has Task.Run for the same purpose; Node.js uses worker threads. The opposite bridge, calling async code from sync code, should happen once at the program's entry point with asyncio.run. Calling it deep inside a library, or blocking on .Result in C#, is the classic sync-over-async deadlock and pool-starvation recipe.
Async streams and backpressure
When input is unbounded, such as messages from a queue or lines of a large file, a fixed pool of workers reading from a bounded queue gives you bounded concurrency and backpressure at once. The producer awaits queue.put when the queue is full, so it slows down to the speed of the consumers instead of buffering without limit.
async def pipeline(source, n_workers: int = 32) -> None:
q: asyncio.Queue = asyncio.Queue(maxsize=n_workers * 4) # bounded => backpressure
async def producer() -> None:
async for msg in source: # async iterator over the input
await q.put(msg) # waits when consumers fall behind
for _ in range(n_workers):
await q.put(None) # one stop marker per worker
async def worker() -> None:
while (msg := await q.get()) is not None:
try:
await process(msg)
except Exception:
await dead_letter(msg) # isolate poison messages
async with asyncio.TaskGroup() as tg:
tg.create_task(producer())
for _ in range(n_workers):
tg.create_task(worker())The same shape appears in JavaScript with for await...of over an async iterable and in C# with IAsyncEnumerable and System.Threading.Channels. The reactive streams backpressure article covers the protocol-level version of the same idea, where demand is signalled upstream explicitly.
Worked example: a batch enrichment service
Consider an enrichment service that receives batches of up to 2,000 product IDs. For each ID it calls a catalogue API (median 40 ms, rate limit 100 concurrent requests) and a pricing API (median 30 ms, rate limit 50 concurrent), and the whole batch must answer within 2 seconds.
- Naive version. Two sequential awaits per item, items in sequence: 2,000 x 70 ms = 140 seconds. Completely unusable.
- Unbounded fan-out. 4,000 simultaneous requests: both APIs return rate-limit errors and the batch fails.
- Per-item TaskGroup plus two semaphores. Each item fetches catalogue and pricing concurrently, guarded by semaphores of 100 and 50. Pricing is the bottleneck: 2,000 calls at 50 in flight and 30 ms each is about 40 rounds, roughly 1.2 seconds. Catalogue finishes in about 20 rounds of 40 ms, 0.8 seconds, overlapping. Expected batch time is around 1.2 to 1.4 seconds.
- Deadline. Wrap the batch in
asyncio.timeout(1.9)and, per item, fall back to cached prices on a pricing timeout, so a slow pricing service degrades freshness rather than failing the batch. - Measure. Record in-flight counts per semaphore and loop lag (how late a 100 ms periodic timer fires). If loop lag rises above a few milliseconds, something is blocking the loop; find it with debug mode before raising any limit.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Latency is the sum of calls | Awaiting independent calls one by one | Start them together with a TaskGroup or gather |
| Every request slows at once | Blocking call on the event loop | Move it to to_thread or a process pool; enable debug mode |
| Dependency rate-limits or falls over | Unbounded fan-out | Semaphore or worker pool sized to its capacity |
| Work continues after a timeout | Cancellation swallowed or not propagated | Never catch BaseException; pass tokens or signals down |
Task was destroyed but it is pending warnings | Fire-and-forget task with no owner | Create tasks inside a TaskGroup or keep a reference and await it |
| Unhandled rejection or lost exception | A started task or promise nobody awaited | Await every task; use structured scopes |
| Deadlock calling async from sync | Blocking on a task from a thread the task needs | Async all the way down; one asyncio.run at the entry point |
| Memory grows under load | Unbounded queue between producer and consumer | Bounded queue so the producer waits |
Trade-offs against threads and virtual threads
Async/await is a good fit for I/O-bound services with many concurrent connections, because a waiting task costs a small heap object rather than a thread stack. It has costs: function colour (async functions can only be awaited from async functions, so the model spreads through a codebase), a separate ecosystem of async drivers, and the blocking-call hazard. Threads are simpler for modest concurrency and for libraries that are only synchronous. Java's virtual threads take a third route: blocking-style code with a cheap scheduler underneath, avoiding function colour; the virtual threads in production article covers where that model has its own pitfalls. Rust's async adds explicit runtimes and Send bounds; see the Rust async runtime picture. None of these models makes CPU-bound work faster; for that you need parallelism across cores.
What to do next
- Search your code for consecutive awaits on independent calls and convert them to a TaskGroup,
Promise.allorTask.WhenAll. - Put a semaphore or a worker pool in front of every fan-out, sized from the dependency's limits.
- Give every network await a deadline, preferably one budget per request covering retries.
- Audit exception handlers for anything that catches cancellation, and pass cancellation tokens or abort signals through every layer.
- Turn on asyncio debug mode (or the equivalent) in staging and fix every slow-callback warning.
- Measure event loop lag in production and alert when it exceeds a few milliseconds.
- Replace fire-and-forget tasks with tasks owned by a structured scope.
- Use bounded queues between producers and consumers so backpressure reaches the source.