A coroutine is a function that can stop in the middle, hand control back to someone else, and later continue from exactly where it stopped with its local variables intact. The idea is old. Melvin Conway described it in 1963, decades before async/await. It matters now because almost every modern runtime uses some form of it to run very large numbers of concurrent tasks on a few operating system threads.

The conceptual case for coroutines, cheap suspension and the never-block rule, is covered in Coroutines: cooperative suspension for concurrency. This article goes one level down, into how suspension is actually implemented. It explains what state a coroutine must save, the two families of implementation (stackless state machines and stackful coroutines with their own stacks), builds a working scheduler from Python generators in under fifty lines, shows what asyncio adds on top, and then covers the failure modes and trade-offs that follow from each design. The scheduler and asyncio examples were run on CPython 3.13; the state-machine lowering is pseudocode.

What a coroutine must save

An ordinary function call is strictly nested. The caller pushes a frame on the thread's stack, the callee runs to completion, and its frame is popped. Nothing about the callee survives the return. A coroutine breaks that nesting: it can return control while its frame must survive, so that a later call resumes it. To make that possible, an implementation has to save three things at every suspension point:

  1. The resume point. Which suspension point it stopped at, so the next resume jumps to the instruction after it.
  2. Live local state. Every variable and temporary that is used after the suspension point.
  3. The path back. Who resumes it and where a result or exception should be delivered.

There are two places to keep that state. Either the compiler gathers only the live variables into a heap object, or the runtime gives the coroutine an entire separate stack and simply leaves the frames on it while another coroutine runs. These are the stackless and stackful designs, and almost every behaviour difference between, say, Python's asyncio and Go's goroutines follows from that one choice.

Stackless coroutines: the compiler builds a state machine

Stackless: frame object on the heapStackful: a stack per coroutineThread stackscheduler loop frame onlyresume(frame)Coroutine Astate=2, locals x, yCoroutine Bstate=1, locals nSuspend = return to looponly at await in this functionCallers must also be asyncevery frame up the chainSchedulerswap stack pointer and registersStack Amain, handler, parse, readStack Bmain, worker, sleepSuspend from any depthno async keyword neededCost: stack memory per tasksmall and growable in Go
Left: a stackless coroutine is one heap object per async function, resumed by the loop. Right: a stackful coroutine keeps a whole call chain on its own stack, and the scheduler switches stacks.

In the stackless design, the compiler rewrites each async function into a state machine. Every suspension point becomes a numbered state, every variable live across a suspension becomes a field, and the body becomes a resume method that switches on the state number. Here is the lowering of a two-await function, in Python-like pseudocode:

# source
async def fetch_both(a, b):
    x = await get(a)
    y = await get(b)
    return x + y

# what the compiler effectively builds
class FetchBoth:
    state = 0
    def resume(self, sent):
        if self.state == 0:
            self.pending = get(self.a); self.state = 1
            return SUSPEND(self.pending)
        if self.state == 1:
            self.x = sent                      # result of the first await
            self.pending = get(self.b); self.state = 2
            return SUSPEND(self.pending)
        if self.state == 2:
            return DONE(self.x + sent)

The languages differ in the details but share the shape. CPython keeps the whole frame object for a generator or coroutine alive on the heap and resumes it with send; an empty coroutine object measured with sys.getsizeof is under 200 bytes plus its frame, which undercounts but shows the order of magnitude. C# and JavaScript compile async functions into similar state machines. Rust turns each async fn into an anonymous type implementing Future, whose poll method advances the state machine; because the state can hold references into itself, it must be pinned in memory before it is polled. C++20 coroutines allocate a coroutine frame, usually on the heap unless the compiler can prove it is safe to elide the allocation, and the library author supplies the promise type that decides what suspension means. Kotlin compiles suspend functions with an extra hidden continuation parameter and a state label.

The consequence that everyone meets is that a stackless coroutine can only suspend in its own body. If fetch_both calls an ordinary function that wants to wait, that function has no state machine and cannot suspend, so it has to become async too, and so do its callers. That is the origin of function colouring: async and sync functions form two worlds, and calling across the boundary needs an explicit bridge such as running the loop or offloading to a thread.

Stackful coroutines: a stack per task

A stackful coroutine gets its own stack. Suspending saves the CPU registers, including the stack pointer and instruction pointer, and loads another coroutine's; resuming does the reverse. Because the entire call chain stays on the coroutine's stack, any function at any depth can suspend, and nothing needs an async keyword.

Go's goroutines are the best-known example. Each starts with a small stack that the runtime grows by copying to a larger allocation when needed, and the scheduler multiplexes goroutines onto a pool of OS threads; blocking-looking calls such as a network read park the goroutine and let the thread run another. Since Go 1.14 the runtime can also preempt a goroutine that runs a long loop without function calls, so goroutines are not purely cooperative. Lua's coroutines are stackful as well, and are explicitly resumed and yielded by the program. Java's virtual threads take a hybrid route: when a virtual thread blocks, its frames are copied off the carrier thread into heap objects and copied back on resume, which gives stackful semantics without reserving a full stack per thread. Goroutines at scale and Java virtual threads in production cover those two runtimes in depth.

The price is memory and runtime machinery. Each stackful coroutine needs at least a small stack and a way to grow it, and the runtime must be able to intercept every blocking operation, or a blocking system call pins the underlying thread.

Worked example: a scheduler from Python generators

The quickest way to understand a coroutine scheduler is to build one. Python generators are stackless coroutines: yield suspends and send resumes. The scheduler below keeps a ready queue and a heap of sleeping tasks. A task suspends by yielding a request object; the scheduler interprets the request, parks the task, and resumes it later.

import heapq, time
from collections import deque

class Sleep:
    def __init__(self, seconds):
        self.seconds = seconds

class Scheduler:
    def __init__(self):
        self.ready = deque()          # tasks that can run now
        self.sleeping = []            # heap of (wake_time, seq, task)
        self.seq = 0

    def spawn(self, gen):
        self.ready.append((gen, None))

    def run(self):
        while self.ready or self.sleeping:
            if not self.ready:        # nothing runnable: wait for the next timer
                wake, _, gen = heapq.heappop(self.sleeping)
                time.sleep(max(0, wake - time.monotonic()))
                self.ready.append((gen, None))
            while self.sleeping and self.sleeping[0][0] <= time.monotonic():
                _, _, gen = heapq.heappop(self.sleeping)
                self.ready.append((gen, None))
            gen, value = self.ready.popleft()
            try:
                request = gen.send(value)   # resume until the next yield
            except StopIteration:
                continue                    # coroutine finished
            if isinstance(request, Sleep):
                self.seq += 1
                heapq.heappush(self.sleeping,
                               (time.monotonic() + request.seconds, self.seq, gen))
            else:
                self.ready.append((gen, None))  # plain yield: go to the back

def worker(name, delay, steps):
    for i in range(steps):
        print(f"{time.monotonic() - T0:4.2f}s {name} step {i}")
        yield Sleep(delay)
    print(f"{time.monotonic() - T0:4.2f}s {name} done")

T0 = time.monotonic()
s = Scheduler()
s.spawn(worker("fast", 0.1, 3))
s.spawn(worker("slow", 0.25, 2))
s.run()

Running it prints an interleaved trace. Both workers start at 0.00s; fast runs its steps at 0.10s and 0.20s; slow wakes at 0.25s; fast finishes at 0.30s and slow at 0.50s. The whole program took half a second on one thread, which is the sum of the longest chain, not of all the sleeps. Walk the data flow once: send resumes a generator frame on the heap; the frame runs until yield Sleep(...); the request object flows back to the scheduler, which files the frame under its wake time; nothing runs that frame again until the heap says so. That is the entire mechanism. A real event loop replaces time.sleep with a call that waits for timers and I/O readiness together.

What asyncio adds to the toy

The same program in asyncio looks almost identical, and produces the same timeline:

import asyncio, time

async def worker(name, delay, steps):
    for i in range(steps):
        print(f"{time.monotonic() - T0:4.2f}s {name} step {i}")
        await asyncio.sleep(delay)
    print(f"{time.monotonic() - T0:4.2f}s {name} done")

async def main():
    async with asyncio.TaskGroup() as tg:     # Python 3.11+
        tg.create_task(worker("fast", 0.1, 3))
        tg.create_task(worker("slow", 0.25, 2))

T0 = time.monotonic()
asyncio.run(main())

What asyncio adds is everything around the toy. The loop waits on timers and sockets together through the operating system's readiness API via the selectors machinery, or on Windows through I/O completion ports. A Task wraps a coroutine and drives it; await on a Future suspends the task until someone sets the future's result, and the loop schedules the task's next step as a callback. TaskGroup adds structured concurrency: if one child fails, the others are cancelled and the error propagates out of the async with block, which is the discipline described in structured concurrency.

Failure modes

Each design has characteristic bugs. These are the ones that reach production.

  • A blocking call stalls everything. In a stackless runtime, a synchronous requests.get or a heavy CPU loop inside a coroutine holds the loop's only thread. asyncio's debug mode logs any callback that runs longer than loop.slow_callback_duration, 0.1 seconds by default; use it in tests, and move blocking work out with asyncio.to_thread or an executor.
  • The forgotten await. Calling an async function without awaiting it creates a coroutine object and runs nothing. Python warns that the coroutine was never awaited, but only when the object is garbage collected, which may be far from the bug.
  • Tasks that vanish. The event loop keeps only weak references to tasks, so a task created with create_task and not stored anywhere can be garbage collected before it finishes. Keep a reference, or use a TaskGroup.
  • Swallowed cancellation. Cancellation is delivered as CancelledError at the next suspension point. A broad exception handler that catches it and carries on makes the task impossible to stop.
  • Unbounded fan-out. Coroutines are cheap, the services they call are not. Spawning one per item for a million items opens a million requests; bound concurrency with a semaphore or a fixed worker pool.
  • Context that does not follow. Thread-local storage is per OS thread, and many coroutines share one thread, so request-scoped data leaks between them. Use contextvars in Python, or the runtime's coroutine context elsewhere.
  • Pinned carriers in stackful runtimes. A stackful coroutine that enters a blocking system call or native code the runtime cannot intercept occupies its OS thread until the call returns.

Trade-offs: stackless versus stackful

PropertyStackless (async/await)Stackful (goroutines, Lua, virtual threads)
Where suspension can happenOnly at await in an async functionAnywhere in the call chain
Memory per suspended taskOnly live variables, often a few hundred bytesA stack, small and growable, or copied frames
Function colouringYes: async spreads up the call graphNo: ordinary functions can block
Suspension points visible in codeYes, every await is markedNo, any call may switch
Integration with existing blocking librariesNeeds async versions or thread offloadRuntime must intercept blocking calls
Cost of a switchA function return and callSaving and loading registers and stack pointer

Visible suspension points are the underrated advantage of the stackless model: between two awaits, no other coroutine on the same loop can run, so shared state touched only between awaits needs no lock. Stackful models give that up in exchange for not splitting the world in two. Neither is free; pick the one whose failure modes your team can see and test for. For the history of how languages arrived here, see the history of green threads.

What to do next

To turn this into working knowledge:

  1. Run the generator scheduler above, then add a Channel request type so one task can wait for a value another sends.
  2. Rewrite one blocking client in your codebase with asyncio and run it with PYTHONASYNCIODEBUG=1 to catch slow callbacks.
  3. Audit every create_task call and make sure each task is referenced or owned by a TaskGroup.
  4. Search for broad exception handlers inside coroutines and confirm they re-raise CancelledError.
  5. Put a semaphore on every fan-out loop and choose its size from the downstream service's capacity.
  6. If you use Go or Java virtual threads, find the calls that pin carrier threads, such as native calls, and measure them under load.
Key takeaway: A coroutine must save its resume point, its live locals and a path back to whoever resumes it. Stackless runtimes such as asyncio, Rust, C# and C++20 compile async functions into heap state machines that can only suspend at await, which makes suspension visible but colours functions. Stackful runtimes such as Go, Lua and Java virtual threads give each task a stack so any call can suspend, at the cost of per-task memory and runtime interception of blocking calls. Either way, never block the loop, keep task references, re-raise cancellation and bound fan-out.