Rust's async fn does nothing on its own. Calling one returns a value, a future, that sits inert until something polls it. The language defines what a future is and how to wake one up; it deliberately does not ship the thing that runs them. That job belongs to a runtime library, and in practice usually to Tokio. Most confusing behaviour in async Rust, from a server that stalls under load to a timeout that never fires, comes from not having a picture of what that runtime is doing.

This page builds that picture from the bottom: the Future trait and its state machines, wakers, the I/O and timer reactor, Tokio's multi-threaded scheduler with its per-worker queues and LIFO slot, cooperative budgeting, the blocking thread pool, and cancellation by drop. It ends with a worked diagnosis of a stalled service and a checklist. Generic work stealing is covered in Work-Stealing Scheduler Architecture and is only summarised here.

The picture in one diagram

Hold this diagram in mind for the rest of the page. A small number of worker threads each run a loop: take a ready task, poll it until it returns Pending or Ready, repeat. Tasks that are not ready sit nowhere at all; they are parked inside the I/O driver or timer driver, which own the operating system's readiness mechanism. When a socket becomes readable or a deadline passes, the driver calls the task's waker, which puts the task back on a run queue. Work that would block a thread is shipped to a separate, much larger pool so the workers keep cycling.

Async worker threads (default: one per core)Worker 1LIFO slot + local queueWorker 2LIFO slot + local queueWorker Nlocal queue, 256 tasksGlobal inject queuespawns from outside, overflowI/O driverepoll / kqueue / IOCPTimer driversleep, timeoutBlocking poolspawn_blocking, max 512Task = pinned future + headerpolled by a workerWakerre-queues the taskstealreadypollwakehand off
Tokio's multi-thread runtime: workers poll tasks from local queues, the drivers wake tasks when the OS reports readiness, and blocking work goes to a separate pool.

Futures are state machines that get polled

A future is anything that implements one method:

pub trait Future {
    type Output;
    fn poll(self: Pin<&mut Self>, cx: &mut Context<'_>) -> Poll<Self::Output>;
}
// enum Poll<T> { Ready(T), Pending }

The compiler turns each async fn into an anonymous type that implements this trait: a state machine with one variant per .await point, holding the local variables that are live across that point. Polling resumes from the last state, runs until it hits a sub-future that returns Pending, stores where it was, and returns Pending itself. Because the state machine may contain references into itself, it must not move in memory once polling starts; that is what Pin<&mut Self> promises.

Two rules follow, and almost every runtime bug breaks one of them. A future that returns Pending must have arranged to be woken, or it will never be polled again. And poll must return quickly; a future that computes for 200 ms or calls a blocking function inside poll holds the worker thread for that long, and every other task queued on that worker waits.

The smallest possible executor makes the contract concrete. It needs nothing beyond the standard library:

use std::future::Future;
use std::pin::pin;
use std::sync::Arc;
use std::task::{Context, Poll, Wake, Waker};
use std::thread::{self, Thread};

struct ThreadWaker(Thread);
impl Wake for ThreadWaker {
    fn wake(self: Arc<Self>) { self.0.unpark(); }   // waking = unpark the thread that is waiting
}

fn block_on<F: Future>(fut: F) -> F::Output {
    let mut fut = pin!(fut);                          // pin it: it must not move after the first poll
    let waker = Waker::from(Arc::new(ThreadWaker(thread::current())));
    let mut cx = Context::from_waker(&waker);
    loop {
        match fut.as_mut().poll(&mut cx) {
            Poll::Ready(v) => return v,
            Poll::Pending => thread::park(),          // sleep until some leaf future calls wake()
        }
    }
}

Everything Tokio adds is about doing this for hundreds of thousands of futures on a few threads: replacing "park this thread" with "put this task on a queue", and supplying leaf futures (sockets, timers, channels) that know how to register a waker with the operating system.

Wakers and the reactor

Leaf futures are where waiting actually happens. When a TcpStream read returns Pending, Tokio has registered the socket with the I/O driver, which wraps epoll on Linux, kqueue on BSD and macOS, and IOCP on Windows (through the mio crate), and stored the task's waker against that registration. The driver is polled by the workers themselves: a worker with nothing to do blocks in the OS readiness call, and a busy worker checks the driver periodically. Tokio's event_interval setting, 61 scheduler ticks by default, bounds how many tasks a worker runs before it checks for I/O and timer events.

Timers work the same way: tokio::time::sleep registers a deadline in a hierarchical timer wheel, and the driver fires wakers whose deadlines have passed. The consequence worth remembering is that both drivers only make progress when a worker gets back to them. A worker stuck inside one long poll delays I/O and timers for everything scheduled on it, which is why a timeout can fire late or appear not to fire at all.

Tokio&amp;amp;amp;#x27;s multi-thread scheduler

The default multi-thread runtime starts one worker per core. Each worker owns a fixed-size local run queue that holds up to 256 tasks, plus a single LIFO slot. Tasks spawned from outside the runtime, and overflow from full local queues, go to a shared global inject queue. A worker's loop, simplified:

loop {
    tick += 1;
    let task = if tick % global_queue_interval == 0 {
        inject_queue.pop().or_else(|| local.pop())      // fairness: don't starve the global queue
    } else {
        lifo_slot.take().or_else(|| local.pop()).or_else(|| inject_queue.pop())
    };
    let task = task.or_else(|| steal_half_from_random_sibling());
    match task {
        Some(t) => t.poll(),                            // runs until Pending/Ready or budget is spent
        None => park_on_io_driver_or_sleep(),           // also where I/O and timers are processed
    }
    if tick % event_interval == 0 { poll_drivers_without_blocking(); }
}

The LIFO slot is an optimisation for message passing. When task A wakes task B, for example by sending on a channel, B goes into A's worker's LIFO slot and runs next, while its data is still in that core's cache. Without it, B would go to the back of the queue. Tokio limits how many times in a row the LIFO slot is used so two tasks ping-ponging messages cannot monopolise a worker. Stealing takes half of a sibling's local queue when a worker runs dry, which is the standard approach explained in the work-stealing article. Go's scheduler has the same broad shape, described in Goroutines at Scale; the key difference is that goroutines can be preempted and Rust tasks cannot.

Cooperative scheduling and the task budget

Rust tasks are cooperative: the runtime cannot interrupt a poll. That creates a subtle starvation case even for well-behaved code. A task looping over a channel that always has a message ready never returns Pending, so it never yields, and everything else on that worker waits. Tokio defends against this with a per-task budget. Each time a task is polled it gets a budget, 128 units in the current source, and Tokio's own resources (channel receives, socket reads, and similar) consume a unit each. When the budget is exhausted, those resources return Pending even if they are ready, and immediately wake the task, forcing it back through the scheduler.

Two practical points. The budget only covers Tokio's resources; a loop over your own CPU work or a third-party future that never touches them is not protected, so long computation needs explicit tokio::task::yield_now().await points or, better, a move off the async workers. And when you write your own leaf future or a busy select! loop, test it under load for fairness, not only for correctness.

Blocking work: spawn_blocking and block_in_place

Some work cannot be made async: a synchronous database driver, file system calls on platforms without an async file API, compression of a 50 MB buffer. Tokio offers two escape hatches.

// 1. Ship the blocking closure to the blocking pool; await its result.
let digest = tokio::task::spawn_blocking(move || sha256_of_file(&path)).await??;

// 2. Turn the current worker into a blocking thread for this call (multi-thread runtime only;
//    it panics on a current_thread runtime). Other tasks on this worker move elsewhere.
let rows = tokio::task::block_in_place(|| legacy_client.query_sync(&sql))?;

// CPU-heavy work belongs on a sized pool (e.g. rayon), bridged back with a oneshot channel.
let (tx, rx) = tokio::sync::oneshot::channel();
rayon::spawn(move || { let _ = tx.send(render_report(data)); });
let report = rx.await?;

The blocking pool grows on demand up to max_blocking_threads, 512 by default, and idle threads exit after thread_keep_alive, 10 seconds by default. It is sized for threads that mostly wait, not for CPU work: 512 compute-bound closures on a 16-core machine just thrash. Use it for blocking I/O, and a separate fixed-size pool for computation. The same split between cheap concurrency and real CPU parallelism appears on the JVM in Java Virtual Threads in Production.

Cancellation is drop

In Rust, cancelling a future means dropping it. There is no exception and no flag; the state machine is destroyed at whichever .await it last stopped on, and destructors run for whatever it held. tokio::time::timeout cancels its inner future this way when the deadline wins, and tokio::select! drops every branch that did not complete.

This is powerful and easy to misuse. A future dropped half way through writing a frame to a socket leaves a half-written frame. Code is cancel-safe if dropping it at any await point loses no data; Tokio documents cancel safety per method, and select! loops should only use cancel-safe operations in their branches. Note also that dropping a JoinHandle does not cancel the spawned task, it detaches it; call abort() for that, or use a JoinSet, which aborts its tasks when dropped. Structured lifetimes for groups of tasks are the subject of Structured Concurrency.

Worked example: a gateway that stalls at peak

Worked example. An HTTP gateway on an 8-core host proxies requests to a backend and, for some routes, signs the response body. Under normal load p99 latency is 15 ms. At peak it jumps to 900 ms, health checks time out, and the instance is restarted, even though CPU sits at 40%.

The investigation follows the picture. CPU at 40% with huge latency means tasks are waiting for a worker, not for CPU in aggregate. The runtime metrics show a few workers with long busy durations and growing local queues while others idle; tokio-console shows the signing task with polls lasting 60 to 120 ms. The signing code called a synchronous HSM client library inside an async fn. Each such call held a worker for the whole round trip, during which that worker's queue, its LIFO slot and the I/O events it would have processed all waited. Health checks queued behind it missed their deadline.

// Before: blocks a runtime worker for the full HSM round trip.
async fn sign(body: Bytes) -> Result<Signature> {
    Ok(hsm.sign_sync(&body)?)
}

// After: the blocking call runs on the blocking pool; the worker keeps polling other tasks.
async fn sign(hsm: Arc<HsmClient>, body: Bytes) -> Result<Signature> {
    let sig = tokio::task::spawn_blocking(move || hsm.sign_sync(&body)).await??;
    Ok(sig)
}

After the change, peak p99 dropped back to about 20 ms. The team then added a semaphore limiting concurrent HSM calls to what the device can serve, so a slow HSM would queue requests visibly instead of piling hundreds of threads into the blocking pool, and added a CI lint that flags known blocking calls inside async code.

Trade-offs: multi-thread or current-thread

Tokio offers two schedulers. The multi-thread runtime, the default for #[tokio::main], spreads tasks across workers and requires spawned futures to be Send, because a task may be resumed on a different thread after any await. That is why holding a std::sync::MutexGuard or an Rc across an .await fails to compile in a spawned task. The current-thread runtime, #[tokio::main(flavor = "current_thread")], runs everything on one thread: no stealing, no cross-thread synchronisation, and with a LocalSet it can run non-Send futures. It suits CLIs, tests, and designs that run one runtime per core with sharded state, but any blocking call stalls the whole runtime.

Failure modes

  • Blocking inside async. Synchronous I/O, std::thread::sleep, heavy CPU in poll. Symptom: latency spikes with moderate CPU. Fix: spawn_blocking or a compute pool.
  • Lost wake-ups in hand-written futures. Returning Pending without storing the waker. Symptom: tasks hang forever with no error. Fix: always register cx.waker() before returning Pending.
  • Unbounded spawning. Spawning a task per message with no limit turns backpressure into memory growth. Fix: bounded channels and semaphores; see Reactive Streams Backpressure.
  • Holding a lock across await. A std mutex held across an await can deadlock or block a worker; use Tokio's async mutex only when you truly must hold it across awaits, and keep critical sections short otherwise.
  • Cancel-unsafe select loops. Data lost when a branch is dropped. Fix: only cancel-safe operations in select!, or pin the in-progress future outside the loop.
  • Nested runtimes. Calling Runtime::block_on from inside a runtime panics. Keep one runtime and pass handles.

What to do next

  1. Draw the picture for your service: how many workers, which tasks are long-lived, and where blocking calls happen.
  2. Search the codebase for synchronous I/O, sleeps and heavy computation inside async functions and move them to spawn_blocking or a compute pool.
  3. Enable tokio-console in a staging build and look for tasks with long poll durations.
  4. Export runtime metrics (worker busy time, queue depths) to your dashboards.
  5. Put bounded channels and semaphores in front of every fan-out and every blocking resource.
  6. Review every select! loop for cancel safety and every spawned task for an owner that aborts it.
  7. Pick the runtime flavour deliberately: multi-thread by default, current_thread for CLIs, tests and shard-per-core designs.
Key takeaway: An async Rust program is a set of state machines that do nothing until a runtime polls them. Tokio's workers poll ready tasks from per-worker queues, steal when idle, and rely on the I/O and timer drivers to wake waiting tasks. Because scheduling is cooperative, any long poll or blocking call stalls everything on that worker, so move blocking and CPU-heavy work off the async threads, bound concurrency, treat cancellation as drop, and watch poll durations and queue depths in production.