Kotlin coroutines, Go goroutines and Rust async tasks all let you run a hundred thousand concurrent operations on a handful of threads. That shared result hides three quite different architectures, and they behave differently when something goes wrong. Cancellation does not reach a goroutine unless the goroutine asks for it. A blocking JDBC call can quietly eat a Kotlin dispatcher. In Rust, dropping a future halfway through an await can leave a half-written buffer behind.
This article compares them as architectures. The internals of each are covered elsewhere on this site: stackless versus stackful coroutines, the Go scheduler at scale and the Rust async runtime picture. Here we write one realistic program in all three languages, then use it to answer the questions that decide your design: what is saved when a task pauses, who decides what runs next, what a blocking call costs, and how a stop signal travels. At the end you should be able to pick a model for a service and know which failures to test for.
Three questions that define a concurrency model
Any concurrency model answers three questions. Where does a paused task's state live? Who picks the next task to run? What happens when a task makes a call that blocks the OS thread? OS threads answer: in a kernel-managed stack, the kernel scheduler, and nothing special because only that thread waits. Everything else in this article is a different set of answers.
Kotlin compiles every suspend function into a state machine. The compiler adds a hidden Continuation parameter and stores the locals that live across a suspension point in a heap object. A coroutine is that chain of continuation objects, plus a Job for its lifecycle. A CoroutineDispatcher decides which thread resumes it. The design is stackless: a suspension can only happen at a call to another suspend function, so the compiler knows every place a task can pause.
Go gives every goroutine its own small, growable stack. The runtime copies the stack to a larger allocation when it fills. Because the stack is real, any function can block, and the runtime can park the goroutine wherever it is. The scheduler multiplexes goroutines (G) onto OS threads (M) through processors (P), where the number of Ps is GOMAXPROCS. Since Go 1.14 the runtime can also preempt a goroutine stuck in a tight loop by sending its thread a signal.
Rust turns every async fn into a type that implements Future: an enum of states whose size is fixed at compile time. Nothing runs until an executor calls poll. When a future cannot make progress it returns Pending and registers a Waker, and the executor polls it again once woken. The language provides only the trait; the scheduler is a library you choose, usually Tokio.
One program, three languages
The test program is the shape most backend services contain: fetch N URLs concurrently, at most 16 in flight, give up on all of them after two seconds, and if any fetch fails, cancel the others and return the error. It exercises fan-out, a concurrency limit, a deadline and failure propagation.
Kotlin, using a suspending HTTP client (Ktor with expectSuccess = true, so non-2xx responses throw):
suspend fun fetchAll(client: HttpClient, urls: List<String>): List<String> {
val limit = Semaphore(16) // kotlinx.coroutines.sync
return withTimeout(2.seconds) { // throws TimeoutCancellationException
urls.map { url ->
async { limit.withPermit { client.get(url).bodyAsText() } }
}.awaitAll() // first failure cancels the scope
}
}Structured concurrency does most of the work here. withTimeout opens a child scope, and every async is a child of it. If one child throws, the scope cancels its siblings and rethrows. If the deadline passes, the whole subtree is cancelled. No task outlives the call.
Go, with errgroup from golang.org/x/sync:
func fetchAll(parent context.Context, urls []string) ([]string, error) {
ctx, cancel := context.WithTimeout(parent, 2*time.Second)
defer cancel()
g, ctx := errgroup.WithContext(ctx) // ctx is cancelled on the first error
g.SetLimit(16)
out := make([]string, len(urls))
for i, url := range urls { // per-iteration variables since Go 1.22
g.Go(func() error {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return err
}
resp, err := http.DefaultClient.Do(req)
if err != nil {
return err
}
defer resp.Body.Close()
if resp.StatusCode >= 400 {
return fmt.Errorf("%s: %s", url, resp.Status)
}
b, err := io.ReadAll(resp.Body)
out[i] = string(b)
return err
})
}
if err := g.Wait(); err != nil {
return nil, err
}
return out, nil
}The shape matches, but cancellation is a value you pass around. If you forget NewRequestWithContext and call http.Get(url), the request ignores the deadline. The group then waits for it and the function takes as long as the slowest server.
Rust, on Tokio with reqwest and anyhow:
async fn fetch_all(urls: Vec<String>) -> anyhow::Result<Vec<String>> {
let client = reqwest::Client::new();
let limit = Arc::new(Semaphore::new(16)); // tokio::sync::Semaphore
let mut set = JoinSet::new();
for (i, url) in urls.into_iter().enumerate() {
let (client, limit) = (client.clone(), limit.clone());
set.spawn(async move {
let _permit = limit.acquire_owned().await?;
let body = client.get(&url).send().await?
.error_for_status()?.text().await?;
anyhow::Ok((i, body))
});
}
let mut out = vec![String::new(); set.len()];
let collect = async {
while let Some(joined) = set.join_next().await {
let (i, body) = joined??; // JoinError, then task error
out[i] = body;
}
anyhow::Ok(())
};
tokio::time::timeout(Duration::from_secs(2), collect).await??;
Ok(out)
}Here cancellation comes from ownership. If the timeout fires or ? returns early, set is dropped, and dropping a JoinSet aborts every task still in it. Each aborted task's future is dropped at whatever await it was parked on.
Cancellation: exception, context or drop
The three programs look alike, but the way a stop signal reaches running code differs in each language. This is the difference you will meet most often in incident reviews.
| Kotlin | Go | Rust (Tokio) | |
|---|---|---|---|
| Signal | Job cancelled; next suspension point throws CancellationException | ctx.Done() closes; code must check it | Future dropped at its current await |
| Reaches CPU loops? | Only if the loop calls ensureActive() or yield() | Only if the loop checks ctx.Err() | Only at the next await; a busy loop never yields |
| Cleanup | finally blocks run; suspending cleanup needs withContext(NonCancellable) | defer runs when the goroutine returns | Drop impls run; there is no async drop |
| Classic bug | Catching Exception swallows the cancellation | Goroutine blocked on a channel nobody reads: a leak | Not cancel-safe: dropped mid-write loses data |
Kotlin and Rust cancel by default: a parent can stop children without their cooperation, at least at suspension points. Go cancels by agreement: a goroutine cannot be killed from outside, so every blocking operation in a long-lived goroutine needs a select that includes ctx.Done(). The Go approach is more verbose but has no surprise exits. The Rust approach is the most abrupt: code after an await simply never runs. That is why Tokio documents which methods are cancel-safe, and why using select! on a non-cancel-safe future in a loop is a known source of lost messages.
Blocking calls and the carrier threads
Every M:N system has to deal with code that blocks the OS thread: a synchronous database driver, file I/O, a JNI or cgo call, or a long CPU computation. The three runtimes respond very differently.
Go handles it for you. Network I/O goes through the runtime's netpoller, so a goroutine waiting on a socket is parked and its thread is reused. When a goroutine enters a blocking system call, the runtime detaches its P and hands it to another M, creating a thread if needed, so other goroutines keep running. The cost is threads: a burst of slow file or cgo calls can push the process to thousands of OS threads. The runtime stops the process when it passes the limit set by debug.SetMaxThreads (10,000 by default).
Kotlin leaves it to you. Dispatchers.Default has as many threads as CPU cores, with a minimum of two. If a coroutine on it calls a blocking driver, it takes one of those few threads for the duration. Eight such calls on an eight-core box stall every other coroutine on the dispatcher. The fix is withContext(Dispatchers.IO) around the blocking call. That dispatcher allows up to 64 threads by default, or more on machines with more cores. Use Dispatchers.IO.limitedParallelism(n) to give a specific dependency its own bounded slice.
Rust also leaves it to you, and enforces it the least. A blocking call inside a Tokio task holds a worker thread, and the tasks queued on that worker wait. Because of work stealing it can take a while before anyone notices. Wrap blocking work in tokio::task::spawn_blocking, which runs it on a separate pool with a high default cap, or move it to a dedicated thread. A related compile-time trap: holding a std::sync::MutexGuard across an await makes the future !Send, so tokio::spawn rejects it.
Java virtual threads behave like Go here; see Java virtual threads in production.
Memory per task: a worked estimate
Memory per task decides how many concurrent connections a box can hold, so it is worth estimating rather than quoting slogans. The figures below are orders of magnitude. Measure your own workload before relying on them.
| Model | What a parked task holds | Typical cost | Grows with |
|---|---|---|---|
| OS thread | Kernel structures plus a stack reservation (often 8 MB of virtual address space on Linux, committed lazily) | Tens of KB resident at minimum, more as the stack is touched | Deepest call stack reached |
| Go goroutine | Its own stack, starting at a few KB, copied larger on demand | A few KB for idle handlers | Deepest call stack reached; never shrinks below what GC allows |
| Kotlin coroutine | One continuation object per suspended frame, holding only live locals | Hundreds of bytes to a few KB | Number of suspended frames and captured locals |
| Rust future | One struct sized for its largest state, nested futures inline | Exactly what the type says; can be large | Large locals held across await; box big ones |
A worked example: a websocket gateway with 200,000 idle connections. At 4 KB each, goroutines cost about 800 MB of stack. Kotlin continuations at around 1 KB cost about 200 MB of heap, which the garbage collector must trace. Rust futures cost whatever the connection-handler type declares, and Tokio allocates each spawned task on the heap. If that type is 2 KB because someone held a buffer across an await, you pay 400 MB. With 200,000 OS threads, the kernel scheduler becomes the limit before memory does.
Function colouring and interop
Stackless designs pay for their efficiency with function colouring. In Kotlin and Rust, a function that suspends has a different type from one that does not, and only suspending code can call it. A library is either async or sync. Bridging them takes runBlocking or block_on, both disastrous inside the async runtime.
Go has no colouring: every function can block, and any function can be called from any goroutine. That is why existing synchronous Go code scales without rewriting, and why Java chose a stackful design for virtual threads. The cost is visibility: a Go signature does not say whether a call blocks, so by convention such functions take a context.Context first.
When plain threads are the better choice
Plain OS threads are still the right choice more often than conference talks suggest:
- Few, long, CPU-bound tasks. A video encoder or a training data loader with eight workers gains nothing from M:N scheduling. Use a fixed thread pool sized to cores.
- Mostly blocking dependencies. If every call goes through a synchronous driver you cannot replace, async in Kotlin or Rust just moves the blocking into an I/O pool. You get the complexity without the scaling.
- Hard latency isolation. A thread with its own priority or CPU pinning gets kernel guarantees that a cooperative task on a shared worker cannot.
- Concurrency in the hundreds. 200 concurrent requests fit in 200 threads, with ordinary stack traces and profilers.
Failure modes
- Dispatcher starvation (Kotlin). Blocking calls on
Dispatchers.Default: throughput collapses at exactly core-count concurrent requests. Detect it with thread dumps showing allDefaultDispatcher-workerthreads in socket reads. - Goroutine leaks (Go). A sender blocks forever on an unbuffered channel after the receiver returned early. The goroutine count climbs without bound. Watch
runtime.NumGoroutine()and the goroutine pprof profile. - Worker stalls (Rust). A synchronous call or long loop inside a task delays unrelated requests;
tokio-consoleshows tasks with long poll times. - Swallowed cancellation (Kotlin).
catch (e: Exception)around a suspending call catchesCancellationException, so the coroutine keeps running after its scope died. Rethrow it, or catch narrower types. - Unbounded fan-out (all three). Without a semaphore or
SetLimit, a 50,000-item batch opens 50,000 sockets and the downstream service falls over. M:N makes spawning cheap, not the work the tasks do. See async/await patterns for bounded fan-out designs.
What to do next
- Write down your service's peak concurrent operations and what fraction is blocking I/O, non-blocking I/O and CPU. Under a few hundred, threads are a valid answer.
- Port the fetch-all program to your stack, add a deliberately slow and a deliberately failing URL, and verify that the deadline and cancel-on-first-error both work.
- Inventory every blocking dependency and confirm each runs on
Dispatchers.IO,spawn_blocking, or (in Go) a bounded worker pool. - Add a lint or review rule for your language's classic bug: broad catches in Kotlin, missing
ctxplumbing in Go, non-cancel-safe futures inselect!in Rust. - Export task counts (goroutines, active coroutines or Tokio tasks) and dispatcher or worker saturation to your dashboards, and alert on sustained growth.
- Load-test at 2x expected concurrency and capture a thread dump or task dump at peak, so you know what normal looks like before an incident.