A green thread is a thread the operating system does not know about. The runtime keeps the stack, the saved registers and the run queue, and switches threads with an ordinary function call, while the kernel sees a handful of real threads doing all the work. The name comes from Sun's Green project, whose early Java virtual machine scheduled threads this way, and the idea never went away: Go goroutines, Erlang processes and Java virtual threads are green threads under other names.
Green threads were the normal way to get concurrency in the 1990s, were abandoned by Java, Linux, Solaris and Rust, and then came back as the default for high-concurrency servers. Both moves were right for their time, and the reasons are what you need to judge a runtime today. This article covers the three threading models, that history, a working runtime in C, and the hard parts: blocking, preemption, stacks and failure modes. For the JVM's version in detail, read JVM virtual threads architecture.
Three ways to map threads onto kernel threads
Every threading system answers one question: how many kernel-scheduled entities back how many program threads? The answer gives three models.
| Model | Mapping | Switch cost | Uses many cores | A blocking call stalls | Examples |
|---|---|---|---|---|---|
| 1:1 | each thread is a kernel thread | kernel switch | yes | only that thread | pthreads on Linux (NPTL), Windows threads, Java platform threads |
| N:1 | all threads on one kernel thread | user-space call | no | every thread | Java 1.1 green threads, GNU Pth, early Ruby |
| M:N | M threads on N kernel threads | user-space call | yes | one carrier, unless the runtime hands off | Go, Erlang/BEAM, Java virtual threads |
N:1 is the simplest and most limited: one core, and any blocking system call freezes every thread because the only kernel thread is asleep. 1:1 handles blocking perfectly and uses every core, but each thread costs a kernel stack and a reserved user stack, and every switch goes through the kernel. M:N keeps the cheap switches of N:1 and the multicore use of 1:1, at the price of a user-space scheduler that must cooperate with the kernel's.
Why the industry abandoned green threads, then rebuilt them
The first wave of green threads existed because kernels had poor or no thread support. Java 1.1 on Solaris ran all Java threads on a single kernel thread, an N:1 model, so a multithreaded Java program could not use a second processor and a blocking native call stopped the virtual machine. Sun moved to native threads and dropped green threads soon after.
Operating systems tried M:N themselves. Solaris had a two-level scheduler, and IBM's NGPT library proposed M:N threads for Linux. Both lost. NPTL arrived on Linux in 2003 with a strict 1:1 model and fast futexes, and it was simpler and fast enough; Solaris 9 made 1:1 its default. The lesson was that two schedulers which cannot see each other make bad decisions: the user-level one cannot tell that a carrier is blocked in the kernel, and the kernel cannot tell which user-level thread matters. Rust drew the same conclusion in 2014 and removed its green-thread runtime before version 1.0.
Then the workload changed. A server holding 100,000 mostly idle connections needs 100,000 concurrent waits, and 1:1 threads are expensive at that count. Event loops with callbacks solved the cost but made code hard to write. The second wave of green threads keeps the event loop's cost model and the straight-line code of threads, and fixes the old flaw in one specific way: the runtime owns all I/O. Go, Erlang and Java virtual threads park a thread that would block on a socket, register the socket with the kernel's readiness API, and run something else, so a carrier never blocks on network I/O.
The architecture of an M:N runtime
Every M:N runtime has five parts. A thread record holds a stack and saved registers. Run queues hold ready threads, usually one per carrier with work stealing between them. Carriers are kernel threads, about one per core, that loop: take a ready thread, switch to it, regain control when it yields, parks or finishes. The poller turns readiness events into ready threads. A blocking-call pool absorbs operations with no non-blocking form, such as file reads, C-library DNS and foreign calls.
Build one: a green-thread runtime with ucontext
The fastest way to understand green threads is to write a small N:1 runtime. The POSIX ucontext functions do the hard part: getcontext captures registers, makecontext points a context at a new stack and entry function, and swapcontext saves the current context and resumes another. They were removed from POSIX.1-2008 but glibc still ships them, so this compiles on Linux with gcc green.c -o green.
#define _XOPEN_SOURCE 700
#include <stdio.h>
#include <stdlib.h>
#include <ucontext.h>
#define MAX_G 64
#define STACK_SIZE (64 * 1024)
typedef enum { G_FREE, G_READY, G_DONE } gstate;
typedef struct { ucontext_t ctx; gstate state; void *stack; void (*fn)(void); } green_t;
static green_t gs[MAX_G];
static ucontext_t sched_ctx; /* where every green thread returns to */
static int current = -1;
void g_yield(void) { /* give the carrier back to the scheduler */
swapcontext(&gs[current].ctx, &sched_ctx);
}
static void trampoline(void) {
gs[current].fn();
gs[current].state = G_DONE; /* returning follows uc_link to the scheduler */
}
int g_spawn(void (*fn)(void)) {
for (int i = 0; i < MAX_G; i++) {
if (gs[i].state != G_FREE) continue;
getcontext(&gs[i].ctx);
gs[i].stack = malloc(STACK_SIZE);
gs[i].ctx.uc_stack.ss_sp = gs[i].stack;
gs[i].ctx.uc_stack.ss_size = STACK_SIZE;
gs[i].ctx.uc_link = &sched_ctx;
gs[i].fn = fn;
makecontext(&gs[i].ctx, trampoline, 0);
gs[i].state = G_READY;
return i;
}
return -1;
}
void g_run(void) { /* round-robin until no thread is ready */
int live = 1;
while (live) {
live = 0;
for (int i = 0; i < MAX_G; i++) {
if (gs[i].state != G_READY) continue;
live = 1;
current = i;
swapcontext(&sched_ctx, &gs[i].ctx);
if (gs[i].state == G_DONE) { free(gs[i].stack); gs[i].state = G_FREE; }
}
}
}
static void worker(void) {
int me = current;
for (int step = 0; step < 3; step++) {
printf("g%d step %d\n", me, step);
g_yield();
}
}
int main(void) {
g_spawn(worker);
g_spawn(worker);
g_run();
return 0;
}Run it and the workers interleave: g0 step 0, g1 step 0, g0 step 1, and so on. Trace one switch. The scheduler calls swapcontext(&sched_ctx, &gs[0].ctx), which stores its own registers, including stack pointer and return address, in sched_ctx and loads g0's. The CPU now runs on g0's malloc'd stack. When g0 calls g_yield, the reverse happens and the scheduler's call simply returns. When a worker finishes, uc_link returns control to the scheduler, which frees the stack once it is no longer running on it.
This toy has every limitation of the N:1 era: one core, starvation by any thread that never yields, a socket read that blocks the only kernel thread, and a fixed stack with no guard page. The rest of this article covers how production runtimes remove each limit.
What a context switch actually has to save
A green-thread switch is cheap because it happens at a function call, so the calling convention has already told the compiler which registers the caller does not expect to survive. On x86-64 with the System V ABI, a switch only needs to save the stack pointer and the six callee-saved registers (rbx, rbp, r12-r15) into the old thread's record, load the new thread's values, and execute ret, which pops the new thread's return address off its own stack. Production runtimes write those fourteen moves in assembly, adding the MXCSR and x87 control words for completeness.
swapcontext saves every register and also calls sigprocmask to save and restore the signal mask, a system call on every switch, which is the main reason real runtimes avoid it. A kernel thread switch costs more again: kernel entry, a scheduler decision, and often cache and TLB disruption. Measure on your hardware, but the ordering is stable: a hand-written user switch is cheapest, then swapcontext, then a kernel switch.
Blocking: the problem every runtime must solve
A green thread that makes a blocking system call blocks its carrier, and with it every green thread queued on that carrier. Runtimes use three techniques, usually together.
- Wrap I/O in non-blocking calls. Sockets are set non-blocking. A read that returns
EAGAINregisters the descriptor with epoll, kqueue or an I/O completion port, parks the green thread and switches away. The poller makes it ready when data arrives. This covers network I/O, timers and channel operations. - Hand off when a call must block. Go lets another kernel thread take over the run queue of a carrier stuck in a system call. Java virtual threads compensate for some blocking operations by temporarily adding a carrier. Other runtimes send the call to a thread pool and park the caller.
- Forbid or detect pinning. Some states make a green thread impossible to unmount from its carrier, such as a native frame on the stack. Before JDK 24, blocking inside a
synchronizedblock also pinned a Java virtual thread; JEP 491 removed that case. Pinning turns an M:N runtime back into a 1:1 one for as long as it lasts.
The practical rule: every library you call must use the runtime's I/O or block only briefly. A C driver doing its own blocking reads from 10,000 green threads can occupy every carrier.
Preemption and fairness
Cooperative scheduling switches only at yield points such as I/O, channel operations and locks. Switches stay cheap, but a CPU-bound loop with no yield point holds its carrier forever. Runtimes differ most here.
- Erlang's BEAM counts reductions, roughly function calls, and preempts a process that uses its budget, so one busy process cannot stall the rest.
- Go was cooperative until Go 1.14 added asynchronous preemption: the runtime signals a carrier whose goroutine has run too long and switches at a safe point.
- Java virtual threads are not time-sliced. A CPU-bound virtual thread keeps its carrier until it blocks or finishes, which is why the JDK guidance is to use virtual threads for waiting-heavy work and a bounded pool of platform threads for compute.
Stacks, and a worked memory example
Stack strategy decides how many green threads fit in memory. Fixed stacks force a guess. Segmented stacks add a chunk on overflow; Go dropped them in Go 1.3 because a hot loop crossing a chunk boundary allocated and freed a segment on every call. Go now starts each goroutine with a small contiguous stack and copies it to a larger one on overflow. Java virtual threads keep their frames on the heap while unmounted.
Now size a server holding 100,000 idle connections, one thread waiting on each. With 1:1 threads on Linux, each thread has a kernel stack, usually 16 KiB on x86-64, so the kernel alone holds about 1.5 GiB. Each user stack and its guard page are separate memory mappings, and the default vm.max_map_count of 65,530 stops you near 32,000 threads unless you raise it, along with kernel.threads-max and kernel.pid_max. With goroutines at a 2 KiB starting stack, 100,000 waits cost about 195 MiB of ordinary heap and no extra kernel objects. 1:1 can reach 100,000 with tuning; the green-thread version simply costs memory in proportion to what each wait uses.
Failure modes
- Carrier starvation. CPU-bound or pinned green threads leave no carrier for ready work; latency spikes while CPU looks normal. Move compute to a bounded pool.
- Hidden blocking. A library does blocking I/O the runtime cannot see; throughput collapses near the carrier count. Replace it or wrap it in the blocking pool.
- Thread-local assumptions. Code that caches large objects per thread allocates one per green thread, and there may be a million of them. See ThreadLocal architecture for the safer alternatives.
- Unbounded concurrency. Cheap threads make it easy to start 50,000 database queries at once. Bound fan-out with a semaphore at each scarce resource.
- Stack overflow without a guard page. Hand-built runtimes with malloc'd stacks corrupt adjacent memory instead of crashing. Use
mmapwith a protected guard page.
Trade-offs
| Concern | 1:1 kernel threads | M:N green threads |
|---|---|---|
| Thread count | thousands to tens of thousands with tuning | hundreds of thousands to millions |
| Switch cost | kernel entry and scheduler | user-space register swap |
| Blocking calls | always safe | safe only through the runtime's I/O |
| CPU fairness | kernel time-slicing | depends on the runtime's preemption |
| Native interop | direct | pins or hands off a carrier |
| Tooling | every OS tool works | needs runtime-aware tools |
Choose green threads when concurrency is mostly waiting on networks and you can audit the libraries that do I/O. Choose a core-sized pool of kernel threads for compute and native code. Most systems use both.
What to do next
- Compile and run the ucontext runtime above, then add a third worker that never yields and watch it starve the others.
- Replace
swapcontextwith a hand-written switch and anmmap'd stack with a guard page, and measure switches per second before and after. - List every library in one of your services that performs I/O and confirm each one goes through your runtime's I/O layer or a blocking pool.
- Enable your runtime's pinning or blocking diagnostics, such as JFR pinning events or Go goroutine profiles, in staging.
- Put a semaphore sized to each scarce downstream resource before raising concurrency.
- Read coroutines to see the stackless alternative, then Java virtual threads in production for a full runbook on one modern M:N runtime.