Go makes concurrency cheap enough that a single process can hold hundreds of thousands of goroutines, and it is common to see services that start one per request, one per downstream call and one per stream. Cheap is not free. At scale goroutines cost memory, garbage-collector time and scheduler work, and the classic production incidents in Go services, memory that grows until the pod is killed or latency that rises under load, are usually caused by goroutines that were never bounded or never finished.

This article explains how the runtime schedules goroutines, what each one costs, and how to bound, observe and test them. Runtime details were checked against the Go source and the Go 1.25 and 1.26 release notes. If you come from Java, virtual threads in production covers the same problem on the JVM.

Advertisement

What a goroutine costs

A goroutine is a runtime structure (the g) plus a stack. The minimum stack is 2 KiB (stackMin = 2048 in runtime/stack.go). Since Go 1.19 the runtime picks the starting size from the average stack use it observed at recent garbage collections, so a program whose goroutines use deep stacks starts new ones larger; GODEBUG=adaptivestackstart=0 turns that off. When a function needs more stack than is left, the runtime allocates a stack twice as large, copies the frames and adjusts pointers into the stack; during garbage collection a stack using less than a quarter of its size can be shrunk.

So the floor is a few kilobytes per goroutine, and 100,000 idle goroutines need at least about 200 MiB of stack. In practice the stack is the small part. What a goroutine references, request bodies, buffers, decoded JSON, a TLS connection, is what fills the heap, and every live goroutine stack is a root the garbage collector must scan. A goroutine count is therefore a proxy for retained memory and GC work, not just for scheduler load.

How the scheduler runs them

The scheduler has three kinds of object. A G is a goroutine. An M is an operating-system thread. A P is a logical processor, a token that an M must hold to run Go code; there are GOMAXPROCS of them. Each P owns a local run queue of 256 slots plus a single runnext slot, which holds a goroutine just made runnable by the current one so that a producer and consumer pair runs back to back.

Global run queueoverflow and fairness; checked every 61st schedule tickP0runnext + local queue (256)P1runnext + local queue (256)P2 (idle)steals half of a victim queuehalf on overflowstealM0: OS threadrunning GM1: OS threadrunning GM2: in syscallP handed offNetpoller (epoll, kqueue, IOCP)parks Gs waiting on sockets; readies them in batchessysmonretakes Ps from long syscalls; preempts Gs running over 10 msready Gs
The Go scheduler. GOMAXPROCS logical processors (P) each own a local run queue; OS threads (M) execute goroutines (G) only while holding a P. Idle Ps steal work, blocking system calls hand their P to another thread, the netpoller turns socket waits into parked goroutines, and the sysmon thread enforces preemption.

When a P's local queue is full, half of it is moved to the global run queue. To stop the global queue from starving, a P checks it on every 61st scheduling tick even when it has local work. A P with nothing to do looks in the global queue, polls the network, and then tries to steal half of another P's queue, the work-stealing scheme described in work-stealing schedulers.

Blocking is handled by category. Channel operations, mutexes and select park the goroutine in the runtime and free the thread immediately; the mechanics of channels are covered in channels as a concurrency primitive. Network I/O uses non-blocking sockets and the netpoller (epoll on Linux, kqueue on BSD and macOS, IOCP on Windows): a goroutine waiting on a socket costs no thread. A blocking system call, such as a file read on Linux or a cgo call, does hold its thread; if it lasts long, the background sysmon thread retakes the P and gives it to another M, which is why a Go process can have many more threads than GOMAXPROCS.

Since Go 1.14 preemption is asynchronous: sysmon marks a goroutine that has run for more than 10 ms, and the runtime interrupts it (with a signal on Unix-like systems), so a tight loop no longer starves its P or stalls garbage collection.

Advertisement

GOMAXPROCS in containers

The number of Ps is the parallelism of your Go code. Before Go 1.25 it defaulted to the number of logical CPUs on the machine, so a container limited to 2 CPUs on a 64-core node ran 64 Ps, burned its CPU quota in a burst and was throttled for the rest of each period, which shows up as tail latency. Many services fixed this with a third-party library that read the cgroup limit.

From Go 1.25 the runtime does it itself on Linux: if the cgroup CPU bandwidth limit is lower than the number of logical CPUs, GOMAXPROCS defaults to the limit, with fractions rounded up and a floor of 2, and the runtime re-checks periodically if the limit or CPU count changes. This follows the Kubernetes CPU limit, not the CPU request; a pod with a request and no limit still sees every CPU on the node. Setting GOMAXPROCS explicitly, by environment variable or runtime.GOMAXPROCS, disables the automatic behaviour, and GODEBUG=containermaxprocs=0 and updatemaxprocs=0 turn off each part. Check your base images: a hard-coded value left over from older tooling now overrides the better default.

Bounding concurrency: a worked fan-out

The arithmetic of goroutine count is Little's law: concurrency equals arrival rate times time in system. Take an aggregation service at 2,000 requests per second, where each request fans out to 50 backend calls of 200 ms each, issued in parallel. Steady state is 2,000 x 0.2 x 50 = 20,000 concurrent backend goroutines, plus 400 request goroutines. That is fine. Now the backend slows to 5 seconds during an incident: concurrency becomes 500,000 goroutines, each holding a request context, a connection attempt and a response buffer. The process runs out of memory, restarts, and every client retries, so the next instance sees the same load. Unbounded fan-out converts downstream latency into upstream memory.

The fix is to bound concurrency at each level and to make waiting time-limited. Per request, cap the fan-out with errgroup from golang.org/x/sync, which provides SetLimit and TryGo; across requests, share a weighted semaphore per backend so the total is bounded, and fail fast when it is exhausted.

var backendSlots = semaphore.NewWeighted(2000) // whole-process cap for this backend

func aggregate(ctx context.Context, keys []string) ([]Item, error) {
    ctx, cancel := context.WithTimeout(ctx, 800*time.Millisecond)
    defer cancel()

    g, ctx := errgroup.WithContext(ctx)
    g.SetLimit(16)                                 // per-request fan-out cap
    out := make([]Item, len(keys))
    for i, k := range keys {
        g.Go(func() error {                        // Go 1.22+: i and k are per-iteration
            if err := backendSlots.Acquire(ctx, 1); err != nil {
                return err                         // deadline hit while queued: shed load
            }
            defer backendSlots.Release(1)
            item, err := fetch(ctx, k)             // must honour ctx
            out[i] = item
            return err
        })
    }
    return out, g.Wait()                           // every goroutine has returned here
}

Three properties make this safe. The context deadline bounds how long any goroutine can live. The two limits bound how many exist. And g.Wait guarantees that no goroutine outlives the function, so there is nothing to leak. For simple cases without errors, Go 1.25 added sync.WaitGroup.Go, which combines Add and the goroutine start.

Leaks: goroutines that never finish

A leaked goroutine is one blocked forever. The usual causes are a send on a channel nobody will receive from, a receive on a channel nobody will close, and a call that ignores its context. A typical case is a function that starts a worker, waits on a result channel with a timeout, and returns on the timeout: the worker finishes later and blocks forever sending its result to an unbuffered channel. Each timeout leaks one goroutine plus whatever it references, and the leak only appears under the slow conditions that trigger timeouts.

// Leaks one goroutine per timeout.
func lookup(ctx context.Context, k string) (string, error) {
    ch := make(chan string)              // unbuffered
    go func() { ch <- slowLookup(k) }()  // blocks forever if nobody receives
    select {
    case v := <-ch:
        return v, nil
    case <-ctx.Done():
        return "", ctx.Err()
    }
}
// Fix: make(chan string, 1) so the send always completes, and pass ctx
// into slowLookup so the work itself stops.

Detection has improved. Go 1.26 added an experimental profile, goroutineleak, enabled at build time with GOEXPERIMENT=goroutineleakprofile. It uses garbage-collector reachability: a goroutine blocked on a channel, mutex or condition variable that no runnable goroutine can reach can never wake, so it is reported as leaked. The Go 1.27 release notes list it as generally available, served by net/http/pprof at /debug/pprof/goroutineleak. On any version, the ordinary goroutine profile grouped by stack shows which code paths are accumulating, and comparing two snapshots taken minutes apart finds the growing ones. In tests, the testing/synctest package (generally available since Go 1.25) runs code in a bubble with a fake clock, so timeout paths can be exercised in milliseconds and goroutines still blocked at the end are visible; Uber's goleak package fails a test that leaves goroutines behind.

Observing goroutines in production

Export a small set of runtime metrics from every service. From runtime/metrics: /sched/goroutines:goroutines for the live count, and since Go 1.26 the per-state breakdown under /sched/goroutines/ (runnable, running, waiting, not-in-go), which distinguishes a CPU backlog from goroutines parked on I/O; /sched/latencies:seconds, a histogram of how long runnable goroutines waited for a P, the most direct signal that you are CPU-bound or that GOMAXPROCS is wrong; and /sched/gomaxprocs:threads to confirm the container default took effect.

Alert on goroutine count growing without matching traffic, and on scheduler latency percentiles. For incidents, the runtime/trace.FlightRecorder API added in Go 1.25 keeps the last few seconds of execution trace in a ring buffer, so you can write it out when a slow request is detected rather than tracing continuously. GODEBUG=schedtrace=1000 prints a line per second with run-queue lengths per P, which is crude but works on any binary. Continuous profiling of goroutine and CPU profiles, covered in profiling in production, makes leaks visible as a slow upward trend days before the out-of-memory kill.

Contention and garbage collection

Running more goroutines does not create more CPU. Past GOMAXPROCS runnable goroutines, extra ones only wait, and shared state becomes the bottleneck. A single sync.Mutex or channel touched by every goroutine serialises them; with heavy contention the mutex enters starvation mode (after a waiter has waited more than 1 ms), switching to strict hand-off, which is fair but slower. Shard hot state by key, keep per-goroutine counters and aggregate them, and use sync.Pool for short-lived buffers so a burst of goroutines does not become a burst of allocations. The lock mechanics are in how mutexes really work.

Garbage collection is the other shared cost. Since Go 1.26 the Green Tea collector is the default, with release notes reporting a 10 to 40 percent reduction in GC overhead for programs that allocate many small objects; it does not change the fact that live goroutines, and everything they reference, must be scanned. Fewer, longer-lived worker goroutines pulling from a queue can beat one goroutine per item when items are tiny and numerous.

Failure modes

SymptomLikely causeFix
Memory grows until OOM kill, count climbsLeak on channel send or ignored contextBuffered result channel, pass ctx, leak profile
Memory spikes when a backend slowsUnbounded fan-outerrgroup limit plus per-backend semaphore and deadline
High p99, CPU throttling in containerGOMAXPROCS above CPU limitGo 1.25+ default or set to the limit
Many more threads than GOMAXPROCSBlocking syscalls or cgo callsBound concurrent file or cgo work
Throughput flat as goroutines increaseShared mutex or channel hotspotShard state, reduce critical sections
Scheduler latency high with idle CPULong GC or lock convoysTrace with the flight recorder

Trade-offs

One goroutine per unit of work gives the simplest code and the best latency when work is I/O-bound and counts are bounded. A fixed pool of workers gives predictable memory and back-pressure, at the cost of queueing and slightly more code. Semaphores sit between: natural code, bounded count. Larger GOMAXPROCS increases parallelism but, inside a CPU limit, causes throttling. Buffered channels decouple producers and consumers and hide leaks; unbuffered channels make blocking visible but make the timeout-leak pattern easy to write. Choose bounded goroutine-per-task with deadlines by default, and pools only where per-item work is tiny.

What to do next

  1. Export the goroutine count, the per-state counts and the scheduler-latency histogram from every service, and alert on unexplained growth.
  2. Upgrade to Go 1.25 or later and remove hard-coded GOMAXPROCS settings, unless you have measured that a different value is better.
  3. Find every place that starts goroutines per request or per item; give each a context deadline and a concurrency limit.
  4. Audit select-with-timeout code for unbuffered result channels and fix them.
  5. Add synctest or goleak tests for timeout and cancellation paths.
  6. Enable the goroutine leak profile in staging and capture flight-recorder traces on slow requests.
  7. Load-test with a slow backend, not only a fast one, and watch memory and goroutine count.
Key takeaway: Goroutines are cheap but not free: each carries a stack, references heap objects and is a root the collector scans. The scheduler multiplexes them over GOMAXPROCS processors with local queues, work stealing, a netpoller and asynchronous preemption, and since Go 1.25 sizes GOMAXPROCS from container CPU limits. At scale, size concurrency with Little's law, bound every fan-out with deadlines and limits, fix select-with-timeout leaks, export scheduler metrics and use leak profiles and synctest to catch goroutines that never finish.