Go makes concurrency cheap enough that a single process can hold hundreds of thousands of goroutines, and it is common to see services that start one per request, one per downstream call and one per stream. Cheap is not free. At scale goroutines cost memory, garbage-collector time and scheduler work, and the classic production incidents in Go services, memory that grows until the pod is killed or latency that rises under load, are usually caused by goroutines that were never bounded or never finished.
This article explains how the runtime schedules goroutines, what each one costs, and how to bound, observe and test them. Runtime details were checked against the Go source and the Go 1.25 and 1.26 release notes. If you come from Java, virtual threads in production covers the same problem on the JVM.
What a goroutine costs
A goroutine is a runtime structure (the g) plus a stack. The minimum stack is 2 KiB (stackMin = 2048 in runtime/stack.go). Since Go 1.19 the runtime picks the starting size from the average stack use it observed at recent garbage collections, so a program whose goroutines use deep stacks starts new ones larger; GODEBUG=adaptivestackstart=0 turns that off. When a function needs more stack than is left, the runtime allocates a stack twice as large, copies the frames and adjusts pointers into the stack; during garbage collection a stack using less than a quarter of its size can be shrunk.
So the floor is a few kilobytes per goroutine, and 100,000 idle goroutines need at least about 200 MiB of stack. In practice the stack is the small part. What a goroutine references, request bodies, buffers, decoded JSON, a TLS connection, is what fills the heap, and every live goroutine stack is a root the garbage collector must scan. A goroutine count is therefore a proxy for retained memory and GC work, not just for scheduler load.
How the scheduler runs them
The scheduler has three kinds of object. A G is a goroutine. An M is an operating-system thread. A P is a logical processor, a token that an M must hold to run Go code; there are GOMAXPROCS of them. Each P owns a local run queue of 256 slots plus a single runnext slot, which holds a goroutine just made runnable by the current one so that a producer and consumer pair runs back to back.
When a P's local queue is full, half of it is moved to the global run queue. To stop the global queue from starving, a P checks it on every 61st scheduling tick even when it has local work. A P with nothing to do looks in the global queue, polls the network, and then tries to steal half of another P's queue, the work-stealing scheme described in work-stealing schedulers.
Blocking is handled by category. Channel operations, mutexes and select park the goroutine in the runtime and free the thread immediately; the mechanics of channels are covered in channels as a concurrency primitive. Network I/O uses non-blocking sockets and the netpoller (epoll on Linux, kqueue on BSD and macOS, IOCP on Windows): a goroutine waiting on a socket costs no thread. A blocking system call, such as a file read on Linux or a cgo call, does hold its thread; if it lasts long, the background sysmon thread retakes the P and gives it to another M, which is why a Go process can have many more threads than GOMAXPROCS.
Since Go 1.14 preemption is asynchronous: sysmon marks a goroutine that has run for more than 10 ms, and the runtime interrupts it (with a signal on Unix-like systems), so a tight loop no longer starves its P or stalls garbage collection.
GOMAXPROCS in containers
The number of Ps is the parallelism of your Go code. Before Go 1.25 it defaulted to the number of logical CPUs on the machine, so a container limited to 2 CPUs on a 64-core node ran 64 Ps, burned its CPU quota in a burst and was throttled for the rest of each period, which shows up as tail latency. Many services fixed this with a third-party library that read the cgroup limit.
From Go 1.25 the runtime does it itself on Linux: if the cgroup CPU bandwidth limit is lower than the number of logical CPUs, GOMAXPROCS defaults to the limit, with fractions rounded up and a floor of 2, and the runtime re-checks periodically if the limit or CPU count changes. This follows the Kubernetes CPU limit, not the CPU request; a pod with a request and no limit still sees every CPU on the node. Setting GOMAXPROCS explicitly, by environment variable or runtime.GOMAXPROCS, disables the automatic behaviour, and GODEBUG=containermaxprocs=0 and updatemaxprocs=0 turn off each part. Check your base images: a hard-coded value left over from older tooling now overrides the better default.
Bounding concurrency: a worked fan-out
The arithmetic of goroutine count is Little's law: concurrency equals arrival rate times time in system. Take an aggregation service at 2,000 requests per second, where each request fans out to 50 backend calls of 200 ms each, issued in parallel. Steady state is 2,000 x 0.2 x 50 = 20,000 concurrent backend goroutines, plus 400 request goroutines. That is fine. Now the backend slows to 5 seconds during an incident: concurrency becomes 500,000 goroutines, each holding a request context, a connection attempt and a response buffer. The process runs out of memory, restarts, and every client retries, so the next instance sees the same load. Unbounded fan-out converts downstream latency into upstream memory.
The fix is to bound concurrency at each level and to make waiting time-limited. Per request, cap the fan-out with errgroup from golang.org/x/sync, which provides SetLimit and TryGo; across requests, share a weighted semaphore per backend so the total is bounded, and fail fast when it is exhausted.
var backendSlots = semaphore.NewWeighted(2000) // whole-process cap for this backend
func aggregate(ctx context.Context, keys []string) ([]Item, error) {
ctx, cancel := context.WithTimeout(ctx, 800*time.Millisecond)
defer cancel()
g, ctx := errgroup.WithContext(ctx)
g.SetLimit(16) // per-request fan-out cap
out := make([]Item, len(keys))
for i, k := range keys {
g.Go(func() error { // Go 1.22+: i and k are per-iteration
if err := backendSlots.Acquire(ctx, 1); err != nil {
return err // deadline hit while queued: shed load
}
defer backendSlots.Release(1)
item, err := fetch(ctx, k) // must honour ctx
out[i] = item
return err
})
}
return out, g.Wait() // every goroutine has returned here
}Three properties make this safe. The context deadline bounds how long any goroutine can live. The two limits bound how many exist. And g.Wait guarantees that no goroutine outlives the function, so there is nothing to leak. For simple cases without errors, Go 1.25 added sync.WaitGroup.Go, which combines Add and the goroutine start.
Leaks: goroutines that never finish
A leaked goroutine is one blocked forever. The usual causes are a send on a channel nobody will receive from, a receive on a channel nobody will close, and a call that ignores its context. A typical case is a function that starts a worker, waits on a result channel with a timeout, and returns on the timeout: the worker finishes later and blocks forever sending its result to an unbuffered channel. Each timeout leaks one goroutine plus whatever it references, and the leak only appears under the slow conditions that trigger timeouts.
// Leaks one goroutine per timeout.
func lookup(ctx context.Context, k string) (string, error) {
ch := make(chan string) // unbuffered
go func() { ch <- slowLookup(k) }() // blocks forever if nobody receives
select {
case v := <-ch:
return v, nil
case <-ctx.Done():
return "", ctx.Err()
}
}
// Fix: make(chan string, 1) so the send always completes, and pass ctx
// into slowLookup so the work itself stops.Detection has improved. Go 1.26 added an experimental profile, goroutineleak, enabled at build time with GOEXPERIMENT=goroutineleakprofile. It uses garbage-collector reachability: a goroutine blocked on a channel, mutex or condition variable that no runnable goroutine can reach can never wake, so it is reported as leaked. The Go 1.27 release notes list it as generally available, served by net/http/pprof at /debug/pprof/goroutineleak. On any version, the ordinary goroutine profile grouped by stack shows which code paths are accumulating, and comparing two snapshots taken minutes apart finds the growing ones. In tests, the testing/synctest package (generally available since Go 1.25) runs code in a bubble with a fake clock, so timeout paths can be exercised in milliseconds and goroutines still blocked at the end are visible; Uber's goleak package fails a test that leaves goroutines behind.
Observing goroutines in production
Export a small set of runtime metrics from every service. From runtime/metrics: /sched/goroutines:goroutines for the live count, and since Go 1.26 the per-state breakdown under /sched/goroutines/ (runnable, running, waiting, not-in-go), which distinguishes a CPU backlog from goroutines parked on I/O; /sched/latencies:seconds, a histogram of how long runnable goroutines waited for a P, the most direct signal that you are CPU-bound or that GOMAXPROCS is wrong; and /sched/gomaxprocs:threads to confirm the container default took effect.
Alert on goroutine count growing without matching traffic, and on scheduler latency percentiles. For incidents, the runtime/trace.FlightRecorder API added in Go 1.25 keeps the last few seconds of execution trace in a ring buffer, so you can write it out when a slow request is detected rather than tracing continuously. GODEBUG=schedtrace=1000 prints a line per second with run-queue lengths per P, which is crude but works on any binary. Continuous profiling of goroutine and CPU profiles, covered in profiling in production, makes leaks visible as a slow upward trend days before the out-of-memory kill.
Contention and garbage collection
Running more goroutines does not create more CPU. Past GOMAXPROCS runnable goroutines, extra ones only wait, and shared state becomes the bottleneck. A single sync.Mutex or channel touched by every goroutine serialises them; with heavy contention the mutex enters starvation mode (after a waiter has waited more than 1 ms), switching to strict hand-off, which is fair but slower. Shard hot state by key, keep per-goroutine counters and aggregate them, and use sync.Pool for short-lived buffers so a burst of goroutines does not become a burst of allocations. The lock mechanics are in how mutexes really work.
Garbage collection is the other shared cost. Since Go 1.26 the Green Tea collector is the default, with release notes reporting a 10 to 40 percent reduction in GC overhead for programs that allocate many small objects; it does not change the fact that live goroutines, and everything they reference, must be scanned. Fewer, longer-lived worker goroutines pulling from a queue can beat one goroutine per item when items are tiny and numerous.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| Memory grows until OOM kill, count climbs | Leak on channel send or ignored context | Buffered result channel, pass ctx, leak profile |
| Memory spikes when a backend slows | Unbounded fan-out | errgroup limit plus per-backend semaphore and deadline |
| High p99, CPU throttling in container | GOMAXPROCS above CPU limit | Go 1.25+ default or set to the limit |
| Many more threads than GOMAXPROCS | Blocking syscalls or cgo calls | Bound concurrent file or cgo work |
| Throughput flat as goroutines increase | Shared mutex or channel hotspot | Shard state, reduce critical sections |
| Scheduler latency high with idle CPU | Long GC or lock convoys | Trace with the flight recorder |
Trade-offs
One goroutine per unit of work gives the simplest code and the best latency when work is I/O-bound and counts are bounded. A fixed pool of workers gives predictable memory and back-pressure, at the cost of queueing and slightly more code. Semaphores sit between: natural code, bounded count. Larger GOMAXPROCS increases parallelism but, inside a CPU limit, causes throttling. Buffered channels decouple producers and consumers and hide leaks; unbuffered channels make blocking visible but make the timeout-leak pattern easy to write. Choose bounded goroutine-per-task with deadlines by default, and pools only where per-item work is tiny.
What to do next
- Export the goroutine count, the per-state counts and the scheduler-latency histogram from every service, and alert on unexplained growth.
- Upgrade to Go 1.25 or later and remove hard-coded GOMAXPROCS settings, unless you have measured that a different value is better.
- Find every place that starts goroutines per request or per item; give each a context deadline and a concurrency limit.
- Audit select-with-timeout code for unbuffered result channels and fix them.
- Add synctest or goleak tests for timeout and cancellation paths.
- Enable the goroutine leak profile in staging and capture flight-recorder traces on slow requests.
- Load-test with a slow backend, not only a fast one, and watch memory and goroutine count.