Scala's ExecutionContext.global runs futures on a pool sized to the number of CPU cores. That is the right size for computation and exactly the wrong size for waiting. Eight threads blocked on a slow JDBC query are eight threads doing nothing while every other future in the application queues behind them. The standard library's answer is a one-word construct, scala.concurrent.blocking, wrapped around code that blocks.

Most developers know it as an incantation. This article explains what it actually does: how the hint travels through BlockContext, how the global pool turns it into extra threads, the cap on how many, and the common situations where it silently does nothing. It then shows a reproducible experiment, an instrumenting BlockContext for finding blocking on the wrong pool, and when a dedicated pool or virtual threads are the better answer. Details are from the Scala 2.13 standard library, which Scala 3 also uses.

Advertisement

Why a core-sized pool cannot wait

A fork-join pool keeps a fixed target parallelism, normally the core count, and assumes each worker is either running or idle. When a task calls Thread.sleep, a JDBC driver, InputStream.read or a lock, the worker is neither: it is parked, holding a slot. The pool cannot tell the difference between a thread that is busy computing and one that is waiting on the network, so it does not start another.

With enough blocking tasks the pool degrades to serial waiting. On an eight-core machine, 64 futures that each block for one second take about eight seconds, as they run in waves of eight. Worse, if one of those tasks waits for the result of another future on the same pool, every worker can end up waiting for work that has no thread to run on. That is the starvation deadlock described in the Await article.

What blocking actually does

The method itself is tiny. Its whole job is to find the current BlockContext and hand it the code:

// scala.concurrent package object (2.13), simplified
def blocking[T](body: => T): T =
  BlockContext.current.blockOn(body)(scala.concurrent.AwaitPermission)

// scala.concurrent.BlockContext
trait BlockContext {
  def blockOn[T](thunk: => T)(implicit permission: CanAwait): T
}

BlockContext.current resolves in a fixed order. A context installed on the current thread with BlockContext.withBlockContext wins. Otherwise, if the current Thread object itself implements BlockContext, that is used. Otherwise the result is DefaultBlockContext, which runs the thunk immediately and does nothing else.

So blocking is a message, not a mechanism. It tells whatever is running the current thread that the enclosed code may block, and the meaning depends entirely on the receiver. Await.result and Await.ready send the same message internally, which is why awaiting on the global pool behaves better than raw blocking.

Advertisement

Inside the global pool: managed blocking with a cap

The global pool creates its workers through ExecutionContextImpl.DefaultThreadFactory, and each worker is a ForkJoinWorkerThread that also implements BlockContext. Its blockOn does three checks before doing anything: the calling thread must be the worker itself, the worker must not already be inside a blocking region, and a permit must be available from a semaphore created with maxBlockers permits. If all three pass, it marks itself blocked, wraps the thunk in a ForkJoinPool.ManagedBlocker and calls ForkJoinPool.managedBlock, releasing the permit in a finally. If any check fails, the thunk simply runs.

Future { blocking { jdbc() } }runs on a global workerBlockContext.currentthread-local, else threadDefaultBlockContextjust runs the thunkGlobal worker threadis a BlockContextPermit available?Semaphore(maxExtraThreads)ForkJoinPool.managedBlockpool may add a spare threadNo permit or nestedrun thunk, no compensationOther work proceedson the spare threadblocking(...)not a BlockContextworker threadblockOnyesnoThe same blocking call compensates on the global pool and does nothing on a plain fixed pool
How scala.concurrent.blocking is resolved. The hint is routed to the current BlockContext; only threads that implement one, such as the global pool's workers, can react by letting the ForkJoinPool add a spare thread.

managedBlock is the JDK's own protocol for this problem. Before the blocker runs, the pool may activate or create a spare worker so that parallelism stays at its target while this one waits. When the blocker finishes, the extra worker eventually retires. The effect is that blocking regions borrow threads instead of stealing slots.

The limits are system properties read when the global pool is created, so they must be set before first use, typically with -D flags:

PropertyDefaultMeaning
scala.concurrent.context.minThreads1Lower bound on parallelism
scala.concurrent.context.numThreadsx1Target parallelism; x1 means one per available processor
scala.concurrent.context.maxThreadsx1Upper bound on parallelism
scala.concurrent.context.maxExtraThreads256Permits for concurrent blocking regions that may get compensation

Two details matter in practice. Nesting does not compound: a blocking inside a blocking on the same worker skips compensation because the worker is already marked blocked. And the 257th concurrent blocking region on a default global pool runs uncompensated, so under a burst the pool can starve exactly as if you had never written the hint.

Where blocking does nothing

Because the behaviour depends on the thread, the same line of code can be a real safeguard in one place and decoration in another:

  • Fixed or cached thread pools. ExecutionContext.fromExecutorService(Executors.newFixedThreadPool(16)) creates ordinary threads that are not BlockContexts, so blocking falls through to the default and does nothing. The pool size you chose is the only protection.
  • Framework dispatchers. Actor and HTTP frameworks provide their own executors. Whether their threads honour the hint depends on the implementation and version; check before relying on it, and prefer the framework's documented blocking dispatcher.
  • Effect runtimes. Cats Effect and ZIO have their own blocking primitives, IO.blocking and ZIO.attemptBlocking, which shift work to a separate blocking pool. Use those rather than the standard-library hint inside effects; the Cats Effect article and the ZIO runtime article cover their pools.
  • Calls that do not block. Wrapping pure computation in blocking makes the global pool add threads for CPU work, oversubscribing the cores. Use it only around code that actually waits.

An experiment you can run in a minute

The effect is easy to see. This program runs 64 one-second sleeps on the global pool, once bare and once wrapped:

import scala.concurrent.*
import scala.concurrent.duration.*
import ExecutionContext.Implicits.global

object BlockingDemo:
  def run(label: String)(task: => Unit): Unit =
    val start = System.nanoTime()
    val all = Future.sequence((1 to 64).map(_ => Future(task)))
    Await.result(all, 5.minutes)
    val secs = (System.nanoTime() - start) / 1e9
    println(f"$label%-10s ${secs}%.1f s, cores=${Runtime.getRuntime.availableProcessors}")

  def main(args: Array[String]): Unit =
    run("bare")(Thread.sleep(1000))
    run("blocking")(blocking(Thread.sleep(1000)))

On an eight-core machine the bare run should take about eight seconds: 64 tasks in waves of eight. The wrapped run should take a little over one second, because each sleeping worker gets a spare and all 64 sleep together, well within the 256-permit cap. The exact figures depend on the machine; the shape is what matters. Now run it with -Dscala.concurrent.context.maxExtraThreads=8 and the wrapped run slows to roughly four seconds, since only eight blocking regions at a time get compensation.

Then try the worst case: change the pool to a fixed pool of eight threads with ExecutionContext.fromExecutorService. Both runs now take about eight seconds, because the hint has no receiver.

While the wrapped run is sleeping, take a thread dump with jcmd <pid> Thread.print. You will see many more threads named scala-execution-context-global-N than you have cores, most in TIMED_WAITING inside Thread.sleep. Wait a minute after the run and dump again: the spares have retired and the pool is back near its target size. That before-and-after picture is the quickest way to confirm, on a real service, that compensation is happening where you think it is.

Instrumenting blocking with your own BlockContext

Because BlockContext.current prefers a thread-installed context, you can wrap a region with your own context that observes blocking and then delegates to whatever was there before, so compensation still happens:

import java.util.concurrent.atomic.LongAdder
import scala.concurrent.{BlockContext, CanAwait}

final class MeteredBlockContext(underlying: BlockContext, calls: LongAdder, nanos: LongAdder)
    extends BlockContext:
  override def blockOn[T](thunk: => T)(implicit permission: CanAwait): T =
    calls.increment()
    val t0 = System.nanoTime()
    try underlying.blockOn(thunk)
    finally nanos.add(System.nanoTime() - t0)

object Metered:
  val calls = LongAdder()
  val nanos = LongAdder()

  def apply[T](body: => T): T =
    val previous = BlockContext.current
    BlockContext.withBlockContext(MeteredBlockContext(previous, calls, nanos))(body)

withBlockContext installs the context only for the synchronous extent of its body on the calling thread. Wrapping a handler that returns a Future therefore measures nothing, because the blocking happens later in tasks on other workers. Wrap the blocking body itself, as in Future(Metered(UserDao.find(id))), or wrap the pool so every task runs inside the context:

final class MeteredEC(underlying: ExecutionContext) extends ExecutionContext:
  def execute(r: Runnable): Unit = underlying.execute(() => Metered(r.run()))
  def reportFailure(t: Throwable): Unit = underlying.reportFailure(t)

Export the two counters as metrics. A sudden rise in blocking time per request points at a new blocking call on a compute pool. You can go further and log a stack trace when previous is not a thread that can compensate, which flags blocking inside a fixed pool where the hint is ignored. Keep the work inside blockOn cheap, because it runs on every blocking call.

The delegation is the important design choice. If MeteredBlockContext ran the thunk directly instead of calling underlying.blockOn, installing it would switch compensation off for every region it wraps, because the thread-local context now wins the lookup and the worker thread is never consulted. An instrumentation change that quietly reintroduces starvation is the worst kind, so test it with the 64-sleep experiment: with each task written as Future(Metered(blocking(Thread.sleep(1000)))), the run must still finish in about one second.

Note what this cannot see: blocking code that never calls blocking or Await. A JDBC call without the hint is invisible to BlockContext entirely. For that, use a thread-dump sampler or a profiler that shows threads in waiting states on the compute pool.

Dedicated pools and virtual threads

blocking on the global pool is a good default for occasional, short waits. For heavy, steady blocking, such as a service whose main job is calling a database, a dedicated pool is clearer and safer, because it bounds concurrency to what the downstream can take and isolates failures:

import java.util.concurrent.{Executors, ThreadFactory}
import java.util.concurrent.atomic.AtomicInteger
import scala.concurrent.{ExecutionContext, Future}

object Pools:
  private def named(prefix: String): ThreadFactory =
    val n = AtomicInteger()
    r => { val t = Thread(r, s"$prefix-${n.incrementAndGet()}"); t.setDaemon(true); t }

  // Sized to the connection pool, not the cores: more threads would only queue on connections.
  val jdbc: ExecutionContext =
    ExecutionContext.fromExecutorService(Executors.newFixedThreadPool(32, named("jdbc")))

def loadUser(id: Long)(using ExecutionContext): Future[User] =
  Future(UserDao.find(id))(Pools.jdbc)       // blocking work on its own bulkhead
    .map(enrich)                              // continue on the caller's compute pool

On JDK 21 or later, ExecutionContext.fromExecutorService(Executors.newVirtualThreadPerTaskExecutor()) gives each blocking task a virtual thread, which is cheap to park. It removes the thread-count problem but not the downstream limit, so keep a semaphore or connection pool in front of the database. Before JDK 24, blocking inside synchronized pinned the carrier thread; JEP 491 removed most of that. The virtual threads article and the Futures and ExecutionContext article cover both pools in more depth.

Failure modes and trade-offs

SymptomLikely causeFix
Latency spikes when a dependency slowsBlocking calls on the global pool without the hintWrap in blocking or move to a dedicated pool
Thread count jumps to hundredsblocking is working, many concurrent waitsBound concurrency with a dedicated pool or semaphore
Starvation despite blockingFixed pool, or more than maxExtraThreads concurrent regionsCheck the thread type; raise the cap or isolate the work
CPU oversubscribed, context switchingblocking wrapped around CPU-bound codeRemove the hint from non-waiting code
Deadlock on Await inside futuresWaiting for a future that needs the same saturated poolCompose with map and flatMap instead of awaiting

The underlying trade-off is between elasticity and bounds. blocking makes the global pool elastic, which is convenient and hides small mistakes, but elasticity also means a slow dependency can drive hundreds of threads and memory with no back-pressure. A dedicated pool is rigid, which is exactly what you want in front of a resource with a hard capacity.

What to do next

  1. Search your codebase for JDBC, file IO, Thread.sleep, HTTP client calls and lock waits inside Future bodies, and list which pool each runs on.
  2. For each on the global pool, add blocking if waits are short and rare; otherwise move it to a dedicated pool sized to its downstream.
  3. For each on a fixed or framework pool, remove any belief that blocking protects it, and size the pool deliberately.
  4. Run the 64-sleep experiment on your production JVM flags to confirm compensation works and to see the effect of maxExtraThreads.
  5. Install the metered BlockContext around blocking bodies or through a wrapping ExecutionContext, and export blocking calls and time as metrics.
  6. Alert on global-pool thread count and on blocking time per request.
  7. Replace Await inside futures with composition, and inside effect code use IO.blocking or ZIO.attemptBlocking.
  8. On JDK 21 or later, trial a virtual-thread ExecutionContext for blocking IO, keeping a connection-pool or semaphore limit in front of each dependency.
Key takeaway: scala.concurrent.blocking passes a hint to the current BlockContext. On the global pool, worker threads answer it by calling ForkJoinPool.managedBlock so a spare thread can keep work moving, up to maxExtraThreads concurrent regions, with no extra effect when nested. On ordinary thread pools it does nothing. Use it for short, occasional waits on the global pool, use dedicated or virtual-thread pools sized to the downstream for heavy blocking, and measure blocking with an instrumenting BlockContext.