Every map, flatMap, recover and onComplete on a Scala Future takes an implicit ExecutionContext, and most developers pass it the way they pass a logger: by importing whatever compiles. That habit is why thread-pool incidents dominate Future-based services. The ExecutionContext is the one decision that determines which thread runs your code, how many of them exist, what happens when they are all busy, and whether your request ID and tracing context survive the hop.
The companion article on Futures and ExecutionContext covers Futures as a whole, and the blocking deep dive covers managed blocking in the global pool. This page is about the abstraction itself: the two-method contract and what it does not promise, how global is built and tuned, the parasitic and opportunistic contexts, building your own from a Java executor, propagating thread-local context, testing with a deterministic context, and designing a service with separate pools as bulkheads.
The contract: two methods and four non-promises
The trait is tiny. An ExecutionContext must implement execute(runnable: Runnable): Unit, which schedules a piece of work, and reportFailure(cause: Throwable): Unit, which receives exceptions that escape that work and have nowhere else to go. A third method, prepare(), is deprecated since 2.12.0 and should be neither called nor overridden. Everything else is a convention of particular implementations.
What the contract leaves unsaid is the important part. execute does not promise a thread: the work may run on another thread, on the calling thread, or later in a batch. It does not promise ordering: two runnables submitted in sequence may run concurrently or in either order. It does not promise capacity: an implementation may queue without bound, reject, or block the caller. And it does not promise anything about thread-local state: a ThreadLocal set before execute is not visible inside the runnable unless the context copies it. Code that relies on any of these is relying on one implementation, and will break when someone passes a different one.
reportFailure deserves attention because it is where callback exceptions end up. If the function you pass to onComplete or foreach throws, there is no Future left to fail, so the exception goes to the context's reporter. The default reporter prints the stack trace to System.err, which in a container with structured logging means it is effectively lost. Every context you build should route failures to your logger and increment a metric.
Inside global: sizing, extra threads and daemons
ExecutionContext.global is backed by a work-stealing ForkJoinPool. Its size is controlled by four system properties, read once when the pool is first used: scala.concurrent.context.minThreads (default 1), scala.concurrent.context.numThreads (default x1), scala.concurrent.context.maxThreads (default x1) and scala.concurrent.context.maxExtraThreads (default 256). A value prefixed with x is a multiplier of available processors, so x1 means one thread per core and x2 means two. In a container, "available processors" is what the JVM reports from its CPU quota, so check Runtime.getRuntime.availableProcessors inside the container rather than assuming the host's core count.
The extra-threads limit exists for blocking { }: when code inside a blocking block runs on global, the pool may add a compensating thread so that the remaining work still makes progress, up to maxExtraThreads. Outside a blocking block, a blocked thread is simply lost capacity. Global's threads are daemon threads, so a main that starts Futures and returns can let the JVM exit before they finish; an application entry point must await its top-level work or install its own non-daemon pool.
// Set before anything touches ExecutionContext.global, e.g. in JAVA_OPTS:
// -Dscala.concurrent.context.numThreads=x1
// -Dscala.concurrent.context.maxThreads=x1
// -Dscala.concurrent.context.maxExtraThreads=64
import scala.concurrent.ExecutionContext
import scala.concurrent.ExecutionContext.Implicits.global // same pool as ExecutionContext.global
println(s"cores seen by this JVM: ${Runtime.getRuntime.availableProcessors}")Lowering maxExtraThreads from 256 is often worth considering: it caps how much damage unmarked-then-marked blocking can do to memory and scheduling, at the cost of queueing blocking work sooner. Either way, the robust fix for blocking is a dedicated pool, shown below.
parasitic and opportunistic
Scala 2.13 added two specialised contexts, and Scala 3 uses the same standard library. ExecutionContext.parasitic runs each runnable on the thread that calls execute, trampolining nested submissions so the stack does not grow without bound. It is the right choice for tiny, non-blocking transformations such as future.map(_.id)(ExecutionContext.parasitic), where a pool hop would cost more than the work. The standard library documentation is blunt about the risk: do not call blocking code in it, because you are stealing time from whichever thread completed the Future, and that may be a Netty event loop or a database driver's I/O thread.
ExecutionContext.opportunistic, added in 2.13.4, batches nested tasks and runs them on the same thread as the enclosing task, which makes chains of short callbacks cheaper. It is declared private[scala] for binary compatibility, so application code reaches it through a structural type, as the standard library's own documentation shows. Libraries should fall back to global if it is missing. Because a batch is pinned to one thread, long or blocking work inside it must be wrapped in blocking so pending tasks can move to another thread.
import scala.concurrent.{ExecutionContext, ExecutionContextExecutor}
import scala.language.reflectiveCalls // Scala 3: import scala.reflect.Selectable.reflectiveSelectable instead
val opportunistic: ExecutionContextExecutor =
(ExecutionContext: { def opportunistic: ExecutionContextExecutor }).opportunistic
// cheap projection: no pool hop
val ids = usersF.map(_.map(_.id))(ExecutionContext.parasitic)
Building your own context
Any Java Executor becomes an ExecutionContext through ExecutionContext.fromExecutor(e, reporter), and an ExecutorService through fromExecutorService(e, reporter), which returns an ExecutionContextExecutorService you can shut down. Building your own is how you get the properties global does not have: named threads for thread dumps, a bounded queue, an explicit rejection policy, and a reporter wired to logging.
import java.util.concurrent._
import java.util.concurrent.atomic.AtomicInteger
import scala.concurrent.{ExecutionContext, ExecutionContextExecutorService}
def namedFactory(prefix: String): ThreadFactory = new ThreadFactory {
private val n = new AtomicInteger()
def newThread(r: Runnable): Thread = {
val t = new Thread(r, s"$prefix-${n.incrementAndGet()}")
t.setDaemon(false)
t
}
}
def boundedPool(name: String, threads: Int, queue: Int): ExecutionContextExecutorService = {
val exec = new ThreadPoolExecutor(
threads, threads, 60L, TimeUnit.SECONDS,
new ArrayBlockingQueue[Runnable](queue),
namedFactory(name),
new ThreadPoolExecutor.AbortPolicy()) // overload fails fast, visibly
ExecutionContext.fromExecutorService(exec, t => {
log.error(s"[$name] uncaught in callback", t)
metrics.counter(s"ec.$name.reported").increment()
})
}
val jdbcEc = boundedPool("jdbc", threads = 16, queue = 1000)
sys.addShutdownHook { jdbcEc.shutdown(); jdbcEc.awaitTermination(30, TimeUnit.SECONDS) }The rejection policy is a design decision, not a detail. With AbortPolicy, a full queue throws RejectedExecutionException from execute. When that happens inside Future.apply, the Future fails and the caller sees an error it can map to HTTP 503; when it happens while scheduling a callback, the standard library catches it and the derived Future generally fails with the rejection instead, which is easy to misread as a downstream bug, so count rejections at the pool. CallerRunsPolicy applies backpressure by running the work on the submitting thread, which is safe only if that thread is allowed to block. An unbounded LinkedBlockingQueue never rejects and turns overload into latency and, eventually, memory exhaustion.
Carrying context across threads
Thread-local state such as SLF4J's MDC or a tracing span does not follow a Future across contexts, because each callback may run on a different thread. The fix is a wrapping context that captures the state at execute time and restores it around the runnable:
import org.slf4j.MDC
import scala.concurrent.ExecutionContext
final class MdcPropagatingEc(underlying: ExecutionContext) extends ExecutionContext {
def execute(r: Runnable): Unit = {
val captured = MDC.getCopyOfContextMap // may be null
underlying.execute { () =>
val previous = MDC.getCopyOfContextMap
if (captured == null) MDC.clear() else MDC.setContextMap(captured)
try r.run()
finally if (previous == null) MDC.clear() else MDC.setContextMap(previous)
}
}
def reportFailure(t: Throwable): Unit = underlying.reportFailure(t)
}Restoring the previous map in finally matters: pool threads are reused, and without it one request's ID leaks into the next request's log lines, which is worse than having no ID at all. Wrap every pool you create, not just global, or context disappears at exactly the hop into the JDBC pool where you most need it. Most tracing libraries ship an equivalent wrapper; prefer theirs when you use one.
Worked example: pools as bulkheads
A worked example: an order service handles HTTP requests, parses and validates JSON, queries Postgres through JDBC, and calls a pricing service over a non-blocking HTTP client. Treat each kind of work as a separate failure domain with its own pool, a bulkhead, so that one cannot starve the others:
| Work | Context | Sizing rule | Overload behaviour |
|---|---|---|---|
| Parsing, validation, composition | global (CPU-bound) | One thread per core | Queues briefly; never blocked |
| JDBC queries | jdbcEc, bounded | Equal to the connection-pool size; more threads only wait for connections | Reject when the queue fills; return 503 |
| Tiny projections of results | parasitic | No threads of its own | Inherits the completing thread |
| Pricing client callbacks | global | The client's own I/O threads complete Futures | Client-side timeouts and limits |
import scala.concurrent.{ExecutionContext, Future}
final class OrderService(repo: OrderRepo, pricing: PricingClient)(implicit cpu: ExecutionContext) {
private val db: ExecutionContext = new MdcPropagatingEc(jdbcEc)
def quote(req: QuoteRequest): Future[Quote] =
for {
valid <- Future(validate(req)) // cpu
order <- Future(repo.loadOrder(valid.orderId))(db) // blocking JDBC on its own pool
prices <- pricing.prices(order.skus) // non-blocking client
} yield priceOrder(order, prices) // cpu
}Two details carry the design. The JDBC pool is sized to the connection pool, because a thread without a connection only waits; if the connection pool has 16 connections, 64 JDBC threads add 48 parked threads and no throughput. And the explicit (db) on the single blocking call keeps the implicit context as the CPU pool for everything else, so the boundary is visible in code review. Under a database slowdown, JDBC requests queue and then fail fast with 503, while health checks and requests that never touch the database still run on global.
Testing with a deterministic context
Tests that use global are nondeterministic. A context that only queues work, and runs it when the test says so, makes asynchronous code step-by-step reproducible:
import scala.collection.mutable
import scala.concurrent.{ExecutionContext, Future}
final class ManualEc extends ExecutionContext {
private val queue = mutable.Queue.empty[Runnable]
val failures = mutable.Buffer.empty[Throwable]
def execute(r: Runnable): Unit = queue.enqueue(r)
def reportFailure(t: Throwable): Unit = failures += t
def runAll(): Int = { var n = 0; while (queue.nonEmpty) { queue.dequeue().run(); n += 1 }; n }
}
// in a test
implicit val ec: ManualEc = new ManualEc
val f = Future(21).map(_ * 2)
assert(!f.isCompleted) // nothing has run yet
ec.runAll()
assert(f.value.contains(scala.util.Success(42)))
assert(ec.failures.isEmpty)Because nothing runs until runAll, you can assert intermediate states, check that a callback exception reached reportFailure, and prove ordering assumptions false by running queued tasks in a different order.
Failure modes and alternatives
| Failure | Cause | Fix |
|---|---|---|
| Service freezes with idle CPU | Blocking calls on global without a blocking block, or more blocking than extra threads | Dedicated bounded pool for blocking work |
| Callback exceptions vanish | Default reporter prints to System.err | Reporter wired to logging and a metric |
| Request IDs missing or wrong in logs | Thread-locals not propagated, or not restored | Wrapping context with capture and restore |
| Latency climbs, memory grows | Unbounded queue in a custom pool | Bounded queue and an explicit rejection policy |
| Event loop stalls | parasitic callback did real work on an I/O thread | Keep parasitic for trivial projections only |
| Program exits before work finishes | Global threads are daemon | Await top-level Futures or use a non-daemon pool |
| Too many threads in a container | Pool sized from host cores or blind multipliers | Size from the JVM's view of the CPU quota |
Futures are not the only way to run asynchronous Scala. Cats Effect and ZIO run lazy effects on their own fiber runtimes, with separate compute and blocking pools built in and propagation of fiber-local state; the trade-offs are compared in the effect systems comparison, and ZIO's scheduler is covered in the ZIO runtime article. If you stay on Futures, the ExecutionContext is the runtime, and designing it is your job.
What to do next
- List every ExecutionContext in your service: search for
Implicits.global,fromExecutorandExecutionContextparameters, and write down what each pool is for. - Print
availableProcessorsinside your production container and check global's sizing properties against it. - Give every custom pool named threads, a bounded queue, an explicit rejection policy and a reporter that logs and counts.
- Move every blocking call (JDBC, file I/O, legacy clients) to a dedicated pool sized to the resource it waits on.
- Wrap each pool with MDC or tracing propagation and verify with a log line on both sides of a pool hop.
- Use parasitic only for trivial, non-blocking projections, and say so in a comment.
- Write one deterministic test with a manual context for your most important asynchronous flow.