Scala is a strict language: arguments are evaluated before a call, and a val is computed where it is declared. Laziness is something you ask for, in four distinct ways: by-name parameters, lazy val, lazy collections, and effect values such as IO that describe work without doing it. Each one changes when code runs, how often it runs and which thread runs it, and each has a cost that shows up in production rather than in the REPL.
This page treats evaluation strategy as a language mechanism. It explains the three strategies from first principles, what the compiler generates for by-name parameters and lazy vals, how the Scala 3.3 lazy val scheme differs from Scala 2's, where deadlocks and repeated work come from, and how effect systems take the idea further. Lazy collections are covered in Scala Streams and lazy sequences, and the basics of val, var and initialization order are in Scala val vs var, in depth.
Three evaluation strategies
An evaluation strategy answers two questions about an argument: when is it evaluated, and how many times? Call-by-value evaluates it once, before the call. Call-by-name evaluates it at every use, possibly never. Call-by-need, Haskell's default, delays evaluation but caches the first result, so it runs at most once.
Scala gives you all three. Ordinary parameters are strict. A parameter typed => T is by-name. Call-by-need is not a parameter mode, but you get it by capturing a by-name parameter in a lazy val. The example below makes the difference visible with a side effect.
// Strict: the argument is evaluated once, before the call.
def twiceStrict(x: Int): Int = x + x
// By-name: x is a thunk; every use re-runs it.
def twiceByName(x: => Int): Int = x + x
// Call-by-need: capture the thunk in a lazy val to run it at most once.
def twiceByNeed(x: => Int): Int =
lazy val v = x
v + v
def noisy(): Int = { println("evaluated"); 21 }
twiceStrict(noisy()) // prints once, returns 42
twiceByName(noisy()) // prints twice, returns 42
twiceByNeed(noisy()) // prints once, returns 42| Strategy | Scala form | Runs | Typical use |
|---|---|---|---|
| Call-by-value | x: Int | exactly once, before the call | the default; predictable and cheapest |
| Call-by-name | x: => Int | zero or more times, at each use | control structures, retries, lazy log messages |
| Call-by-need | lazy val v = x | zero or one time, on first use | expensive values that may not be needed |
What a by-name parameter really is
A by-name parameter is compiled to a function of no arguments, a Function0. At the call site the compiler wraps the argument expression in a closure; inside the method every mention of the parameter becomes a call to that closure's apply. Nothing is cached, which is why the by-name version above prints twice.
Two consequences follow. First, the closure allocation costs something on hot paths, although the JIT often removes it when the call is inlined. Second, by-name parameters let ordinary methods behave like language syntax. The block passed to retry below is not run at the call site; retry runs it, and re-runs it on failure. Closures in general are covered in Scala Closures.
import scala.util.control.NonFatal
// A control structure: the block is not run until retry decides to run it.
def retry[A](attempts: Int, backoffMs: Long)(block: => A): A =
try block
catch
case NonFatal(e) if attempts > 1 =>
Thread.sleep(backoffMs)
retry(attempts - 1, backoffMs * 2)(block) // pass the thunk on, unevaluated
// A logger that never builds the message when the level is off.
final class Log(enabled: Boolean):
def debug(msg: => String): Unit = if enabled then println(msg)
val user = retry(3, 100)(httpGetUser(42)) // re-runs the HTTP call on failure
Log(enabled = false).debug(s"state = ${expensiveDump()}") // expensiveDump never runsThe logger pattern is the most common real-world win. A strict debug(msg: String) builds the interpolated string, and whatever expensiveDump() does, even when debug logging is off. The by-name version skips all of it. The standard library already uses this shape in require(cond, message), whose message is by-name.
Where the standard library is already lazy
Several everyday operations are lazy by design, and reading their signatures tells you which. The boolean operators && and || evaluate their right operand only when needed, and if evaluates only one branch. Option.getOrElse, Map.getOrElse and Option.fold's empty case take their default by-name, so cache.getOrElse(k, loadFromDb(k)) only hits the database on a miss. Try.apply takes its body by-name so it can catch what the body throws.
One signature misleads people. Future.apply also takes its body by-name, but submits it to the execution context immediately: the parameter is by-name so the work can run on another thread, not to postpone it. A Future is eager and memoized.
lazy val internals: Scala 2
A lazy val is call-by-need attached to a field: computed on first read, cached, and safe to read from many threads. Thread safety is the expensive part. In Scala 2 the compiler adds a bitmap field with one bit per lazy val and a private initializer method. A read checks the bit; if it is clear, the initializer enters synchronized on the enclosing instance, checks the bit again, runs the right-hand side, stores the value and sets the bit. This is double-checked locking, made correct by the memory-model guarantees of the monitor and the volatile bitmap.
The flaw is the lock's scope. The monitor is the whole enclosing object, held for the entire time the initializer runs. Any other lazy val in the same object, and any other code that synchronizes on that object, waits behind a slow initializer. Worse, if an initializer touches a lazy val in another object while a second thread does the reverse, each thread holds one monitor and waits for the other.
// Scala 2: each lazy val initializer runs inside the enclosing object's monitor.
object A { lazy val a: Int = { Thread.sleep(50); B.b + 1 } }
object B { lazy val b: Int = { Thread.sleep(50); A.a + 1 } }
// Thread 1 reads A.a: locks A, sleeps, then needs B's lock.
// Thread 2 reads B.b: locks B, sleeps, then needs A's lock. -> deadlock.
// Even single-threaded this is a cycle with no answer; in Scala 3 the
// reference documents recursive lazy val behaviour as undefined.Objects add a separate hazard: an object is initialized by JVM class initialization, which has its own lock, so two objects whose constructors reference each other from different threads can deadlock even without lazy vals.
lazy val internals: Scala 3.3 and later
Scala 3.0 to 3.2 already avoided Scala 2's monitor with a bitmap and compare-and-swap scheme; Scala 3.3.0 replaced it with a lighter implementation, on by default. Each lazy val gets one volatile field of type Object whose contents encode its state. null means not initialized. The first reader uses compare-and-swap to install a marker called Evaluating and runs the initializer. A second reader that finds Evaluating swaps in a Waiting object, which is a CountDownLatch(1), and parks on it. When the initializer returns, the value is stored (a NullValue marker stands in for a genuine null) and any latch is released. The marker types live in scala.runtime.LazyVals.
// Roughly what Scala 3.3+ generates for: class Conf { lazy val db: Db = connect() }
class Conf:
@volatile private var db$lzy: AnyRef = null // null, Evaluating, Waiting, NullValue or the value
def db: Db =
val cur = db$lzy
if cur.isInstanceOf[Db] then cur.asInstanceOf[Db] // fast path: one volatile read
else db$lzyINIT() // slow path: CAS loop below
private def db$lzyINIT(): Db =
// loop:
// null -> CAS to Evaluating; run connect(); store value; if CAS finds Waiting, countDown()
// Evaluating -> CAS to a new Waiting latch, then await()
// Waiting -> await() on the existing latch
// value -> return it
// if connect() throws: reset to null, release any latch, rethrow
???Because no monitor is held while the initializer runs, unrelated lazy vals in the same instance no longer block each other, and code synchronizing on the instance is not affected. The Scala 3 reference describes the aim as reducing the possibility of deadlock, not eliminating it: a thread waiting on a latch is still waiting, so a cycle of lazy vals that depend on each other across threads can still hang, and the reference states that recursive lazy val behaviour is undefined. Keep initializers acyclic.
The 3.3.0 release notes also record a compatibility escape hatch, -Ylegacy-lazy-vals, for environments where the new scheme caused trouble, GraalVM native images being the case reported at the time. For a lazy val that only ever one thread touches, Scala 3 provides @scala.annotation.threadUnsafe, which generates a plain check-and-assign with no atomics at all.
What laziness costs
Every read of a thread-safe lazy val performs at least one volatile read, which on x86 is nearly free but still prevents some compiler and JIT optimizations, and on weakly ordered CPUs such as ARM costs a load-acquire. In a tight loop, copy the lazy val into a local val once. The first read is far more expensive: a compare-and-swap, the initializer itself, and possibly parking threads.
Laziness also changes memory and failure timing. A pending lazy val or stored by-name thunk keeps everything it captures alive. A strict val that fails to load configuration fails at startup, where an operator sees it; a lazy one fails on the first request that touches it, and because a throwing initializer is not cached, every later request retries the failing work.
Laziness as suspended effects
Effect systems such as Cats Effect and ZIO generalize the idea. An IO[A] is a value that describes a computation; nothing happens until the runtime interprets it. That is call-by-name for whole programs: running the same IO twice performs its effects twice. The benefit is referential transparency, explained in Referential transparency in Scala.
The bug to watch for is accidentally strict construction. IO.pure takes an already-computed value, so any side effect inside its argument runs once, when the IO is built. IO.delay takes its argument by-name and runs it each time. When you do want call-by-need inside an effect system, ask for it explicitly with memoize, which returns an IO[IO[A]]: running the outer effect sets up the cache, and the inner effect computes once and then reuses the result. More on the runtime in Cats Effect.
import cats.effect.{IO, IOApp}
object Demo extends IOApp.Simple:
val eager: IO[Long] = IO.pure(System.nanoTime()) // nanoTime ran ONCE, when eager was built
val delayed: IO[Long] = IO.delay(System.nanoTime()) // nanoTime runs each time the IO runs
def run: IO[Unit] =
for
a <- eager; b <- eager // a == b
c <- delayed; d <- delayed // c != d
once <- delayed.memoize // IO[IO[Long]]: first run computes, later runs reuse
e <- once; f <- once // e == f
_ <- IO.println((a == b, c == d, e == f)) // (true,false,true)
yield ()
Worked example: a service that starts fast and fails loudly
Consider a service with three dependencies: a configuration file, a database pool and a rarely used PDF renderer that is slow to warm up. A first version makes all three lazy vals to speed up startup. In production the pool's configuration is wrong, the deploy looks healthy, the first request fails, and every later request retries the pool initializer and floods the database with connection attempts.
The fix is a strategy per dependency. Configuration and the pool are built strictly at startup, and the readiness probe fails until the pool can connect, so a bad deploy never receives traffic. Only the PDF renderer stays lazy, with its failure cached for thirty seconds instead of retried per request. Debug logging takes by-name messages, so the diagnostic dump costs nothing when the level is off.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| A side effect runs twice | by-name parameter used more than once | bind it to a lazy val, or make the parameter strict |
| Threads hang on startup | lazy val or object cycle across threads | break the cycle; initialize eagerly in a defined order |
| A failing dependency is hit on every request | lazy initializer throws and is never cached | fail fast at startup or cache the failure explicitly |
| An effect runs once instead of each time | work placed inside IO.pure | use IO.delay; memoize only on purpose |
Trade-offs
Strict evaluation is the easiest to reason about: code runs where it is written, once, and fails at a predictable time. Call-by-name buys control structures and free disabled logging for repeated evaluation and closure allocation. Lazy vals skip unneeded work but add per-read cost and late failures. Effect types make laziness the default and recover predictability through types, at the cost of a runtime.
What to do next
- Search your codebase for
lazy valand classify each one: needed for cost, needed for initialization order, or habit. Make the habitual ones strict. - Move configuration and connection pools to strict startup initialization behind a readiness check.
- Change debug and trace logging helpers to take by-name messages, or use a logging library that already does.
- Check every by-name parameter for multiple uses and bind it to a lazy val where it should run once.
- If you are still on Scala 2, look for lazy vals in objects that reference other objects' lazy vals and remove the cycles before upgrading to Scala 3.3 or later.
- In effectful code, grep for
IO.pure(with a side effect in its argument, and usememoizeonly where caching is intended.