Scala is a strict language: arguments are evaluated before a call, and a val is computed where it is declared. Laziness is something you ask for, in four distinct ways: by-name parameters, lazy val, lazy collections, and effect values such as IO that describe work without doing it. Each one changes when code runs, how often it runs and which thread runs it, and each has a cost that shows up in production rather than in the REPL.

This page treats evaluation strategy as a language mechanism. It explains the three strategies from first principles, what the compiler generates for by-name parameters and lazy vals, how the Scala 3.3 lazy val scheme differs from Scala 2's, where deadlocks and repeated work come from, and how effect systems take the idea further. Lazy collections are covered in Scala Streams and lazy sequences, and the basics of val, var and initialization order are in Scala val vs var, in depth.

Advertisement

Three evaluation strategies

An evaluation strategy answers two questions about an argument: when is it evaluated, and how many times? Call-by-value evaluates it once, before the call. Call-by-name evaluates it at every use, possibly never. Call-by-need, Haskell's default, delays evaluation but caches the first result, so it runs at most once.

Scala gives you all three. Ordinary parameters are strict. A parameter typed => T is by-name. Call-by-need is not a parameter mode, but you get it by capturing a by-name parameter in a lazy val. The example below makes the difference visible with a side effect.

// Strict: the argument is evaluated once, before the call.
def twiceStrict(x: Int): Int = x + x

// By-name: x is a thunk; every use re-runs it.
def twiceByName(x: => Int): Int = x + x

// Call-by-need: capture the thunk in a lazy val to run it at most once.
def twiceByNeed(x: => Int): Int =
  lazy val v = x
  v + v

def noisy(): Int = { println("evaluated"); 21 }

twiceStrict(noisy())   // prints once, returns 42
twiceByName(noisy())   // prints twice, returns 42
twiceByNeed(noisy())   // prints once, returns 42
StrategyScala formRunsTypical use
Call-by-valuex: Intexactly once, before the callthe default; predictable and cheapest
Call-by-namex: => Intzero or more times, at each usecontrol structures, retries, lazy log messages
Call-by-needlazy val v = xzero or one time, on first useexpensive values that may not be needed

What a by-name parameter really is

A by-name parameter is compiled to a function of no arguments, a Function0. At the call site the compiler wraps the argument expression in a closure; inside the method every mention of the parameter becomes a call to that closure's apply. Nothing is cached, which is why the by-name version above prints twice.

Two consequences follow. First, the closure allocation costs something on hot paths, although the JIT often removes it when the call is inlined. Second, by-name parameters let ordinary methods behave like language syntax. The block passed to retry below is not run at the call site; retry runs it, and re-runs it on failure. Closures in general are covered in Scala Closures.

import scala.util.control.NonFatal

// A control structure: the block is not run until retry decides to run it.
def retry[A](attempts: Int, backoffMs: Long)(block: => A): A =
  try block
  catch
    case NonFatal(e) if attempts > 1 =>
      Thread.sleep(backoffMs)
      retry(attempts - 1, backoffMs * 2)(block)   // pass the thunk on, unevaluated

// A logger that never builds the message when the level is off.
final class Log(enabled: Boolean):
  def debug(msg: => String): Unit = if enabled then println(msg)

val user = retry(3, 100)(httpGetUser(42))         // re-runs the HTTP call on failure
Log(enabled = false).debug(s"state = ${expensiveDump()}")  // expensiveDump never runs

The logger pattern is the most common real-world win. A strict debug(msg: String) builds the interpolated string, and whatever expensiveDump() does, even when debug logging is off. The by-name version skips all of it. The standard library already uses this shape in require(cond, message), whose message is by-name.

Advertisement

Where the standard library is already lazy

Several everyday operations are lazy by design, and reading their signatures tells you which. The boolean operators && and || evaluate their right operand only when needed, and if evaluates only one branch. Option.getOrElse, Map.getOrElse and Option.fold's empty case take their default by-name, so cache.getOrElse(k, loadFromDb(k)) only hits the database on a miss. Try.apply takes its body by-name so it can catch what the body throws.

One signature misleads people. Future.apply also takes its body by-name, but submits it to the execution context immediately: the parameter is by-name so the work can run on another thread, not to postpone it. A Future is eager and memoized.

lazy val internals: Scala 2

A lazy val is call-by-need attached to a field: computed on first read, cached, and safe to read from many threads. Thread safety is the expensive part. In Scala 2 the compiler adds a bitmap field with one bit per lazy val and a private initializer method. A read checks the bit; if it is clear, the initializer enters synchronized on the enclosing instance, checks the bit again, runs the right-hand side, stores the value and sets the bit. This is double-checked locking, made correct by the memory-model guarantees of the monitor and the volatile bitmap.

The flaw is the lock's scope. The monitor is the whole enclosing object, held for the entire time the initializer runs. Any other lazy val in the same object, and any other code that synchronizes on that object, waits behind a slow initializer. Worse, if an initializer touches a lazy val in another object while a second thread does the reverse, each thread holds one monitor and waits for the other.

// Scala 2: each lazy val initializer runs inside the enclosing object's monitor.
object A { lazy val a: Int = { Thread.sleep(50); B.b + 1 } }
object B { lazy val b: Int = { Thread.sleep(50); A.a + 1 } }

// Thread 1 reads A.a: locks A, sleeps, then needs B's lock.
// Thread 2 reads B.b: locks B, sleeps, then needs A's lock.  -> deadlock.
// Even single-threaded this is a cycle with no answer; in Scala 3 the
// reference documents recursive lazy val behaviour as undefined.

Objects add a separate hazard: an object is initialized by JVM class initialization, which has its own lock, so two objects whose constructors reference each other from different threads can deadlock even without lazy vals.

lazy val internals: Scala 3.3 and later

Scala 3.0 to 3.2 already avoided Scala 2's monitor with a bitmap and compare-and-swap scheme; Scala 3.3.0 replaced it with a lighter implementation, on by default. Each lazy val gets one volatile field of type Object whose contents encode its state. null means not initialized. The first reader uses compare-and-swap to install a marker called Evaluating and runs the initializer. A second reader that finds Evaluating swaps in a Waiting object, which is a CountDownLatch(1), and parks on it. When the initializer returns, the value is stored (a NullValue marker stands in for a genuine null) and any latch is released. The marker types live in scala.runtime.LazyVals.

Scala 3.3+ lazy val: one Object field, compare-and-swap transitionsnullnot yet initializedEvaluatingone thread runs the RHSWaitinga CountDownLatch(1)CAS by first readerCAS by 2nd readervalueor NullValue markerRHS returns: storestore, then countDownback to nullRHS threwexceptionwaiters released; next read retriesFast pathone volatile read sees valueNo lock on the enclosing instanceScala 2 instead ran the RHS inside synchronized(this)Readers that find Evaluating park on the latch; nobody holds a monitor while the initializer runs.
State transitions of one lazy val field under the Scala 3.3+ scheme. Only the thread that won the first compare-and-swap runs the initializer.
// Roughly what Scala 3.3+ generates for:  class Conf { lazy val db: Db = connect() }
class Conf:
  @volatile private var db$lzy: AnyRef = null      // null, Evaluating, Waiting, NullValue or the value

  def db: Db =
    val cur = db$lzy
    if cur.isInstanceOf[Db] then cur.asInstanceOf[Db]   // fast path: one volatile read
    else db$lzyINIT()                                   // slow path: CAS loop below

  private def db$lzyINIT(): Db =
    // loop:
    //   null       -> CAS to Evaluating; run connect(); store value; if CAS finds Waiting, countDown()
    //   Evaluating -> CAS to a new Waiting latch, then await()
    //   Waiting    -> await() on the existing latch
    //   value      -> return it
    // if connect() throws: reset to null, release any latch, rethrow
    ???

Because no monitor is held while the initializer runs, unrelated lazy vals in the same instance no longer block each other, and code synchronizing on the instance is not affected. The Scala 3 reference describes the aim as reducing the possibility of deadlock, not eliminating it: a thread waiting on a latch is still waiting, so a cycle of lazy vals that depend on each other across threads can still hang, and the reference states that recursive lazy val behaviour is undefined. Keep initializers acyclic.

The 3.3.0 release notes also record a compatibility escape hatch, -Ylegacy-lazy-vals, for environments where the new scheme caused trouble, GraalVM native images being the case reported at the time. For a lazy val that only ever one thread touches, Scala 3 provides @scala.annotation.threadUnsafe, which generates a plain check-and-assign with no atomics at all.

What laziness costs

Every read of a thread-safe lazy val performs at least one volatile read, which on x86 is nearly free but still prevents some compiler and JIT optimizations, and on weakly ordered CPUs such as ARM costs a load-acquire. In a tight loop, copy the lazy val into a local val once. The first read is far more expensive: a compare-and-swap, the initializer itself, and possibly parking threads.

Laziness also changes memory and failure timing. A pending lazy val or stored by-name thunk keeps everything it captures alive. A strict val that fails to load configuration fails at startup, where an operator sees it; a lazy one fails on the first request that touches it, and because a throwing initializer is not cached, every later request retries the failing work.

Laziness as suspended effects

Effect systems such as Cats Effect and ZIO generalize the idea. An IO[A] is a value that describes a computation; nothing happens until the runtime interprets it. That is call-by-name for whole programs: running the same IO twice performs its effects twice. The benefit is referential transparency, explained in Referential transparency in Scala.

The bug to watch for is accidentally strict construction. IO.pure takes an already-computed value, so any side effect inside its argument runs once, when the IO is built. IO.delay takes its argument by-name and runs it each time. When you do want call-by-need inside an effect system, ask for it explicitly with memoize, which returns an IO[IO[A]]: running the outer effect sets up the cache, and the inner effect computes once and then reuses the result. More on the runtime in Cats Effect.

import cats.effect.{IO, IOApp}

object Demo extends IOApp.Simple:
  val eager: IO[Long]   = IO.pure(System.nanoTime())   // nanoTime ran ONCE, when eager was built
  val delayed: IO[Long] = IO.delay(System.nanoTime())  // nanoTime runs each time the IO runs

  def run: IO[Unit] =
    for
      a <- eager;   b <- eager      // a == b
      c <- delayed; d <- delayed    // c != d
      once <- delayed.memoize       // IO[IO[Long]]: first run computes, later runs reuse
      e <- once;    f <- once       // e == f
      _ <- IO.println((a == b, c == d, e == f))   // (true,false,true)
    yield ()

Worked example: a service that starts fast and fails loudly

Consider a service with three dependencies: a configuration file, a database pool and a rarely used PDF renderer that is slow to warm up. A first version makes all three lazy vals to speed up startup. In production the pool's configuration is wrong, the deploy looks healthy, the first request fails, and every later request retries the pool initializer and floods the database with connection attempts.

The fix is a strategy per dependency. Configuration and the pool are built strictly at startup, and the readiness probe fails until the pool can connect, so a bad deploy never receives traffic. Only the PDF renderer stays lazy, with its failure cached for thirty seconds instead of retried per request. Debug logging takes by-name messages, so the diagnostic dump costs nothing when the level is off.

Failure modes

SymptomCauseFix
A side effect runs twiceby-name parameter used more than oncebind it to a lazy val, or make the parameter strict
Threads hang on startuplazy val or object cycle across threadsbreak the cycle; initialize eagerly in a defined order
A failing dependency is hit on every requestlazy initializer throws and is never cachedfail fast at startup or cache the failure explicitly
An effect runs once instead of each timework placed inside IO.pureuse IO.delay; memoize only on purpose

Trade-offs

Strict evaluation is the easiest to reason about: code runs where it is written, once, and fails at a predictable time. Call-by-name buys control structures and free disabled logging for repeated evaluation and closure allocation. Lazy vals skip unneeded work but add per-read cost and late failures. Effect types make laziness the default and recover predictability through types, at the cost of a runtime.

What to do next

  1. Search your codebase for lazy val and classify each one: needed for cost, needed for initialization order, or habit. Make the habitual ones strict.
  2. Move configuration and connection pools to strict startup initialization behind a readiness check.
  3. Change debug and trace logging helpers to take by-name messages, or use a logging library that already does.
  4. Check every by-name parameter for multiple uses and bind it to a lazy val where it should run once.
  5. If you are still on Scala 2, look for lazy vals in objects that reference other objects' lazy vals and remove the cycles before upgrading to Scala 3.3 or later.
  6. In effectful code, grep for IO.pure( with a side effect in its argument, and use memoize only where caching is intended.
Key takeaway: Scala is strict by default, and every form of laziness is a choice about when code runs, how often and on which thread. By-name parameters are closures re-run at each use; lazy vals are call-by-need with thread-safe initialization that Scala 3.3 rebuilt around compare-and-swap and latches instead of a monitor on the whole instance, which reduces but does not remove deadlock risk. Lazy initializers that throw are retried on every access, and lazy failures surface late. Use laziness deliberately for skipped work and control structures, keep initializers acyclic, and fail fast for anything an operator needs to see at deploy time.