Referential transparency is the property that makes functional programming worth its learning curve, and it is usually explained badly, either as a slogan (no side effects) or as a definition nobody can apply. The working definition is a test you can run in your head. An expression is referentially transparent if you can replace it with the value it evaluates to, or a name with its definition, anywhere in the program without changing what the program does.

If every expression in a piece of code passes that test, you can refactor it mechanically. Extracting a variable, inlining a helper, caching a result or reordering two independent lines cannot change behaviour. If an expression fails, every such refactor needs to be checked by hand, and the bugs it introduces are the kind that only appear under load. This article shows the test, the common ways Scala code fails it, how effect types such as Cats Effect IO restore it, a worked retry example where the difference is concrete, and where purity is worth the cost and where it is not. It assumes basic Scala; monads in Scala is useful background.

Advertisement

The substitution test

Take any program that names a value and uses the name twice, then write a second program with the name replaced by its defining expression. Compare what they do, meaning both the result and anything observable such as output, writes, network calls or thrown exceptions.

// P
val x = 2 + 3
val r = (x, x)

// P' : x inlined
val r = (2 + 3, 2 + 3)

Both give (5, 5) and nothing else happens, so 2 + 3 is transparent. Now try an expression with an effect.

// P
val s = { println("hello"); "hi" }
val r = (s, s)          // prints once

// P'
val r = ({ println("hello"); "hi" }, { println("hello"); "hi" })   // prints twice

The results are equal but the behaviour differs, because the program prints a different number of times. The block is not transparent. That is the whole idea. Notice that the test is about expressions, not functions or languages. A Scala program is a mix, and the useful question is always which expressions in this code pass.

The substitution test: replace a name with its definition and compare behaviourProgram Pval x = expr; (x, x)Program P'(expr, expr)inline xexpr = 2 + 3P and P' both yield (5, 5): transparentexpr = Future(println("hi"))P prints once, P' prints twice: opaqueexpr = IO.println("hi")both are descriptions; nothing runs yet(x, x).tupled.voidrunning P or P' prints twice in botheffects as valuessame behaviourEnd of the worldIOApp runs the one description you builtTransparent expressions can be inlined, extracted, cached or reordered freely; opaque ones cannot.
The same inlining refactor applied to three expressions. Arithmetic and IO descriptions pass the substitution test; an eagerly started Future fails it, because the number of effects depends on how many times the expression is written.

Pure functions and transparent expressions

A function is pure when calling it with the same arguments always returns the same result and does nothing else observable. Calls to pure functions are transparent expressions: math.max(a, b) can be replaced by its result. Purity has three practical parts.

  • Deterministic: the result depends only on the arguments. Reading the clock, a random generator, a mutable global or an environment variable breaks this.
  • No observable effects: no printing, logging, writing files, mutating shared state or sending requests.
  • Total: it returns a value for every input rather than throwing. A thrown exception is control flow the type does not mention, and moving the call can move where it escapes.

Totality is the part people forget. list.head throws on an empty list, so val h = xs.head moved above an if (xs.nonEmpty) check changes behaviour. xs.headOption returns an Option and stays transparent wherever it goes.

Advertisement

How everyday Scala breaks it

Most transparency bugs in Scala come from a short list of constructs.

ConstructWhy it fails the testTransparent alternative
var and mutable collections shared across callsThe same expression reads different values at different timesReturn new values; keep mutation local (see below) or use Ref
println, logging, file and network callsEffects happen at evaluation, so duplicating or removing a call changes behaviourDescribe them as IO values
throw and partial methods such as .head, .getMoving the expression moves where the exception escapesOption, Either, IO.raiseError
System.currentTimeMillis, Random.nextIntDifferent result on every evaluationPass time and randomness in, or use Clock[IO] and std.Random
Future { ... }Starts running at construction, and memoises its resultIO, or a function returning Future that is called deliberately
lazy val with effectsThe effect happens on first access, wherever that isMake the effect explicit and run it once at startup

The Future row deserves its own section, because it is where Scala developers most often meet the problem in production code. Futures and execution contexts covers the mechanics of Future itself.

Future: eager and memoised

A Future begins executing on its execution context the moment it is constructed, and caches its result. Both properties break substitution.

import scala.concurrent.{Future, ExecutionContext}
given ExecutionContext = ExecutionContext.global

def charge(id: String): Future[Receipt] = Future { gateway.charge(id) }  // network call

// P: one charge, awaited twice
val f = charge("order-1")
for { a <- f; b <- f } yield (a, b)

// P': inlined, two charges
for { a <- charge("order-1"); b <- charge("order-1") } yield (a, b)

In P the customer is charged once; in P' twice. A reviewer seeing P' may well extract the duplicate into a val, a refactor that looks like tidying and changes billing. And if the two calls in P' were first bound to separate vals above the for-comprehension, both charges would start in parallel at declaration, before the comprehension sequences anything. With Future, where an expression is written decides when its effect happens.

Effects as values: Cats Effect IO

An effect system restores transparency by separating describing an effect from running it. An IO[A] is a value that describes a computation which, when run, may perform effects and produce an A. Building it does nothing. Composing it with flatMap builds a larger description. Only the runtime, at the edge of the program, executes it.

import cats.effect.{IO, IOApp}
import cats.syntax.all.*

val hello: IO[Unit] = IO.println("hello")

// P and P' now behave identically: each prints twice when run
val p1 = (hello, hello).tupled.void
val p2 = (IO.println("hello"), IO.println("hello")).tupled.void

def charge(id: String): IO[Receipt] = IO.blocking(gateway.charge(id))

object Main extends IOApp.Simple:
  def run: IO[Unit] = charge("order-1").flatMap(r => IO.println(s"charged $r"))

Now hello means the same thing wherever it appears: a program that prints once each time it is sequenced. Using it twice prints twice; naming it changes nothing. Inlining charge("order-1") or extracting it into a val cannot change how many charges happen, because the count is decided by how the description is sequenced, which is visible in the code. Cats Effect and ZIO cover the runtimes; the principle is the same in both.

Note what has and has not changed. The program still charges a card; purity did not remove the effect. What changed is that every expression in the code is transparent, and the single impure act, running the final IO, happens once in IOApp.

Worked example: retry

Retrying a failed operation shows the practical payoff. With a value that describes an action, retrying means running the same value again, so a generic retry combinator is a few lines.

import scala.concurrent.duration.*

def retry[A](io: IO[A], attempts: Int, delay: FiniteDuration): IO[A] =
  io.handleErrorWith { e =>
    if attempts <= 1 then IO.raiseError(e)
    else IO.sleep(delay) *> retry(io, attempts - 1, delay * 2)
  }

val fetchRate: IO[BigDecimal] = IO.blocking(fxClient.rate("EUR", "USD"))
val robust: IO[BigDecimal] = retry(fetchRate, attempts = 4, delay = 200.millis)

Each retry re-runs fetchRate because fetchRate is a recipe, not a result. Try the same thing with Future.

def retryF[A](f: Future[A], attempts: Int): Future[A] =
  f.recoverWith { case _ if attempts > 1 => retryF(f, attempts - 1) }   // BUG

retryF(Future(fxClient.rate("EUR", "USD")), 4)

This compiles and never retries. The Future ran once and memoised its failure, so every recoverWith sees the same failed value. The fix is to pass () => Future[A] and call it per attempt, which is precisely the move from value to description that IO makes the default. Under the transparent version the combinator is reusable for any action, and its behaviour can be read off its definition.

One caveat applies to both: retry is only safe for idempotent operations. Transparency tells you exactly how many times charge will run; it does not make running it twice harmless.

Local mutation is allowed

Transparency is a property of an expression as seen from outside. A function may use a mutable buffer internally and still be pure, as long as the mutation cannot escape or be observed.

def csvLine(fields: Seq[String]): String =
  val sb = new StringBuilder          // local, never escapes
  fields.zipWithIndex.foreach { (f, i) =>
    if i > 0 then sb.append(',')
    sb.append('"').append(f.replace("\"", "\"\"")).append('"')
  }
  sb.toString

Callers cannot tell that csvLine mutated anything, so calls to it are transparent. Use local mutation for performance in hot paths. When state really must be shared, as with a counter used by many fibres, put it in Ref[IO, A]: Ref.of[IO, Int](0) is itself an IO that creates a fresh reference each time it runs, which is transparent, and updates through ref.update(_ + 1) are effects described as values.

Testing and reasoning benefits

Pure functions test with plain assertions: inputs in, outputs compared, no mocks for time or randomness because those are parameters. Effectful code written as IO can be run under a test runtime, such as Cats Effect's TestControl, which advances virtual time, so the retry above can be tested with a four-attempt failure in microseconds rather than in real seconds. See testing in Scala for the frameworks.

Reasoning improves in the same way. Code review becomes local: a reviewer can understand a function from its signature and body, because nothing outside the arguments influences it and nothing it does is hidden. Equational reasoning, rewriting code step by step by replacing equals with equals, becomes a legitimate refactoring tool rather than a hope.

Trade-offs and failure modes

Purity has costs. Effect types add allocation, and tight numeric loops written as IO chains are slower than plain loops, so keep hot inner loops pure and strict and wrap only the boundary. The learning curve is real for teams new to flatMap-heavy code. Stack traces through effect runtimes are harder to read, although both major runtimes now add tracing.

There are also ways to appear pure while not being so:

  • Calling unsafeRunSync() in the middle of library code. It reintroduces eager effects at an arbitrary point; keep it to the program's edge and tests.
  • Wrapping side effects in IO.pure(sideEffect()). The argument is evaluated eagerly; use IO.delay or IO.blocking.
  • Logging with a plain logger inside pure functions. Harmless in practice, but it breaks the test; accept it consciously or use an effectful logger.
  • Mutable Java objects such as java.util.Date returned from a pure function and later mutated by the caller.
  • By-name parameters and lazy val that hide when, or whether, an effect runs.

A pragmatic position for most Scala teams is a pure core with an effectful shell: domain logic as pure functions over immutable data, effects described in IO at the edges, and one runtime entry point. Effect systems compared helps choose the runtime.

What to do next

  1. Pick one module and run the substitution test on every val that holds a Future, a mutable collection or the result of an I/O call; list the failures.
  2. Replace partial calls (.head, .get, throw in domain code) with Option, Either or raiseError.
  3. Pass the clock and randomness in as parameters or capabilities, so time-dependent logic becomes testable.
  4. Convert one Future-based client to IO and rewrite its retry with the combinator above; test it with virtual time.
  5. Search for unsafeRun outside main and tests, and remove each occurrence.
  6. Keep performance-critical inner loops pure and strict; measure before wrapping them in effects.
Key takeaway: An expression is referentially transparent when replacing it with its value or definition cannot change what the program does. In Scala, var, println, exceptions, clocks, randomness and eagerly started Futures fail that test, which makes ordinary refactors silently change behaviour. Effect types such as Cats Effect IO restore it by making effects descriptions that run only at the program's edge, which turns retries, caching and reordering into safe, local edits. Aim for a pure core and an effectful shell, keep mutation local, and run the substitution test whenever a refactor makes you nervous.