Scala gives you strong types, immutable data and pure functions, and people sometimes conclude that it needs fewer tests. It needs different ones. The compiler proves that a function returning Long returns a Long; it says nothing about whether the discount is applied at ten units or eleven, whether a retry loop gives up, or whether a query returns rows in the order the caller assumes. Tests are how you pin down behaviour the types cannot express.

This article builds a real test stack from first principles: how sbt runs tests, choosing a framework, example-based tests, ScalaCheck properties, effectful code without real sleeps, and integration tests that stay out of the fast loop. One small pricing module is the running example. The code is Scala 3, but everything except the syntax applies to Scala 2.13.

Advertisement

What a test run actually is

When you type test in sbt, three layers cooperate. sbt compiles src/test/scala against the main classes plus the Test-scoped dependencies. It then asks every test framework on the classpath, through the small sbt test-interface API, for its fingerprints: rules such as "any class that extends munit.FunSuite". Every compiled class that matches becomes a task, the framework runs the task, and sbt writes JUnit-style XML reports for CI.

Two defaults matter more than people expect. First, unless you set Test / fork := true, tests run inside sbt's own JVM, sharing its heap, class loaders and system properties, and suites run in parallel with each other. Forking costs a JVM start-up per run but isolates tests from the build tool, and it is the only way to set JVM flags for tests. Second, forked suites run sequentially unless you also set Test / testForkedParallel := true. Once suites run in parallel, any mutable global state, such as a shared in-memory database, a static cache or a System.setProperty call, becomes a race between them.

How sbt finds and runs Scala testssbt Test configcompiles src/test/scalatest-interfacefingerprints: subclass, annotationFrameworkScalaTest / MUnit / ZIO TestdiscovertasksUnit suitespure functions, ms eachProperty suitesScalaCheck: 100+ cases eachEffect suitesIO / ZIO, virtual clockForked test JVM (Test / fork := true)own heap, own system propertiesIntegration suitescontainers, real DB, run separatelysequential unless testForkedParallelslow, tagged, excluded from the fast looprunReportsJUnit XML in target/test-reports, failing seed printed for property tests
sbt discovers suites through framework fingerprints, runs them (in parallel by default) in a forked JVM, and keeps slow integration suites out of the fast loop.
// build.sbt -- versions as published in each project's docs at time of writing; pin to current releases
ThisBuild / scalaVersion := "3.3.4"

libraryDependencies ++= Seq(
  "org.scalameta" %% "munit"            % "1.0.4" % Test,
  "org.scalameta" %% "munit-scalacheck" % "1.2.0" % Test
)

// Run tests in a separate JVM: isolates system properties, heap and static state from sbt itself
Test / fork := true
Test / javaOptions ++= Seq("-Xmx1g", "-Duser.timezone=UTC")
// Forked suites run sequentially by default; opt in to parallelism once suites are isolated
Test / testForkedParallel := true

Choosing a framework

Three frameworks cover almost all Scala codebases, and they differ more in philosophy than in capability. ScalaTest is the oldest and most flexible. It offers many interchangeable styles (AnyFunSuite, AnyFlatSpec, AnyWordSpec, AnyFreeSpec and others) plus a large matcher DSL (result should be (42)). The flexibility is a cost in a team setting, because two people can write two styles in the same repository. Pick one style and enforce it.

MUnit deliberately does less. There is one style: test("name") { ... } with plain assertions such as assertEquals, which print a readable diff of the two values when they differ. It has fixtures, tags, .only and .ignore markers on test names, and integrations for ScalaCheck and cats-effect. ZIO Test is the natural choice if your code is written in ZIO: tests are ZIO values, the environment (including a test clock and test console) is provided as layers, and property testing is built in rather than bolted on. See the ZIO introduction for the effect model it assumes.

ScalaTestMUnitZIO Test
StyleMany styles, matcher DSLOne style, plain assertionsTests are ZIO effects
Property testsVia scalatestplus-scalacheckmunit-scalacheckBuilt in (Gen, check)
Async and effectsAsyncFunSuite for FutureFuture built in; IO via munit-cats-effectNative ZIO, TestClock
Best fitLarge legacy suites, BDD-style specsNew projects, cats-effect stacksZIO applications
Advertisement

Example-based tests that document intent

An example-based test fixes one input and asserts one output. Its main job is to document a decision. The worked example is a pricing function with a bulk discount: fifteen percent off a line once the quantity reaches ten, with all amounts in integer cents so rounding is explicit. Each test below names a rule, not a scenario number, so a failure message tells the reader which business decision broke.

import munit.FunSuite

final case class LineItem(sku: String, unitCents: Long, qty: Int)

object Pricing:
  /** Bulk discount: 15% off a line once qty reaches 10; totals are in cents, discount rounded down. */
  def lineTotal(i: LineItem): Long =
    val gross = i.unitCents * i.qty
    if i.qty >= 10 then gross - gross * 15 / 100 else gross

  def orderTotal(items: List[LineItem]): Long = items.map(lineTotal).sum

class PricingSuite extends FunSuite:
  test("no discount below the threshold") {
    assertEquals(Pricing.lineTotal(LineItem("a", 250, 9)), 2250L)
  }
  test("fifteen percent off at exactly ten units") {
    assertEquals(Pricing.lineTotal(LineItem("a", 250, 10)), 2125L)
  }
  test("empty order costs nothing") {
    assertEquals(Pricing.orderTotal(Nil), 0L)
  }

Notice what the examples cover: the boundary on both sides (9 and 10 units) and the empty case. Boundaries and empties are where defects cluster. Notice also what they do not cover: very large quantities, prices of zero, or whether the order of lines matters. You could keep adding examples, but you would only ever test the cases you thought of. That is the gap property-based testing fills.

Property-based testing with ScalaCheck

A property states something that must hold for every input in a space, and the framework tries to disprove it. ScalaCheck has three parts. A Gen[A] is a recipe for producing random values of A; generators compose with for comprehensions, as the item generator below shows. A Prop is a boolean claim over generated values, built with forAll. The runner evaluates the property on many generated inputs (ScalaCheck's default minimum is 100 successful cases, and you can raise it per suite) and stops at the first counterexample.

import munit.ScalaCheckSuite
import org.scalacheck.{Gen, Prop}
import org.scalacheck.Prop.forAll

class PricingProps extends ScalaCheckSuite:
  // Generators describe the input space; keep them realistic so failures mean something
  val item: Gen[LineItem] = for
    sku  <- Gen.identifier
    unit <- Gen.chooseNum(0L, 1_000_000L)
    qty  <- Gen.chooseNum(0, 10_000)
  yield LineItem(sku, unit, qty)

  property("total never exceeds the undiscounted price") {
    forAll(Gen.listOf(item)) { items =>
      Pricing.orderTotal(items) <= items.map(i => i.unitCents * i.qty).sum
    }
  }

  // Two Int generators rather than one LineItem generator: Int has a Shrink instance, so a failure shrinks
  property("buying more never costs less") {
    forAll(Gen.chooseNum(0L, 1_000_000L), Gen.chooseNum(0, 20)) { (unit, qty) =>
      Pricing.lineTotal(LineItem("x", unit, qty + 1)) >= Pricing.lineTotal(LineItem("x", unit, qty))
    }
  }

  override def scalaCheckTestParameters =
    super.scalaCheckTestParameters.withMinSuccessfulTests(500)

Good properties are rarely "reimplement the function and compare". Useful shapes are: bounds (the discounted total never exceeds the gross), invariance (reordering lines does not change the total), monotonicity (buying one more unit never costs less), round trips (decode after encode returns the original), and oracle comparison against a slow but obviously correct version.

Run the monotonicity property against the pricing function and it fails within a few dozen cases, but only because the quantity generator stops at 20. Drawn from 0 to 10,000, a quantity of exactly 9 would come up about once in 10,000 cases and the bug would usually slip through: a generator has to reach the region where the bug lives, so bias it toward thresholds and edges. The reported counterexample is unit = 2, qty = 9: nine units at 2 cents cost 18 cents, but ten units cost 20 − 3 = 17. Every unit price of 2 cents or more has the same cliff, because a 15% discount on ten units is worth more than one unit. None of the example tests caught it, since each checked one side of the threshold in isolation. The property does not tell you whether the cliff is a bug (customers can pad orders to pay less) or an accepted rule; it forces the team to decide explicitly, and then to encode the decision, either by fixing the pricing or by narrowing the property.

When a property fails, ScalaCheck shrinks the counterexample: it repeatedly tries simpler versions (shorter lists, smaller numbers) that still fail, which is how a failure found at unit = 734,211, qty = 9 gets reported as unit = 2, qty = 9. Shrinking needs a Shrink instance for the input type: numbers, strings and collections have one, but a generator for your own case class does not shrink unless you supply one, which is why the monotonicity property draws two numbers instead of a LineItem. Shrunk values can also fall outside a generator's range, so keep the property itself valid for any input of the type. And because inputs are random, the run prints its seed on failure; MUnit prints a ready-made override def scalaCheckInitialSeed = "..." line that you paste into the suite to replay the exact failure deterministically. Commit the fix, then remove the override, and add the minimal counterexample as a plain example test so it is guarded forever.

Testing effectful code without real time

Code that retries, times out, polls or schedules is the classic source of slow and flaky tests, because the obvious test sleeps for real. With cats-effect, the fix is to run the program on a virtual clock. TestControl from the cats-effect-testkit module executes an IO program on a simulated runtime where IO.sleep advances time instantly and deterministically. The retry test below checks both the outcome and the exact backoff, and it finishes in milliseconds.

import cats.effect.IO
import cats.effect.testkit.TestControl
import scala.concurrent.duration.*

// A retry policy under test: three attempts, doubling backoff from 1 second
def retrying[A](op: IO[A], attempts: Int = 3, delay: FiniteDuration = 1.second): IO[A] =
  op.handleErrorWith { e =>
    if attempts <= 1 then IO.raiseError(e)
    else IO.sleep(delay) *> retrying(op, attempts - 1, delay * 2)
  }

// Inside a munit-cats-effect CatsEffectSuite: TestControl runs the program on a virtual clock,
// so "1s + 2s of backoff" completes instantly and the elapsed time is exact, not approximate.
test("gives up after three attempts and three seconds of backoff") {
  val failing = IO.raiseError[Int](new RuntimeException("down"))
  val program = IO.monotonic.flatMap { start =>
    retrying(failing).attempt.flatMap(r => IO.monotonic.map(end => (r.isLeft, end - start)))
  }
  TestControl.executeEmbed(program).map { case (failed, elapsed) =>
    assert(failed)
    assertEquals(elapsed, 3.seconds)
  }
}

The equivalent in ZIO is TestClock.adjust. The principle is the same in both: make time an input you control. Do the same for randomness and the current date: pass them in rather than reading them inside the function. For how IO schedules fibers and why this simulation is possible, see cats-effect in depth.

Integration tests and the fast loop

Unit and property tests should run in seconds on every save. Tests that talk to a real database, a message broker or an HTTP service belong in a separate tier. The common Scala pattern is a real dependency started in a throwaway container (the testcontainers-scala library wraps Testcontainers for both ScalaTest and MUnit), a suite-level fixture that starts it once and tears it down after the last test, and a tag or separate sbt configuration so the fast loop excludes it.

Two rules keep this tier honest. First, each suite owns its data: create a unique schema or key prefix per suite, so parallel suites never see each other's rows. Second, keep a small number of end-to-end paths rather than duplicating every unit case against the real system; the integration tier exists to prove the wiring, SQL and serialisation, not the business rules already covered below it. If you use Doobie, its check support can type-check your SQL against the live schema in exactly this tier; see Doobie in depth.

Failure modes that make suites flaky or slow

SymptomUsual causeFix
Passes alone, fails in the full runShared mutable state between parallel suites (static caches, one database)Isolate data per suite; use fixtures; disable parallelism only to confirm the diagnosis
Fails around midnight or in CI onlyReads the system clock or default time zoneInject the clock; set -Duser.timezone=UTC in forked tests
Random property failure nobody can reproduceSeed not capturedCopy the printed seed into scalaCheckInitialSeed, fix, then add an example test
Test hangs instead of failingAwait on a starved execution context, missing timeoutReturn Future or IO from the test; set a suite timeout
Green suite, broken productionMocks that encode wrong assumptionsPrefer fakes with real behaviour; add a thin integration test for each boundary

Mocking deserves its own warning. Mocking libraries exist for Scala, but idiomatic Scala usually avoids them: if a service depends on a trait such as trait PriceRepo { def find(sku: String): IO[Option[Long]] }, an in-memory implementation backed by a Map takes five lines, behaves like the real thing, and cannot silently drift into asserting call counts. Tagless-final code makes this especially natural; see tagless final in Scala.

What to do next

  1. Pick one framework and one style per module, and write the choice into the contributing guide.
  2. Set Test / fork := true and a fixed time zone in Test / javaOptions.
  3. For each core pure function, write the boundary and empty-input examples, then add at least one property: a bound, an invariance or a round trip.
  4. When a property fails, replay it with the printed seed, fix the bug, and keep the shrunk counterexample as an example test.
  5. Replace every real sleep in tests with a virtual clock (TestControl or TestClock) and inject clocks and randomness into production code.
  6. Move container-backed tests behind a tag or configuration so the fast loop stays under a minute, and run them on every merge in CI.
Key takeaway: Types prove shapes, tests pin down behaviour. Choose one framework, fork the test JVM, and isolate state so parallel suites cannot interfere. Use example tests to document decisions at boundaries, and ScalaCheck properties to find the inputs you did not imagine, replaying failures by seed and keeping the shrunk case forever. Control time instead of sleeping, and keep integration tests in their own tier. For the build tool that ties it together, see <a href="scala_sbt.html">sbt in depth</a>.