Scala Native is an ahead-of-time compiler and runtime that turns Scala code into a native executable, with no JVM at run time. The binary starts in milliseconds, has a small resident footprint, and can call C functions directly. It is the right tool for command-line programs, small daemons, sidecars and anything that needs to link against a C library; it is usually the wrong tool for a long-running service that the JVM's JIT would make faster over hours of warm execution.
This article explains how the compiler works end to end, because almost every surprise you will meet, from link errors to a missing library, follows from that pipeline. It then covers configuration, C interop, memory, threads, a worked example, failure modes and a checklist. Version-specific details below refer to the 0.5 series; 0.5.12, released in May 2026, was the latest at the time of writing.
What Scala Native is, and what it is not
Scala Native is a separate backend for the Scala compiler, in the same way Scala.js is. Your code is ordinary Scala 2.12, 2.13 or Scala 3; what changes is what the compiler emits and what runtime sits underneath. It is not a JVM, and it is not GraalVM native-image. Native-image takes JVM bytecode and a full JDK class library and compiles them under a closed-world assumption. Scala Native never sees bytecode: it has its own intermediate representation, its own reimplementation of the parts of the Java standard library that Scala programs commonly need (called javalib), its own garbage collectors, and it uses LLVM as the code generator.
That design has three consequences worth holding on to. First, only libraries that publish Scala Native artifacts can be used; a plain JVM jar will not link. Second, the Java library surface is a subset: core parts of java.lang, java.util, java.util.concurrent, java.io and java.nio, but not javax and not general reflection. Third, you gain something the JVM cannot offer cheaply: direct, zero-overhead calls into C, including unmanaged pointers and structs.
The compilation pipeline
Compilation happens in two halves. In the first, scalac runs with the Scala Native compiler plugin, which lowers each class to NIR, the Native Intermediate Representation, and writes .nir files next to the usual class files. NIR is a typed, SSA-style representation that still knows about classes, traits, virtual calls and boxing, so the later optimiser can reason about Scala semantics. Published libraries ship NIR inside their jars; that is what the _native0.5 suffix on an artifact name signals.
The second half starts at nativeLink. The linker loads NIR for your code and every dependency, then walks the program from the entry point, marking every class, method and field that is reachable. Anything unreachable is dropped. This is whole-program analysis: because the linker sees the complete set of classes, it can prove that a trait has one implementation and turn virtual calls into direct calls, inline across library boundaries and remove unused fields. The optimiser then emits LLVM IR as .ll files, clang compiles them according to the build mode and LTO setting, and the result is linked with the native runtime, which contains the garbage collector and C parts of javalib.
Two practical points follow. Incremental compilation helps only the first half; linking reruns on every change and dominates release build time. And because reachability is computed from main, a call to a Java method that javalib does not implement is reported as a missing definition at link time, not at run time, so you find the gap on your machine instead of in production.
Setting up a build
In sbt you add the plugin and enable it on the project. Dependencies use %%% so sbt resolves the artifact compiled for the current platform; the same build can cross-compile a JVM and a Native variant of shared code. The sbt article covers cross-projects in general, and Mill has an equivalent Scala Native module if you prefer it.
// project/plugins.sbt
addSbtPlugin("org.scala-native" % "sbt-scala-native" % "0.5.12")
// build.sbt
import scala.scalanative.build._
lazy val logscan = project
.in(file("."))
.enablePlugins(ScalaNativePlugin)
.settings(
scalaVersion := "3.3.x", // replace with your Scala release
// %%% picks the artifact built for Scala Native, not the JVM one;
// replace x.y.z with the current os-lib release
libraryDependencies += "com.lihaoyi" %%% "os-lib" % "x.y.z",
nativeConfig ~= {
_.withMode(Mode.releaseFast)
.withGC(GC.immix)
.withLTO(LTO.thin)
}
)The nativeConfig setting holds the toolchain options. The three that matter most are the build mode, the garbage collector and link-time optimisation, summarised below. You also need a working clang and clang++ on the path; Scala Native does not bundle LLVM. The build target decides whether you produce an executable or a library that C code can link against.
| Option | Values (default first) | Use it for |
|---|---|---|
| mode | debug, release-fast, release-size, release-full | debug for the edit loop; release-fast for most shipped binaries; release-full when you have measured the gain |
| gc | immix, commix, boehm, none | immix for most programs; commix when collection pauses matter on multi-core; none only for short-lived tools |
| lto | none, thin, full | thin to inline across the Scala/C boundary at acceptable build cost |
| buildTarget | application, libraryDynamic, libraryStatic | libraries when C, Rust or Python code will call your Scala |
sbt also provides nativeLinkReleaseFast and nativeLinkReleaseFull as shortcuts, so a CI job can build a release binary without changing the checked-in default of debug. Scala CLI supports the platform as well, which is the fastest way to try a single-file program.
Calling C: extern objects, pointers and zones
Interop is the feature that most distinguishes Scala Native. You declare C functions as members of an object annotated @extern, with bodies of extern. Types from scala.scalanative.unsafe map directly onto C: CInt is int, CSize is size_t, CString is char*, and Ptr[T] is an unmanaged pointer. Unsigned integers live in scala.scalanative.unsigned. @link("z") tells the linker to add -lz.
import scala.scalanative.unsafe._
import scala.scalanative.unsigned._
@extern
object libc {
def getpid(): CInt = extern
def strlen(s: CString): CSize = extern
def getenv(name: CString): CString = extern
}
@link("z") // pass -lz to the linker
@extern
object zlib {
def zlibVersion(): CString = extern
}
object Main {
def main(args: Array[String]): Unit = {
println(s"pid = ${libc.getpid()}")
println(s"zlib = ${fromCString(zlib.zlibVersion())}")
Zone { // Scala 3 syntax
val key = toCString("HOME") // allocated in the zone
val home = libc.getenv(key)
if (home != null) println(fromCString(home))
} // zone memory freed here
val buf = stackalloc[Byte](256) // valid until main returns
buf(0) = 0.toByte
println(libc.strlen(buf)) // 0
}
}Memory handed to C must not move and must not be collected while C is using it, so interop code allocates outside the garbage-collected heap. There are two tools. stackalloc reserves memory in the current stack frame; it is freed automatically when the method returns, which makes it ideal for small buffers passed to a single call. A Zone is a region: everything allocated inside it, including strings converted with toCString, is freed when the block exits. In Scala 3 you write Zone { ... }; Scala 2 uses Zone.acquire { implicit z => ... }.
The rule that keeps you out of trouble is lifetime discipline. Never return a pointer from a zone or a stack allocation, never store one in a Scala object that outlives the block, and convert C strings to Scala strings with fromCString before the owning memory goes away. For long-lived native resources, wrap the pointer in a class with an explicit close and use it with scala.util.Using, exactly as you would a file handle.
Memory and garbage collection
Managed Scala objects live on a heap owned by one of the runtime's collectors. Immix, the default, is a mark-region collector: the heap is divided into blocks and lines, allocation bumps a pointer through free lines, and collection marks live objects and reclaims whole lines. It is mostly precise, which means it knows where pointers are in heap objects and scans stacks conservatively. Commix is based on Immix but marks in parallel and sweeps concurrently, so it suits multi-threaded programs where pause time matters. Boehm is the conservative collector from the C world, kept mainly for compatibility. none allocates and never frees; for a tool that runs for two seconds and exits, that is legitimately the fastest option.
Two limits on what the collector can see cause real bugs. It does not scan memory you obtained from malloc, a zone or stackalloc, so storing the only reference to a Scala object inside native memory lets the collector free it. And C libraries that call back into Scala from their own threads must do so on threads the runtime knows about. Keep managed and unmanaged worlds separate, and pass plain data across the boundary. JVM flags such as -Xmx do not apply; check your version's documentation for heap settings.
Threads and concurrency since 0.5
Before 0.5, Scala Native programs were single-threaded. The 0.5.0 release in April 2024 added multithreading on platform threads: java.lang.Thread, synchronized with object monitors, JVM-compatible @volatile, and thread-safe implementations of most of java.util.concurrent, including atomics and thread pools, plus most of scala.concurrent. 0.5.12 ships an experimental implementation of virtual threads on Linux and macOS; treat it as experimental. Effect libraries that run on the JVM's thread pools, such as Cats Effect, depend on this support, so check each library's release notes for a native build targeting 0.5.
import java.util.concurrent.{Executors, TimeUnit}
import java.util.concurrent.atomic.LongAdder
// Scans files in parallel on platform threads (Scala Native 0.5+).
def countErrors(paths: Seq[os.Path], threads: Int): Long = {
val total = new LongAdder
val pool = Executors.newFixedThreadPool(threads)
paths.foreach { p =>
pool.execute { () =>
os.read.lines.stream(p).foreach(l => if (l.contains(" ERROR ")) total.increment())
}
}
pool.shutdown()
pool.awaitTermination(10, TimeUnit.MINUTES)
total.sum()
}Code like this is identical on the JVM and on Scala Native, which is the point: concurrency primitives are standard, so shared modules can cross-compile. The difference is in the memory model. Scala Native aims for JVM semantics, but a program that happens to work on the JVM through a data race is not guaranteed to behave the same way under a different compiler and optimiser. Use proper synchronisation and atomics, not luck.
Worked example: a log-scanning CLI
A platform team runs a JVM tool that scans gzip-compressed application logs on every host, looking for error bursts. It runs from cron each minute on a few thousand machines. JVM start-up and warm-up make each run take noticeably longer than the scan itself, and the JVM's resident memory competes with the application on small hosts.
The team moves the tool to Scala Native. The parsing code, written with standard collections and os-lib, cross-compiles unchanged. Decompression goes through zlib by @link, replacing a Java library that had no native artifact. They build with release-fast, thin LTO and Immix, and ship one binary per architecture, built on matching CI runners because the output links against the host C library.
They measure before deciding, and so should you: run both versions under hyperfine on a representative log set and compare wall time and peak resident memory. For short runs the native binary wins on start-up; for an hour-long job the JVM's JIT could win on throughput. Inner loops benefit from the habits in Scala collections performance: avoid boxing, prefer arrays and builders.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Library has no native artifact | Dependency resolution fails for the _native0.5 variant | Find an alternative, contribute a cross-build, or use a C library through interop |
| Missing javalib method | Link reports missing definitions and the call path | Replace the call; javalib is a subset of the JDK |
| Reflection-based library | Link or runtime failure in code using Class.forName or bean introspection | Choose libraries that derive codecs at compile time |
| Undefined C symbol | Native linker error naming the function | Add @link, install the -dev package, set linking options |
| Use after free | Crash or garbage output after a zone exits | Copy data out with fromCString before the zone closes |
| Object held only by native memory | Intermittent corruption after a GC | Keep a managed reference, or pass data, not objects |
| Slow release builds | CI time grows with release-full and full LTO | release-fast and thin LTO by default; full modes only where measured |
| Wrong target libc | Binary fails to start on older hosts | Build on the oldest supported OS or link statically where licensing allows |
Trade-offs against the JVM and native-image
Choose Scala Native when start-up time, memory footprint or C interop dominate: CLIs, hooks, container init steps and wrappers around C libraries. Stay on the JVM for long-running services, where the JIT and mature tooling win. GraalVM native-image sits between them: it keeps the JVM library ecosystem, including much reflective code with configuration, but builds are slower and heavier, and C interop is less direct. If you already depend on many Java libraries, try native-image first; if your code is mostly Scala and you want tight native integration, Scala Native is simpler. The Scala overview places all three platforms in the wider language picture.
What to do next
- Install clang, create a sample project with the sbt plugin, and run a hello-world through nativeLink in debug mode.
- List your dependencies and check each for a published _native0.5 artifact before committing to a port.
- Move shared logic into a cross-project so the JVM and native builds run the same tests.
- Wrap one C function you need with @extern, using Zone and stackalloc, and write a test that runs it in a loop to shake out lifetime bugs.
- Build with release-fast and thin LTO in CI; keep debug as the local default.
- Benchmark start-up, throughput and peak memory against the JVM build with hyperfine before switching production.
- Build release binaries per OS and architecture on the oldest platform you support.