Scala's Array looks like one more collection, with map, filter and foldLeft like List or Vector. It is not a collection class at all. An Array[Int] is exactly a JVM int[], an Array[String] is a String[], and the collection methods are added from outside. That design gives Scala arrays the speed and Java compatibility of raw JVM arrays, but it also brings JVM array behaviour into Scala: identity equality, unreadable toString, mutability and runtime element types that generic code has to supply explicitly.

This article explains how arrays are represented, how creation works in generic code, what happens when you call collection methods on an array or pass one where a Seq is expected, the traps that catch experienced developers, and how to write array code that is actually fast. Snippets are labelled Scala 3 or Scala 2.13 where the syntax differs.

Advertisement

What an Array is on the JVM

The compiler maps Array[T] directly to the JVM array type for T. Primitive element types become primitive arrays with values stored inline: Array[Int] is int[], Array[Double] is double[], Array[Boolean] is boolean[]. Reference types become object arrays holding pointers. Array[Any] and Array[AnyRef] both become Object[], and any Int stored in one is boxed into a java.lang.Integer.

That difference is the most important performance fact about arrays. A million-element Array[Int] is one object of about 4 MB. A million Int values in an Array[Any] are a 4-8 MB pointer array plus up to a million separate Integer objects, each with its own header, scattered across the heap. Reading element i becomes a pointer chase, and the garbage collector has a million more objects to trace. Figure 1 shows the three common layouts.

Same values, different memory: Array[Int] vs Array[Integer-boxed] vs Array[Array[Double]]Array[Int] -> JVM int[]: one object, values inline, 4 bytes eachheaderlength31415Array[Any], or a ClassTag inferred as Any -> Object[]: one pointer per slotheaderlengthrefrefrefrefrefInteger3Integer1Integer4Integer1Integer5each element: separate heap object,pointer chase on every read, GC pressureArray.ofDim[Double](3, 4) -> double[][]: an array of row arraysrows3 refsrow 0: double[4]row 1: double[4]row 2: double[4]rows may sit anywhere on the heap;a flat Array[Double](rows * cols) keepsthem contiguous and cache-friendly
Figure 1. A primitive array stores values inline. An object array stores references to separately allocated boxes. A two-dimensional array is an array of independent row arrays.

Creating arrays

There are several constructors, and choosing the right one avoids both boxing and needless copies:

// Scala 2.13 and Scala 3
val a = Array(3, 1, 4, 1, 5)                 // int[] with these values
val zeros = new Array[Double](1024)          // double[1024], all 0.0
val names = new Array[String](3)             // String[3], all null
val filled = Array.fill(4)(scala.util.Random.nextInt(10))   // evaluates the expression 4 times
val squares = Array.tabulate(10)(i => i * i)
val grid = Array.ofDim[Double](3, 4)         // double[3][4], array of row arrays
val range = Array.range(0, 100, 5)           // 0, 5, ..., 95
val copy = a.clone()                         // shallow copy, same element type

Note that new Array[String](3) is full of null, not empty strings. Reference arrays always start as null, so either fill them immediately or prefer Array.fill or Array.tabulate, which never expose a null slot. clone is shallow: an array of arrays or of mutable objects shares its elements with the copy.

Advertisement

Generic code needs a ClassTag

On the JVM an array must know its element type when it is created, because new int[n] and new String[n] are different bytecode instructions. Scala generics are erased: inside def make[T](n: Int) there is no runtime record of T. So this does not compile:

def make[T](n: Int, x: T): Array[T] = Array.fill(n)(x)
// error: No ClassTag available for T

The fix is a ClassTag context bound. The compiler supplies a value that records the runtime class at each call site, and Array uses it to create the right primitive or reference array:

import scala.reflect.ClassTag

def make[T: ClassTag](n: Int, x: T): Array[T] = Array.fill(n)(x)

make(3, 1.5)        // values stored unboxed in a double[]
make(3, "a")        // String[]

Any method that builds a new array needs one: map on an array, toArray on a collection, Array.ofDim. The tag propagates: if your generic method calls one of these, it needs a ClassTag bound too, and so does its caller, up to the point where the type is concrete.

Generic code that only reads arrays avoids the tag but pays another cost. Inside a method with an unknown T, the compiler cannot emit a primitive load instruction, so arr(i) goes through a runtime helper that checks which kind of array it has and boxes primitive results. In a hot loop that can be much slower than the same loop on a concrete Array[Int]. When performance matters, write the inner loop against concrete element types, or let the JIT inline and specialise it and measure the result.

Collection methods: ArrayOps and wrapping

Because Array is not a Scala collection class, calls such as a.map(_ * 2) work through an implicit conversion to ArrayOps, a value class that implements the collection operations directly on the array. In Scala 2.13 and Scala 3 these operations return arrays: a.map(_ * 2) is a new Array[Int], and a.filter(_ > 2) is a new array too. Each step in a chain allocates a full intermediate array:

// three intermediate arrays
val r1 = a.map(_ * 2).filter(_ > 2).map(_ + 1)

// one pass, one result array, via a lazy view
val r2 = a.view.map(_ * 2).filter(_ > 2).map(_ + 1).toArray

Passing an array where a Seq is expected uses a different route. In Scala 2.13, scala.Seq means immutable.Seq. The implicit conversion from an array to an immutable sequence copies the array and is deprecated, so the compiler warns. To share the array without copying, wrap it explicitly and accept responsibility for not mutating it afterwards:

import scala.collection.immutable.ArraySeq

val safe: Seq[Int] = a.toIndexedSeq                    // copies, safe
val shared: Seq[Int] = ArraySeq.unsafeWrapArray(a)      // no copy; do not mutate a afterwards

Scala 2.12 behaved differently: arrays were wrapped in a mutable WrappedArray without copying, and scala.Seq was the general, possibly mutable, collection.Seq. Code that relied on aliasing behaves differently after migrating to 2.13. For a broader view of the collection hierarchy, see Scala collections.

Equality, printing and hashing traps

JVM arrays inherit equals, hashCode and toString from Object. So == on arrays compares references, not contents:

Array(1, 2) == Array(1, 2)                 // false: two different objects
Array(1, 2).sameElements(Array(1, 2))      // true
java.util.Arrays.equals(Array(1, 2), Array(1, 2))   // true
println(Array(1, 2))                       // something like [I@6d06d69c
println(Array(1, 2).mkString("[", ", ", "]"))       // [1, 2]

case class Point(coords: Array[Double])
Point(Array(1.0)) == Point(Array(1.0))     // false: case-class equality uses the array's equals
Set(Array(1), Array(1)).size               // 2

The case-class trap is the expensive one. Case classes generate equals and hashCode from their fields, and an array field contributes identity, so value-equal records are unequal, hash maps miss, and deduplication silently fails. Keep arrays out of case classes used as values, or replace them with ArraySeq or Vector in the public field. java.util.Arrays.deepEquals handles nested arrays.

Worked example: a deduplication bug

A service receives sensor readings and drops duplicates before storing them. Readings are modelled as case class Reading(sensor: String, values: Array[Double]) and deduplicated with readings.distinct. In testing, a replayed batch of 10,000 readings is stored twice: distinct removed nothing.

The cause is the array field. The generated equals compares values with the array's own equals, which is reference identity, and hashCode uses the identity hash. Two readings decoded from the same bytes hold two different arrays, so they are never equal. The fix is to change the field type, not the dedup code:

import scala.collection.immutable.ArraySeq

case class Reading(sensor: String, values: ArraySeq[Double])

def decode(sensor: String, raw: Array[Double]): Reading =
  Reading(sensor, ArraySeq.unsafeWrapArray(raw))   // raw is not touched after this

// distinct now compares contents, element by element

Hot numeric code can still read values(i) cheaply, because ArraySeq.ofDouble keeps a primitive double[] underneath.

Variance and mutability

Java arrays are covariant: a String[] can be assigned to an Object[] variable, and storing an Integer through that variable compiles and then throws ArrayStoreException at runtime. Scala closes that hole by making Array[T] invariant: an Array[String] is not an Array[Any], so the bad store is a compile error. The price is friction when calling methods that take Array[Any]; you convert explicitly or make the method generic.

Arrays are always mutable, so returning one from an API hands every caller the ability to change your internal state. Scala 3 adds IArray[T], an opaque type over Array[T] with no update method. At runtime it is the same JVM array with no wrapper or copy, but the type system stops writes:

// Scala 3
val xs: IArray[Int] = IArray(1, 2, 3)
xs(0)                   // 1
// xs(0) = 5            // does not compile: IArray has no update
val raw: Array[Int] = Array(4, 5, 6)
val frozen = IArray.unsafeFromArray(raw)   // no copy; the caller must stop mutating raw

In Scala 2.13 the closest equivalent is ArraySeq.unsafeWrapArray, which adds a small wrapper object. In both cases the word unsafe means the original array must not be mutated after wrapping.

Layout and fast loops

Multidimensional arrays are arrays of rows. Array.ofDim[Double](1000, 1000) is one array of 1,000 references plus 1,000 separate row arrays. Each access m(i)(j) loads a row pointer and then an element, and rows need not be contiguous. For numeric work a flat array with manual indexing keeps the data contiguous and makes memory access predictable:

final class Matrix(val rows: Int, val cols: Int):
  val data = new Array[Double](rows * cols)
  inline def apply(i: Int, j: Int): Double = data(i * cols + j)
  inline def update(i: Int, j: Int, v: Double): Unit = data(i * cols + j) = v

def sum(xs: Array[Double]): Double =
  var acc = 0.0
  var i = 0
  while i < xs.length do
    acc += xs(i)
    i += 1
  acc

The while loop above compiles to the same bytecode shape as a Java for loop. xs.sum and xs.foldLeft(0.0)(_ + _) are convenient and often fast enough, but closures over primitive types can box, and the JIT does not always remove that. Iterate rows in the order they are stored: for a flat row-major matrix, keep j in the inner loop. Swapping the loops can make a large matrix traversal several times slower purely from cache misses. Do not trust intuition here: benchmark with JMH, because a naive System.nanoTime loop mostly measures JIT warm-up. The collection performance guide compares array costs with the other collections.

For bulk copying use System.arraycopy or Array.copy, which become a native memory copy for same-typed arrays. For sorting primitives, java.util.Arrays.sort on an Array[Int] sorts in place without boxing, while a.sorted allocates a new array and uses an Ordering. scala.util.Sorting.quickSort also sorts in place.

Java interop and varargs

Since Scala arrays are Java arrays, they pass to Java APIs with no conversion: Files.readAllBytes returns an Array[Byte] and String.split returns an Array[String]. Spreading an array into a varargs parameter has different syntax in the two language versions:

def total(xs: Int*): Int = xs.sum
val nums = Array(1, 2, 3)

total(nums*)        // Scala 3
total(nums: _*)     // Scala 2.13 (also accepted by Scala 3 for migration)

Java methods that take Object... expect a reference array. Passing an Array[Int] to one, for example to String.format, does not spread it; it becomes a single argument. Convert with nums.map(Int.box) first, or pass the values individually.

Failure modes

  • Silent equality failures. Arrays in case classes, set elements or map keys compare by identity, so lookups miss and duplicates survive.
  • Accidental boxing. Array[Any], a ClassTag inferred as Any (for example make(3, if ok then 1 else "a")), or generic hot loops turn numeric code into pointer chasing and GC load.
  • Aliasing bugs. unsafeWrapArray, IArray.unsafeFromArray or returning an internal array lets someone else mutate data you thought was fixed.
  • Nulls. new Array[String](n) is full of nulls that surface as NullPointerException far from the allocation.
  • Intermediate allocations. Long map and filter chains on large arrays allocate one full array per step; use view or a single loop.
  • Shallow clones. clone on an Array[Array[Double]] copies only the row references.
  • Index errors. Out-of-range access throws ArrayIndexOutOfBoundsException; use lift for optional access when the index comes from input.

When to use something else

NeedUseWhy
Numeric hot loops, buffers, Java interopArrayPrimitive storage, no wrapper, native copies
Immutable data in a public APIIArray (Scala 3) or ArraySeqArray speed with no mutation through the API
Growable mutable sequenceArrayBufferAmortised append over an internal array
Immutable sequence with cheap updatesVectorPersistent structure; see Scala Vector
Value semantics in case classesVector or ArraySeqStructural equals and hashCode

Arrays also shape memory use at the JVM level: large primitive arrays are single big objects that the collector does not have to trace element by element, while large object arrays hold many references to trace. See JVM memory layout for how headers, alignment and compressed pointers affect the byte counts above.

What to do next

  1. Search your code for Array[Any], Array[AnyRef] and generic hot loops, and move numeric work to concrete primitive arrays.
  2. Find case classes with array fields and switch the public field to ArraySeq, Vector or IArray with an explicit equality story.
  3. Replace == on arrays with sameElements or java.util.Arrays.equals, and printing with mkString.
  4. Add ClassTag bounds where generic code builds arrays, and keep them out of methods that only read.
  5. Turn long map/filter chains on big arrays into a view or one while loop, then confirm the gain with JMH.
  6. On Scala 3, return IArray from APIs instead of Array; on 2.13, return ArraySeq.
Key takeaway: A Scala Array is a JVM array with collection methods added from outside. That gives inline primitive storage, zero-cost Java interop and the fastest loops in the language, but also identity equality, mutability and a runtime element type that generic code must supply through a ClassTag. Use concrete primitive arrays and plain loops for hot numeric code, wrap arrays as IArray or ArraySeq before exposing them, never put them in value-semantics case classes, and measure performance with JMH rather than guessing.