Scala's Array looks like one more collection, with map, filter and foldLeft like List or Vector. It is not a collection class at all. An Array[Int] is exactly a JVM int[], an Array[String] is a String[], and the collection methods are added from outside. That design gives Scala arrays the speed and Java compatibility of raw JVM arrays, but it also brings JVM array behaviour into Scala: identity equality, unreadable toString, mutability and runtime element types that generic code has to supply explicitly.
This article explains how arrays are represented, how creation works in generic code, what happens when you call collection methods on an array or pass one where a Seq is expected, the traps that catch experienced developers, and how to write array code that is actually fast. Snippets are labelled Scala 3 or Scala 2.13 where the syntax differs.
What an Array is on the JVM
The compiler maps Array[T] directly to the JVM array type for T. Primitive element types become primitive arrays with values stored inline: Array[Int] is int[], Array[Double] is double[], Array[Boolean] is boolean[]. Reference types become object arrays holding pointers. Array[Any] and Array[AnyRef] both become Object[], and any Int stored in one is boxed into a java.lang.Integer.
That difference is the most important performance fact about arrays. A million-element Array[Int] is one object of about 4 MB. A million Int values in an Array[Any] are a 4-8 MB pointer array plus up to a million separate Integer objects, each with its own header, scattered across the heap. Reading element i becomes a pointer chase, and the garbage collector has a million more objects to trace. Figure 1 shows the three common layouts.
Creating arrays
There are several constructors, and choosing the right one avoids both boxing and needless copies:
// Scala 2.13 and Scala 3
val a = Array(3, 1, 4, 1, 5) // int[] with these values
val zeros = new Array[Double](1024) // double[1024], all 0.0
val names = new Array[String](3) // String[3], all null
val filled = Array.fill(4)(scala.util.Random.nextInt(10)) // evaluates the expression 4 times
val squares = Array.tabulate(10)(i => i * i)
val grid = Array.ofDim[Double](3, 4) // double[3][4], array of row arrays
val range = Array.range(0, 100, 5) // 0, 5, ..., 95
val copy = a.clone() // shallow copy, same element typeNote that new Array[String](3) is full of null, not empty strings. Reference arrays always start as null, so either fill them immediately or prefer Array.fill or Array.tabulate, which never expose a null slot. clone is shallow: an array of arrays or of mutable objects shares its elements with the copy.
Generic code needs a ClassTag
On the JVM an array must know its element type when it is created, because new int[n] and new String[n] are different bytecode instructions. Scala generics are erased: inside def make[T](n: Int) there is no runtime record of T. So this does not compile:
def make[T](n: Int, x: T): Array[T] = Array.fill(n)(x)
// error: No ClassTag available for TThe fix is a ClassTag context bound. The compiler supplies a value that records the runtime class at each call site, and Array uses it to create the right primitive or reference array:
import scala.reflect.ClassTag
def make[T: ClassTag](n: Int, x: T): Array[T] = Array.fill(n)(x)
make(3, 1.5) // values stored unboxed in a double[]
make(3, "a") // String[]Any method that builds a new array needs one: map on an array, toArray on a collection, Array.ofDim. The tag propagates: if your generic method calls one of these, it needs a ClassTag bound too, and so does its caller, up to the point where the type is concrete.
Generic code that only reads arrays avoids the tag but pays another cost. Inside a method with an unknown T, the compiler cannot emit a primitive load instruction, so arr(i) goes through a runtime helper that checks which kind of array it has and boxes primitive results. In a hot loop that can be much slower than the same loop on a concrete Array[Int]. When performance matters, write the inner loop against concrete element types, or let the JIT inline and specialise it and measure the result.
Collection methods: ArrayOps and wrapping
Because Array is not a Scala collection class, calls such as a.map(_ * 2) work through an implicit conversion to ArrayOps, a value class that implements the collection operations directly on the array. In Scala 2.13 and Scala 3 these operations return arrays: a.map(_ * 2) is a new Array[Int], and a.filter(_ > 2) is a new array too. Each step in a chain allocates a full intermediate array:
// three intermediate arrays
val r1 = a.map(_ * 2).filter(_ > 2).map(_ + 1)
// one pass, one result array, via a lazy view
val r2 = a.view.map(_ * 2).filter(_ > 2).map(_ + 1).toArrayPassing an array where a Seq is expected uses a different route. In Scala 2.13, scala.Seq means immutable.Seq. The implicit conversion from an array to an immutable sequence copies the array and is deprecated, so the compiler warns. To share the array without copying, wrap it explicitly and accept responsibility for not mutating it afterwards:
import scala.collection.immutable.ArraySeq
val safe: Seq[Int] = a.toIndexedSeq // copies, safe
val shared: Seq[Int] = ArraySeq.unsafeWrapArray(a) // no copy; do not mutate a afterwardsScala 2.12 behaved differently: arrays were wrapped in a mutable WrappedArray without copying, and scala.Seq was the general, possibly mutable, collection.Seq. Code that relied on aliasing behaves differently after migrating to 2.13. For a broader view of the collection hierarchy, see Scala collections.
Equality, printing and hashing traps
JVM arrays inherit equals, hashCode and toString from Object. So == on arrays compares references, not contents:
Array(1, 2) == Array(1, 2) // false: two different objects
Array(1, 2).sameElements(Array(1, 2)) // true
java.util.Arrays.equals(Array(1, 2), Array(1, 2)) // true
println(Array(1, 2)) // something like [I@6d06d69c
println(Array(1, 2).mkString("[", ", ", "]")) // [1, 2]
case class Point(coords: Array[Double])
Point(Array(1.0)) == Point(Array(1.0)) // false: case-class equality uses the array's equals
Set(Array(1), Array(1)).size // 2The case-class trap is the expensive one. Case classes generate equals and hashCode from their fields, and an array field contributes identity, so value-equal records are unequal, hash maps miss, and deduplication silently fails. Keep arrays out of case classes used as values, or replace them with ArraySeq or Vector in the public field. java.util.Arrays.deepEquals handles nested arrays.
Worked example: a deduplication bug
A service receives sensor readings and drops duplicates before storing them. Readings are modelled as case class Reading(sensor: String, values: Array[Double]) and deduplicated with readings.distinct. In testing, a replayed batch of 10,000 readings is stored twice: distinct removed nothing.
The cause is the array field. The generated equals compares values with the array's own equals, which is reference identity, and hashCode uses the identity hash. Two readings decoded from the same bytes hold two different arrays, so they are never equal. The fix is to change the field type, not the dedup code:
import scala.collection.immutable.ArraySeq
case class Reading(sensor: String, values: ArraySeq[Double])
def decode(sensor: String, raw: Array[Double]): Reading =
Reading(sensor, ArraySeq.unsafeWrapArray(raw)) // raw is not touched after this
// distinct now compares contents, element by elementHot numeric code can still read values(i) cheaply, because ArraySeq.ofDouble keeps a primitive double[] underneath.
Variance and mutability
Java arrays are covariant: a String[] can be assigned to an Object[] variable, and storing an Integer through that variable compiles and then throws ArrayStoreException at runtime. Scala closes that hole by making Array[T] invariant: an Array[String] is not an Array[Any], so the bad store is a compile error. The price is friction when calling methods that take Array[Any]; you convert explicitly or make the method generic.
Arrays are always mutable, so returning one from an API hands every caller the ability to change your internal state. Scala 3 adds IArray[T], an opaque type over Array[T] with no update method. At runtime it is the same JVM array with no wrapper or copy, but the type system stops writes:
// Scala 3
val xs: IArray[Int] = IArray(1, 2, 3)
xs(0) // 1
// xs(0) = 5 // does not compile: IArray has no update
val raw: Array[Int] = Array(4, 5, 6)
val frozen = IArray.unsafeFromArray(raw) // no copy; the caller must stop mutating rawIn Scala 2.13 the closest equivalent is ArraySeq.unsafeWrapArray, which adds a small wrapper object. In both cases the word unsafe means the original array must not be mutated after wrapping.
Layout and fast loops
Multidimensional arrays are arrays of rows. Array.ofDim[Double](1000, 1000) is one array of 1,000 references plus 1,000 separate row arrays. Each access m(i)(j) loads a row pointer and then an element, and rows need not be contiguous. For numeric work a flat array with manual indexing keeps the data contiguous and makes memory access predictable:
final class Matrix(val rows: Int, val cols: Int):
val data = new Array[Double](rows * cols)
inline def apply(i: Int, j: Int): Double = data(i * cols + j)
inline def update(i: Int, j: Int, v: Double): Unit = data(i * cols + j) = v
def sum(xs: Array[Double]): Double =
var acc = 0.0
var i = 0
while i < xs.length do
acc += xs(i)
i += 1
accThe while loop above compiles to the same bytecode shape as a Java for loop. xs.sum and xs.foldLeft(0.0)(_ + _) are convenient and often fast enough, but closures over primitive types can box, and the JIT does not always remove that. Iterate rows in the order they are stored: for a flat row-major matrix, keep j in the inner loop. Swapping the loops can make a large matrix traversal several times slower purely from cache misses. Do not trust intuition here: benchmark with JMH, because a naive System.nanoTime loop mostly measures JIT warm-up. The collection performance guide compares array costs with the other collections.
For bulk copying use System.arraycopy or Array.copy, which become a native memory copy for same-typed arrays. For sorting primitives, java.util.Arrays.sort on an Array[Int] sorts in place without boxing, while a.sorted allocates a new array and uses an Ordering. scala.util.Sorting.quickSort also sorts in place.
Java interop and varargs
Since Scala arrays are Java arrays, they pass to Java APIs with no conversion: Files.readAllBytes returns an Array[Byte] and String.split returns an Array[String]. Spreading an array into a varargs parameter has different syntax in the two language versions:
def total(xs: Int*): Int = xs.sum
val nums = Array(1, 2, 3)
total(nums*) // Scala 3
total(nums: _*) // Scala 2.13 (also accepted by Scala 3 for migration)Java methods that take Object... expect a reference array. Passing an Array[Int] to one, for example to String.format, does not spread it; it becomes a single argument. Convert with nums.map(Int.box) first, or pass the values individually.
Failure modes
- Silent equality failures. Arrays in case classes, set elements or map keys compare by identity, so lookups miss and duplicates survive.
- Accidental boxing.
Array[Any], aClassTaginferred asAny(for examplemake(3, if ok then 1 else "a")), or generic hot loops turn numeric code into pointer chasing and GC load. - Aliasing bugs.
unsafeWrapArray,IArray.unsafeFromArrayor returning an internal array lets someone else mutate data you thought was fixed. - Nulls.
new Array[String](n)is full of nulls that surface asNullPointerExceptionfar from the allocation. - Intermediate allocations. Long
mapandfilterchains on large arrays allocate one full array per step; useviewor a single loop. - Shallow clones.
cloneon anArray[Array[Double]]copies only the row references. - Index errors. Out-of-range access throws
ArrayIndexOutOfBoundsException; useliftfor optional access when the index comes from input.
When to use something else
| Need | Use | Why |
|---|---|---|
| Numeric hot loops, buffers, Java interop | Array | Primitive storage, no wrapper, native copies |
| Immutable data in a public API | IArray (Scala 3) or ArraySeq | Array speed with no mutation through the API |
| Growable mutable sequence | ArrayBuffer | Amortised append over an internal array |
| Immutable sequence with cheap updates | Vector | Persistent structure; see Scala Vector |
| Value semantics in case classes | Vector or ArraySeq | Structural equals and hashCode |
Arrays also shape memory use at the JVM level: large primitive arrays are single big objects that the collector does not have to trace element by element, while large object arrays hold many references to trace. See JVM memory layout for how headers, alignment and compressed pointers affect the byte counts above.
What to do next
- Search your code for
Array[Any],Array[AnyRef]and generic hot loops, and move numeric work to concrete primitive arrays. - Find case classes with array fields and switch the public field to
ArraySeq,VectororIArraywith an explicit equality story. - Replace
==on arrays withsameElementsorjava.util.Arrays.equals, and printing withmkString. - Add
ClassTagbounds where generic code builds arrays, and keep them out of methods that only read. - Turn long
map/filterchains on big arrays into aviewor onewhileloop, then confirm the gain with JMH. - On Scala 3, return
IArrayfrom APIs instead ofArray; on 2.13, returnArraySeq.