An Iterator is the smallest abstraction in the Scala collections library. It is a cursor with two methods, hasNext and next(), that hands out elements one at a time and never goes back. That makes it the right tool for data that is too big to hold, arrives over time, or comes from a resource such as a file or a paged API. It also makes it the easiest collection type to misuse. Most iterator bugs produce no exception, only a wrong number.
This page explains what an iterator promises and what it does not, which operations are lazy and which consume, where hidden buffers appear, and how to build and close iterators safely. It ends with a worked example and a checklist. Memoised lazy sequences are covered separately in Scala LazyList, in depth, and the general machinery of laziness in Scala lazy evaluation. Everything here applies to the Scala 2.13 and Scala 3 standard library, which share one collections implementation.
What an Iterator is, and why it exists
An iterator holds a position in some source. hasNext reports whether another element is available, and next() returns it and advances. Calling next() when hasNext is false throws NoSuchElementException. Unlike a List or Vector, an iterator is not a container. It does not remember elements it has returned, so it can only be traversed once. Iterator extends IterableOnce, the type that methods use when they need to traverse input exactly once, and that is why toList, ++ and many builders accept iterators directly.
The payoff is memory. Processing a 40 GB log with Source.getLines() uses memory proportional to one line, not to the file, as long as nothing in the pipeline holds on to elements. The cost is that you, not the type system, have to make sure each element is read once and in order.
The one rule: use only the iterator a method returns
The Scala documentation states the contract plainly. After calling a method on an iterator, you should discard that iterator and use only the one the method returned, unless the method was next or hasNext. Operations such as map, filter, take and drop return a new iterator that shares the underlying source. Advancing either one disturbs the other, and the library makes no promise about the state of the original.
val it = Iterator(1, 2, 3, 4, 5)
val n = it.size // consumes all five elements to count them
val total = it.sum // 0: the iterator is already exhausted
val it2 = Iterator(1, 2, 3, 4, 5)
val firstThree = it2.take(3) // after take, 'it2' itself is in an UNSPECIFIED state
it2.next() // compiles, but the result is not guaranteed: do not do this
// Correct: use the returned iterator, or materialise once when you need two passes.
val xs = Iterator(1, 2, 3, 4, 5).toVector
val (count, sum) = (xs.size, xs.sum)The first bug is the most common in real code: a method logs an iterator's size for debugging, and the rest of the method sees empty input. Anything that traverses the whole input, such as size, sum, toList or foldLeft, exhausts the iterator. For two passes, materialise once into a Vector, or restructure the work into a single fold.
Lazy, consuming and partially consuming operations
| Operation | Behaviour | Buffers |
|---|---|---|
map, filter, collect, flatMap, zip, zipWithIndex | Lazy: returns a new iterator, does no work yet | No |
take, drop, takeWhile, slice | Lazy; drop(n) skips n elements when first pulled | No |
grouped(n), sliding(n) | Lazy, but each group is materialised as a Seq | One group |
buffered | Lazy; adds head for peeking | One element |
find, exists, contains, indexWhere | Consumes up to and including the first match | No |
size, sum, foldLeft, toList, foreach, mkString | Consumes everything | Result only |
duplicate, span, partition | Two iterators sharing one source | Up to the whole input |
Laziness has a consequence that surprises people coming from List: side effects in map run when elements are pulled, not when map is called. it.map(x => { println(x); x }) prints nothing until a consumer runs, and if the consumer only takes three elements, only three lines are printed. That is usually what you want. It does mean you cannot rely on a mapping step having run before a later statement that does not touch the iterator.
Iterator, LazyList and views compared
| Iterator | LazyList | View | |
|---|---|---|---|
| Traversals | One | Many | Many, recomputed each time |
| Remembers elements | No | Yes, memoised once evaluated | No |
| Memory for a huge stream | Constant | Grows if you keep a reference to the head | Constant per pass |
| Good for | Files, sockets, paged APIs, one-shot pipelines | Recursive or self-referential sequences | Fusing transformations over an existing collection |
The practical rule: if the data comes from outside the program and is read once, use an iterator. If you need to revisit elements, use a strict collection, or a LazyList when the sequence is defined in terms of itself. Allocation and boxing costs across these choices are compared in Scala collections performance.
Hidden buffers: duplicate, span and partition
Sometimes you want two results from one pass. The library offers duplicate, which returns two independent iterators over the same elements, and span and partition, which split elements into two iterators. All three keep a single-pass source behind two consumers, so whatever one consumer has read and the other has not must be stored. In 2.13, partition is built on duplicate with a filter on each side, so it has the same buffering profile.
// Memory trap 1: duplicate, then let the copies drift apart.
val (a, b) = lines.duplicate
val errors = a.count(_.contains("ERROR")) // drains 'a' fully...
val warnings = b.count(_.contains("WARN")) // ...so every line was queued for 'b'
// Same answer, one pass, constant memory:
val (errs, warns) = lines.foldLeft((0, 0)) { case ((e, w), l) =>
(e + (if (l.contains("ERROR")) 1 else 0), w + (if (l.contains("WARN")) 1 else 0))
}
// Memory trap 2: span. Reading 'body' first forces 'header' to be buffered.
val (header, body) = lines.span(_.startsWith("#"))
val headerLines = header.toVector // consume the prefix FIRST: no buffering needed
body.foreach(process)If you drain the first copy completely before touching the second, the queue grows to the size of the input, and a streaming pipeline quietly becomes a full in-memory copy. For counting or aggregating, a single foldLeft with a tuple or case class accumulator does the same work in constant memory. For span, consume the prefix first. In the usual header-then-body case, that means buffering nothing.
Peeking with BufferedIterator
Many algorithms need to look at the next element before deciding whether to take it: merging sorted inputs, grouping runs, or parsing records that span several lines. it.buffered returns a BufferedIterator with a head method that returns the next element without consuming it. It buffers at most one element, so it is cheap. Two examples follow, both written against AbstractIterator, the base class to extend when implementing iterators:
import scala.collection.AbstractIterator
// Merge two sorted iterators lazily: head lets us look without consuming.
def mergeSorted(a: Iterator[Int], b: Iterator[Int]): Iterator[Int] =
new AbstractIterator[Int] {
private val x = a.buffered
private val y = b.buffered
def hasNext: Boolean = x.hasNext || y.hasNext
def next(): Int =
if (!y.hasNext || (x.hasNext && x.head <= y.head)) x.next() else y.next()
}
// Group consecutive elements that share a key (input must already be ordered by key).
def runs[A, K](it: Iterator[A])(key: A => K): Iterator[(K, Vector[A])] =
new AbstractIterator[(K, Vector[A])] {
private val b = it.buffered
def hasNext: Boolean = b.hasNext
def next(): (K, Vector[A]) = {
val k = key(b.head) // throws NoSuchElementException when empty: correct
val acc = Vector.newBuilder[A]
while (b.hasNext && key(b.head) == k) acc += b.next()
(k, acc.result())
}
}runs is the building block for sessionising logs sorted by user, collapsing duplicate keys from a sorted file, or turning a stream of rows into groups without loading the file. Only one group is held in memory at a time. On a plain iterator, nextOption() safely takes one element, but it consumes it. Peeking always needs buffered.
Building iterators: unfold, generators and AbstractIterator
Paged APIs are a common source of iterator code. Each call returns some items and a token for the next page. There are two clean ways to turn that into a flat stream of items:
final case class Page[A](items: Seq[A], next: Option[String])
// 1) Iterator.unfold: state in, (element, next state) out, None to stop.
def pages[A](fetch: Option[String] => Page[A]): Iterator[A] =
Iterator.unfold[Seq[A], Option[Option[String]]](Some(None)) {
case None => None // no more pages
case Some(token) =>
val p = fetch(token)
Some((p.items, p.next.map(Some(_)))) // stop after the page with no next token
}.flatten
// 2) The same thing by hand, when you need explicit control (retries, metrics, close).
final class PageIterator[A](fetch: Option[String] => Page[A]) extends AbstractIterator[A] {
private var buf: Iterator[A] = Iterator.empty
private var token: Option[String] = None
private var done = false
def hasNext: Boolean = {
while (!buf.hasNext && !done) { // loop: a page may be empty
val p = fetch(token)
buf = p.items.iterator
token = p.next
done = token.isEmpty
}
buf.hasNext
}
def next(): A = if (hasNext) buf.next() else Iterator.empty.next()
}
// 3) Small generators.
Iterator.from(1) // 1, 2, 3, ... forever
Iterator.iterate(1L)(_ * 2) // 1, 2, 4, 8, ...
Iterator.continually(queue.poll()).takeWhile(_ != null)Iterator.unfold is the declarative version. The state says which token to fetch next, and None ends the stream. The hand-written class is useful when you need to retry, record metrics or close a connection. The while loop in hasNext matters: a page can be empty while later pages are not. Make hasNext safe to call twice without fetching twice, because library methods do call it repeatedly. Iterators like these are also the usual bridge into effectful streaming libraries; fs2 can lift one into a stream with proper resource handling.
Resources: getLines and Using
Source.fromFile(path).getLines() is the textbook example of an iterator over a resource. The file handle stays open until someone closes the Source, and the iterator has no way to do that. scala.util.Using closes the resource when its block exits. That creates a trap, because the block can exit before the lazy iterator has been read:
import scala.io.Source
import scala.util.{Try, Using}
// WRONG: the iterator escapes the block, and the file is closed before anyone reads it.
def errorLinesBad(path: String): Iterator[String] =
Using.resource(Source.fromFile(path))(_.getLines().filter(_.contains("ERROR")))
// RIGHT: finish all consumption inside the block; return a materialised result.
def statusCounts(path: String): Try[Map[Int, Int]] =
Using(Source.fromFile(path, "UTF-8")) { src =>
src.getLines()
.flatMap(parseLine) // Option[LogLine] -> drops bad lines
.map(_.status)
.foldLeft(Map.empty[Int, Int])((m, s) => m.updated(s, m.getOrElse(s, 0) + 1))
}The bad version compiles and returns an iterator. The first next() then fails with a stream-closed error, or reads a partial buffer, depending on timing. The rule is that consumption must finish inside the Using block, and only a materialised result such as a Map, a Vector or a count may leave it. Using(...) returns a Try, so I/O errors become values instead of exceptions thrown from deep inside a pipeline.
Java interop and performance notes
With import scala.jdk.CollectionConverters._, javaIt.asScala wraps a java.util.Iterator without copying, and asJava goes the other way. JDBC result sets and directory streams fit this model, with the same resource rule as files.
On performance, iterators avoid building intermediate collections. A map and filter chain over an iterator allocates one small iterator object per stage, not one collection per stage. Elements of primitive type are still boxed, because Iterator[Int] is generic. For tight numeric loops over arrays, a while loop or an array operation is still faster. knownSize returns -1 for most iterators, so builders cannot pre-size, and toVector on a known-length source is faster if you convert the original collection rather than its iterator.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Second computation returns 0 or empty | Iterator consumed by size, sum or a debug print | Materialise once, or fold everything in one pass |
| Elements skipped or repeated | Original used after take, drop or map | Keep only the returned iterator |
| OutOfMemoryError in a streaming job | duplicate, span or partition with diverging consumers | Single fold, or consume the prefix first |
| Stream closed error | Iterator escaped a Using block | Finish consumption inside the block |
| File handle leak | getLines without closing the Source | Using, or a resource-safe streaming library |
| Stream ends early | hasNext treats an empty page as the end | Loop until a non-empty page or the final token |
Worked example: error report from a large access log
Task: from a multi-gigabyte access log, report the count of each 5xx status and the first 20 error lines, in one pass and in constant memory. Combine the pieces above: Using for the file, flatMap(parseLine) to drop malformed lines, and a single foldLeft whose accumulator holds a Map[Int, Int] of counts and a Vector of samples that stops growing at 20. Do not use duplicate to compute the counts and samples separately, and do not call size for a progress log. Return the accumulator from inside the Using block. The result is a Try holding a small value, and the file is closed whether parsing succeeds or fails.
What to do next
- Search your code for iterators passed to logging or
sizecalls before they are processed. - Find every
duplicate,spanandpartitionon an iterator and check which side is consumed first. - Wrap every
Source.fromFileand similar resource inUsing, and make sure nothing lazy escapes the block. - Rewrite one multi-pass aggregation as a single
foldLeftand compare peak memory. - Turn a paged API client into an iterator with
Iterator.unfold, and test it against an empty middle page. - Review Scala collections to choose the right strict type when you do need to materialise.