An Iterator is the smallest abstraction in the Scala collections library. It is a cursor with two methods, hasNext and next(), that hands out elements one at a time and never goes back. That makes it the right tool for data that is too big to hold, arrives over time, or comes from a resource such as a file or a paged API. It also makes it the easiest collection type to misuse. Most iterator bugs produce no exception, only a wrong number.

This page explains what an iterator promises and what it does not, which operations are lazy and which consume, where hidden buffers appear, and how to build and close iterators safely. It ends with a worked example and a checklist. Memoised lazy sequences are covered separately in Scala LazyList, in depth, and the general machinery of laziness in Scala lazy evaluation. Everything here applies to the Scala 2.13 and Scala 3 standard library, which share one collections implementation.

Advertisement

What an Iterator is, and why it exists

An iterator holds a position in some source. hasNext reports whether another element is available, and next() returns it and advances. Calling next() when hasNext is false throws NoSuchElementException. Unlike a List or Vector, an iterator is not a container. It does not remember elements it has returned, so it can only be traversed once. Iterator extends IterableOnce, the type that methods use when they need to traverse input exactly once, and that is why toList, ++ and many builders accept iterators directly.

The payoff is memory. Processing a 40 GB log with Source.getLines() uses memory proportional to one line, not to the file, as long as nothing in the pipeline holds on to elements. The cost is that you, not the type system, have to make sure each element is read once and in order.

The one rule: use only the iterator a method returns

An Iterator pipeline pulls one element at a time; nothing runs until a consumer asksSourcegetLines(), paging APImap(parse)lazy, 0 bufferfilter(isErr)lazy, 0 buffertake(100)lazyforeach / toListconsumernext()Demand flows right to left (next), elements flow left to right; each element is read onceit.duplicatetwo iterators over one sourcea queue holds everything the faster copy has seenand the slower one has not: unbounded if they divergeit.span(p) / it.partition(p)two results from one passconsume the second result first and the first oneis held in memory while you doThe contractafter calling any method other than next or hasNext, use only the iterator it RETURNEDBreaking the contract does not throw; it silently gives wrong answers
Top: an iterator pipeline pulls elements on demand. Middle: the two operations that buffer. Bottom: the rule that prevents most bugs.

The Scala documentation states the contract plainly. After calling a method on an iterator, you should discard that iterator and use only the one the method returned, unless the method was next or hasNext. Operations such as map, filter, take and drop return a new iterator that shares the underlying source. Advancing either one disturbs the other, and the library makes no promise about the state of the original.

val it = Iterator(1, 2, 3, 4, 5)

val n = it.size          // consumes all five elements to count them
val total = it.sum       // 0: the iterator is already exhausted

val it2 = Iterator(1, 2, 3, 4, 5)
val firstThree = it2.take(3)  // after take, 'it2' itself is in an UNSPECIFIED state
it2.next()                    // compiles, but the result is not guaranteed: do not do this

// Correct: use the returned iterator, or materialise once when you need two passes.
val xs = Iterator(1, 2, 3, 4, 5).toVector
val (count, sum) = (xs.size, xs.sum)

The first bug is the most common in real code: a method logs an iterator's size for debugging, and the rest of the method sees empty input. Anything that traverses the whole input, such as size, sum, toList or foldLeft, exhausts the iterator. For two passes, materialise once into a Vector, or restructure the work into a single fold.

Advertisement

Lazy, consuming and partially consuming operations

OperationBehaviourBuffers
map, filter, collect, flatMap, zip, zipWithIndexLazy: returns a new iterator, does no work yetNo
take, drop, takeWhile, sliceLazy; drop(n) skips n elements when first pulledNo
grouped(n), sliding(n)Lazy, but each group is materialised as a SeqOne group
bufferedLazy; adds head for peekingOne element
find, exists, contains, indexWhereConsumes up to and including the first matchNo
size, sum, foldLeft, toList, foreach, mkStringConsumes everythingResult only
duplicate, span, partitionTwo iterators sharing one sourceUp to the whole input

Laziness has a consequence that surprises people coming from List: side effects in map run when elements are pulled, not when map is called. it.map(x => { println(x); x }) prints nothing until a consumer runs, and if the consumer only takes three elements, only three lines are printed. That is usually what you want. It does mean you cannot rely on a mapping step having run before a later statement that does not touch the iterator.

Iterator, LazyList and views compared

IteratorLazyListView
TraversalsOneManyMany, recomputed each time
Remembers elementsNoYes, memoised once evaluatedNo
Memory for a huge streamConstantGrows if you keep a reference to the headConstant per pass
Good forFiles, sockets, paged APIs, one-shot pipelinesRecursive or self-referential sequencesFusing transformations over an existing collection

The practical rule: if the data comes from outside the program and is read once, use an iterator. If you need to revisit elements, use a strict collection, or a LazyList when the sequence is defined in terms of itself. Allocation and boxing costs across these choices are compared in Scala collections performance.

Hidden buffers: duplicate, span and partition

Sometimes you want two results from one pass. The library offers duplicate, which returns two independent iterators over the same elements, and span and partition, which split elements into two iterators. All three keep a single-pass source behind two consumers, so whatever one consumer has read and the other has not must be stored. In 2.13, partition is built on duplicate with a filter on each side, so it has the same buffering profile.

// Memory trap 1: duplicate, then let the copies drift apart.
val (a, b) = lines.duplicate
val errors   = a.count(_.contains("ERROR"))   // drains 'a' fully...
val warnings = b.count(_.contains("WARN"))    // ...so every line was queued for 'b'

// Same answer, one pass, constant memory:
val (errs, warns) = lines.foldLeft((0, 0)) { case ((e, w), l) =>
  (e + (if (l.contains("ERROR")) 1 else 0), w + (if (l.contains("WARN")) 1 else 0))
}

// Memory trap 2: span. Reading 'body' first forces 'header' to be buffered.
val (header, body) = lines.span(_.startsWith("#"))
val headerLines = header.toVector     // consume the prefix FIRST: no buffering needed
body.foreach(process)

If you drain the first copy completely before touching the second, the queue grows to the size of the input, and a streaming pipeline quietly becomes a full in-memory copy. For counting or aggregating, a single foldLeft with a tuple or case class accumulator does the same work in constant memory. For span, consume the prefix first. In the usual header-then-body case, that means buffering nothing.

Peeking with BufferedIterator

Many algorithms need to look at the next element before deciding whether to take it: merging sorted inputs, grouping runs, or parsing records that span several lines. it.buffered returns a BufferedIterator with a head method that returns the next element without consuming it. It buffers at most one element, so it is cheap. Two examples follow, both written against AbstractIterator, the base class to extend when implementing iterators:

import scala.collection.AbstractIterator

// Merge two sorted iterators lazily: head lets us look without consuming.
def mergeSorted(a: Iterator[Int], b: Iterator[Int]): Iterator[Int] =
  new AbstractIterator[Int] {
    private val x = a.buffered
    private val y = b.buffered
    def hasNext: Boolean = x.hasNext || y.hasNext
    def next(): Int =
      if (!y.hasNext || (x.hasNext && x.head <= y.head)) x.next() else y.next()
  }

// Group consecutive elements that share a key (input must already be ordered by key).
def runs[A, K](it: Iterator[A])(key: A => K): Iterator[(K, Vector[A])] =
  new AbstractIterator[(K, Vector[A])] {
    private val b = it.buffered
    def hasNext: Boolean = b.hasNext
    def next(): (K, Vector[A]) = {
      val k = key(b.head)                       // throws NoSuchElementException when empty: correct
      val acc = Vector.newBuilder[A]
      while (b.hasNext && key(b.head) == k) acc += b.next()
      (k, acc.result())
    }
  }

runs is the building block for sessionising logs sorted by user, collapsing duplicate keys from a sorted file, or turning a stream of rows into groups without loading the file. Only one group is held in memory at a time. On a plain iterator, nextOption() safely takes one element, but it consumes it. Peeking always needs buffered.

Building iterators: unfold, generators and AbstractIterator

Paged APIs are a common source of iterator code. Each call returns some items and a token for the next page. There are two clean ways to turn that into a flat stream of items:

final case class Page[A](items: Seq[A], next: Option[String])

// 1) Iterator.unfold: state in, (element, next state) out, None to stop.
def pages[A](fetch: Option[String] => Page[A]): Iterator[A] =
  Iterator.unfold[Seq[A], Option[Option[String]]](Some(None)) {
    case None        => None                                 // no more pages
    case Some(token) =>
      val p = fetch(token)
      Some((p.items, p.next.map(Some(_))))                   // stop after the page with no next token
  }.flatten

// 2) The same thing by hand, when you need explicit control (retries, metrics, close).
final class PageIterator[A](fetch: Option[String] => Page[A]) extends AbstractIterator[A] {
  private var buf: Iterator[A]        = Iterator.empty
  private var token: Option[String]   = None
  private var done                    = false
  def hasNext: Boolean = {
    while (!buf.hasNext && !done) {          // loop: a page may be empty
      val p = fetch(token)
      buf = p.items.iterator
      token = p.next
      done = token.isEmpty
    }
    buf.hasNext
  }
  def next(): A = if (hasNext) buf.next() else Iterator.empty.next()
}

// 3) Small generators.
Iterator.from(1)                       // 1, 2, 3, ... forever
Iterator.iterate(1L)(_ * 2)            // 1, 2, 4, 8, ...
Iterator.continually(queue.poll()).takeWhile(_ != null)

Iterator.unfold is the declarative version. The state says which token to fetch next, and None ends the stream. The hand-written class is useful when you need to retry, record metrics or close a connection. The while loop in hasNext matters: a page can be empty while later pages are not. Make hasNext safe to call twice without fetching twice, because library methods do call it repeatedly. Iterators like these are also the usual bridge into effectful streaming libraries; fs2 can lift one into a stream with proper resource handling.

Resources: getLines and Using

Source.fromFile(path).getLines() is the textbook example of an iterator over a resource. The file handle stays open until someone closes the Source, and the iterator has no way to do that. scala.util.Using closes the resource when its block exits. That creates a trap, because the block can exit before the lazy iterator has been read:

import scala.io.Source
import scala.util.{Try, Using}

// WRONG: the iterator escapes the block, and the file is closed before anyone reads it.
def errorLinesBad(path: String): Iterator[String] =
  Using.resource(Source.fromFile(path))(_.getLines().filter(_.contains("ERROR")))

// RIGHT: finish all consumption inside the block; return a materialised result.
def statusCounts(path: String): Try[Map[Int, Int]] =
  Using(Source.fromFile(path, "UTF-8")) { src =>
    src.getLines()
      .flatMap(parseLine)                          // Option[LogLine] -> drops bad lines
      .map(_.status)
      .foldLeft(Map.empty[Int, Int])((m, s) => m.updated(s, m.getOrElse(s, 0) + 1))
  }

The bad version compiles and returns an iterator. The first next() then fails with a stream-closed error, or reads a partial buffer, depending on timing. The rule is that consumption must finish inside the Using block, and only a materialised result such as a Map, a Vector or a count may leave it. Using(...) returns a Try, so I/O errors become values instead of exceptions thrown from deep inside a pipeline.

Java interop and performance notes

With import scala.jdk.CollectionConverters._, javaIt.asScala wraps a java.util.Iterator without copying, and asJava goes the other way. JDBC result sets and directory streams fit this model, with the same resource rule as files.

On performance, iterators avoid building intermediate collections. A map and filter chain over an iterator allocates one small iterator object per stage, not one collection per stage. Elements of primitive type are still boxed, because Iterator[Int] is generic. For tight numeric loops over arrays, a while loop or an array operation is still faster. knownSize returns -1 for most iterators, so builders cannot pre-size, and toVector on a known-length source is faster if you convert the original collection rather than its iterator.

Failure modes

SymptomCauseFix
Second computation returns 0 or emptyIterator consumed by size, sum or a debug printMaterialise once, or fold everything in one pass
Elements skipped or repeatedOriginal used after take, drop or mapKeep only the returned iterator
OutOfMemoryError in a streaming jobduplicate, span or partition with diverging consumersSingle fold, or consume the prefix first
Stream closed errorIterator escaped a Using blockFinish consumption inside the block
File handle leakgetLines without closing the SourceUsing, or a resource-safe streaming library
Stream ends earlyhasNext treats an empty page as the endLoop until a non-empty page or the final token

Worked example: error report from a large access log

Task: from a multi-gigabyte access log, report the count of each 5xx status and the first 20 error lines, in one pass and in constant memory. Combine the pieces above: Using for the file, flatMap(parseLine) to drop malformed lines, and a single foldLeft whose accumulator holds a Map[Int, Int] of counts and a Vector of samples that stops growing at 20. Do not use duplicate to compute the counts and samples separately, and do not call size for a progress log. Return the accumulator from inside the Using block. The result is a Try holding a small value, and the file is closed whether parsing succeeds or fails.

What to do next

  1. Search your code for iterators passed to logging or size calls before they are processed.
  2. Find every duplicate, span and partition on an iterator and check which side is consumed first.
  3. Wrap every Source.fromFile and similar resource in Using, and make sure nothing lazy escapes the block.
  4. Rewrite one multi-pass aggregation as a single foldLeft and compare peak memory.
  5. Turn a paged API client into an iterator with Iterator.unfold, and test it against an empty middle page.
  6. Review Scala collections to choose the right strict type when you do need to materialise.
Key takeaway: A Scala Iterator is a single-pass cursor: after calling anything other than next or hasNext, use only the iterator that call returned. Transformations are lazy and buffer nothing, aggregations consume everything, and duplicate, span and partition can buffer the whole input. Peek with buffered, build with unfold or AbstractIterator, and finish consuming any resource-backed iterator inside a Using block.