Scala gives every collection two ways to cut a sequence into pieces. grouped(n) partitions it into consecutive chunks that never overlap. sliding(n, step) passes a window of n elements along the sequence, moving step elements each time, so neighbouring windows can share elements. Both look trivial, and both have edge cases that produce wrong answers quietly: a short last window, a skipped tail, a single window for input that is too short, and a copy of every window that turns a linear job quadratic.
This page pins down the exact behaviour, shows how to control the edges, and then uses the two methods for real work: smoothing a time series, flagging runs, and building windows to train a forecasting model without leaking the test set. Every result quoted below was produced by calling scala-library 2.13.16 directly. Scala 3 uses the same collections library, so the behaviour is identical there. For the wider partition and groupBy family, see partition, span and groupBy; this page is only about chunks and windows.
Chunks and windows in one picture
The exact rules, checked
The rules fit in four sentences. An empty collection yields no windows and no groups. A non-empty collection shorter than the window yields exactly one window, the whole collection. With a step, the last window may be shorter than the size, but only if the window before it did not already reach the end and the step did not jump over the leftover elements. A size or step below 1 throws immediately.
The table shows the actual output for each case. The first two rows are the scaladoc's own examples; the rest were run against 2.13.16.
| Call | Result |
|---|---|
List(1,2,3,4,5).sliding(2, 2) | List(1, 2), List(3, 4), List(5) |
List(1,2,3,4,5,6).sliding(2, 3) | List(1, 2), List(4, 5): 3 and 6 are skipped |
List(1,2,3,4,5).sliding(3, 2) | List(1, 2, 3), List(3, 4, 5): no tail, the window reached the end |
List(1,2,3,4,5,6).sliding(3, 2) | List(1, 2, 3), List(3, 4, 5), List(5, 6) |
(1 to 7).toList.grouped(3) | List(1, 2, 3), List(4, 5, 6), List(7) |
List(1).sliding(2) | List(List(1)): one short window, not zero |
List().sliding(2) | empty |
List(1,2,3,4,5).sliding(4, 10) | List(1, 2, 3, 4): element 5 is dropped |
List(1, 2).grouped(0) | IllegalArgumentException: requirement failed: size=0 and step=0, but both must be positive |
Two rows deserve attention. sliding(4, 10) silently drops the fifth element, because the next window would start beyond the end. When step is larger than size, sliding is a sampler, not a partition, and data between windows is never seen. And List(1).sliding(2) yields one window of length one, so any code that destructures a window as exactly two elements will fail on short input. The exception for zero is thrown when you call the method, not when you consume the iterator, even on an empty list.
What you get back: single-pass iterators
On a collection, both methods return an Iterator of the collection's own type: a List produces an Iterator[List[A]], a Vector an iterator of vectors. Nothing is computed until you consume it, and it can be consumed only once. Calling .size to log how many batches you have, and then looping, leaves the loop with nothing. Convert with .toList if you need the windows twice. The general rules for single-pass iterators are covered in Scala iterators.
On an Iterator, the same methods return a GroupedIterator, whose windows are immutable.Seq values (an ArraySeq in 2.13.16). The GroupedIterator adds two switches that the collection methods do not expose, and they are the clean way to handle the tail.
Controlling the tail: withPartial and withPadding
// Drop a short last window instead of returning it.
Iterator(1, 2, 3, 4, 5, 6, 7).grouped(3).withPartial(false).toList
// List(ArraySeq(1, 2, 3), ArraySeq(4, 5, 6))
// Pad the last window to full size. The argument is by-name, evaluated per fill element.
Iterator(1, 2, 3, 4, 5, 6, 7).grouped(3).withPadding(0).toList
// List(ArraySeq(1, 2, 3), ArraySeq(4, 5, 6), ArraySeq(7, 0, 0))
// The short-input case disappears too.
Iterator(1).sliding(2).withPartial(false).toList
// List()
// From a List, go through .iterator to reach the switches.
readings.iterator.sliding(5, 2).withPartial(false)The two switches interact: withPadding turns partial windows back on, and withPartial clears any padding, so the last call wins. Choose by meaning. Padding is right when a consumer requires fixed-size input, such as a fixed-length frame for a model or a protocol. Dropping is right when a short window would be a different kind of object, such as a training example with no label. Keeping the default is right for batching side effects, where the last small batch still has to be written.
What a window costs, and running aggregates
Each window is a new collection. sliding(k) over n elements builds about n windows of k elements, so it allocates and copies roughly n * k element references. For k = 3 that is irrelevant. For a 60-element moving window over ten million readings it is 600 million copies before you compute anything. And a List source pays again, because each window is built as a fresh linked list.
When the computation over a window can be updated incrementally, keep one buffer and update a running value instead of materialising windows:
import scala.collection.mutable
/** Mean of every full window of size k, in one pass, O(1) work per element. */
def movingMean(xs: Iterator[Double], k: Int): Iterator[Double] = {
require(k > 0, s"k=$k must be positive")
val window = mutable.ArrayDeque.empty[Double]
var sum = 0.0
var seen = 0L
xs.flatMap { x =>
window.append(x); sum += x; seen += 1
if (window.size > k) sum -= window.removeHead()
// Floating-point subtraction drifts; recompute exactly now and then.
if (seen % 100000 == 0) sum = window.sum
if (window.size == k) Some(sum / k) else None
}
}This behaves like sliding(k).withPartial(false): input shorter than k produces nothing, where plain sliding(k) would produce one short window. That difference is deliberate, and worth a test. Sums, counts and means update in constant time. A moving minimum or maximum needs a monotonic deque, and a moving median needs two heaps or an order-statistics structure; none of them need the window copied. Keep sliding for small windows and for logic that genuinely looks at the whole window, such as pattern matching on its shape.
Worked example: forecasting windows and an alert
A team trains a model to predict the next hourly meter reading from the previous four. The raw series has ten readings, indexed 0 to 9 here so the windows are easy to read. Each training example needs five consecutive readings: four inputs and one label. Using a step of 2 halves the overlap between examples.
final case class Example(inputs: Seq[Double], label: Double)
def examples(series: Seq[Double], lookback: Int, step: Int): List[Example] =
series.iterator
.sliding(lookback + 1, step)
.withPartial(false) // a short tail has no label
.map(w => Example(w.init, w.last))
.toListThe first draft used series.sliding(5, 2). On indices 0 to 9 that yields (0..4), (2..6), (4..8) and then a fourth window (6, 7, 8, 9), because the window ending at 8 did not reach index 9 and the step of 2 does not skip it. The draft treated the last element of every window as the label, so the partial window became an example whose label was reading 9 and whose inputs were only three readings. With withPartial(false) the output is exactly the three full windows, which is what 2.13.16 returns.
The second bug was leakage. The draft windowed the whole series and then shuffled the examples into training and test sets. With a step smaller than the window, neighbouring examples share four of five readings, so test labels appear as training inputs and the test score is inflated. The fix is ordering: split the series by time first, then window each side separately, and optionally leave a gap of at least the lookback between them, so the first test inputs are not the readings immediately after, and highly correlated with, the last training labels.
val cut = (series.length * 0.8).toInt
val train = examples(series.take(cut), lookback = 4, step = 2)
val test = examples(series.drop(cut + 4), lookback = 4, step = 1) // 4-step gapThe same series also feeds an alert. Hourly consumption is the difference between consecutive readings, and an alert fires on three consecutive hours above a threshold. Both are window shapes, and matching on the shape keeps the code honest about short input:
val usage = readings.sliding(2).collect { case Seq(a, b) => b - a }.toList
val alerts = usage.zipWithIndex.sliding(3).collect {
case Seq((u1, i), (u2, _), (u3, _)) if u1 > limit && u2 > limit && u3 > limit => i
}.toListcollect with a fixed-arity pattern ignores the one short window that a too-short series produces. Writing map { case Seq(a, b) => ... } instead throws a MatchError on a single reading.
Large and unbounded inputs
Because the methods return iterators, they work on inputs that never fit in memory. Source.fromFile(path).getLines().grouped(1000) reads a file in batches of lines, and LazyList.iterate(0)(_ + 1).sliding(3).take(2).toList is fine on an infinite source because only the windows you take are built. Two limits apply. A single window is always fully materialised, so a window of a million elements is a million-element buffer. And the source iterator is consumed, so you cannot reuse it after windowing; see LazyList and views for the lazy alternatives and their own traps.
Outside the standard library the same two ideas appear under the same names. Akka Streams has grouped(n), sliding(n, step), and groupedWithin(n, duration), which emits a batch when it is full or when time runs out; for an unbounded stream that time bound is usually what you want, because a count-only batch can wait forever for its last element. In Spark, windows over rows are expressed with window functions such as lag or an aggregate over Window.orderBy(...).rowsBetween(-2, 0), which are distributed across partitions; driver-side sliding on collected data does not scale past what the driver can hold.
Failure modes
- Short input.
xs.sliding(2)on one element yields one window of one element. Pattern-match withcollect, or usewithPartial(false). - Unwanted tail. With a step, a short last window can appear or not depending on the length modulo the step, so a test with one convenient length passes and production fails. Test lengths that land on, before and after a boundary.
- Sampling by accident.
step > sizeskips data. It is correct for downsampling and wrong everywhere else. - Consumed twice. Counting batches with
.sizeor logging.toListdrains the iterator. - Quadratic copying. Large windows copy
n * kreferences. Use a running aggregate. - Leakage. Overlapping windows cut across a train and test split. Split first, then window.
- Zero sizes from configuration. A batch size read as 0 from a missing setting throws
IllegalArgumentExceptionwhen the method is called. Validate configuration at startup.
Trade-offs and a decision table
| Need | Use | Why |
|---|---|---|
| Batch side effects (inserts, API pages) | grouped(n) | Every element once; keep the short last batch |
| Fixed-size frames for a consumer | iterator.grouped(n).withPadding(x) | Every frame has the same size |
| Pairwise deltas, small pattern windows | sliding(2) / sliding(3) with collect | Clear code; copying is negligible |
| Training examples | iterator.sliding(k + 1, step).withPartial(false) | No example without a label |
| Moving sum, mean, min or max on large k | Running aggregate over a deque | O(1) per element, no window copies |
| Unbounded stream batches | groupedWithin (Akka Streams) | Bounded latency as well as size |
The deeper trade-off is clarity against cost. sliding states intent in one word and is easy to review; a running aggregate is faster but has more state to get wrong. Start with sliding, measure, and switch where the window size or data volume makes the copies visible. For folding logic over the windows themselves, the patterns in fold, map and filter apply unchanged.
What to do next
- Search your code for
.sliding(and.grouped(. For each call, write down what should happen with input shorter than the window and with a short tail, and add a test for both. - Replace
map { case Seq(a, b) => ... }over windows withcollect, or switch toiterator.sliding(n).withPartial(false). - Flag every
sliding(size, step)withstep > sizeand confirm the skipped data is intended. - Find iterators that are consumed twice, such as a size check followed by a loop, and materialise them once.
- Profile any
slidingwith a window above a few dozen elements on large inputs; replace it with a running aggregate if allocation shows up. - For forecasting datasets, move the train and test split before windowing and add a gap of at least the lookback.
- Validate batch and window sizes from configuration at startup so a zero fails the deploy, not the first request.