Calling .parallel() on a stream is the cheapest way to use more than one core in Java: one method call, no threads, no locks. It is also one of the easiest ways to make code slower or wrong. Whether a parallel stream helps depends on how its source splits, how much work each element needs, whether the operations respect a few algebraic rules, and what else is running in the same thread pool. None of that is visible in the code.
This page builds the model you need to decide. It starts with what a parallel pipeline actually does, then covers the pool it runs on, which sources split well, the rules that keep results correct (with a demo whose output is shown as it ran), ordering and collectors, when parallelism pays and how to measure it, and a worked example. The API itself is covered in the Java Streams API guide and the internals of stages and sinks in stream operations.
What .parallel() actually does
A stream pipeline does nothing until its terminal operation runs. .parallel() and parallelStream() only set a flag on the pipeline, and the last call to parallel() or sequential() wins for the whole pipeline: you cannot run half a pipeline in parallel.
When the terminal operation runs in parallel mode, the library takes the source's Spliterator and splits it recursively with trySplit(), each call handing off roughly half of the remaining elements. Splitting stops when pieces are small enough; OpenJDK's implementation aims for about four leaf tasks per unit of pool parallelism, but that is an implementation detail. Each leaf then runs the whole pipeline sequentially over its own elements, producing a partial result: a partial sum, a partial list, a partial map. Finally the partial results are combined pairwise back up the tree.
That structure explains almost everything else on this page. Speed comes from leaves running at the same time, so it depends on whether the source splits evenly and whether each leaf has enough work. Correctness depends on the combine step producing the same answer as a single sequential pass, which is only true if the operations obey the rules described below.
Where the threads come from
Parallel streams run on ForkJoinPool.commonPool(), a single pool shared by the whole JVM. Its default parallelism is the number of available processors minus one, because the thread that called the terminal operation also takes part in the work. You can change it at startup with the system property java.util.concurrent.ForkJoinPool.common.parallelism; the JVM must see the correct processor count, which modern JDKs derive from container CPU limits.
Every parallel stream and every CompletableFuture async stage without an explicit executor competes for the same workers, so a stream that blocks on I/O inside map stalls every other user of the pool. Parallel streams are for CPU-bound work; for many concurrent blocking calls, use virtual threads or a dedicated executor.
You will see code that runs a stream inside customPool.submit(() -> list.parallelStream()...).get() to keep it off the common pool. It works in current OpenJDK because tasks forked from inside a ForkJoinPool worker go to that worker's pool, but this is implementation behaviour, not something the Stream specification promises. Work-stealing and the pool itself are covered in the ForkJoinPool guide.
Sources decide how well a stream splits
A leaf can only start once its piece of the source exists, so the source's spliterator sets the ceiling on speed-up. Sources with known size that split into exact halves in constant time are ideal; sources that must be walked to split are poor.
| Source | How it splits | Parallel suitability |
|---|---|---|
Arrays, ArrayList, IntStream.range | Index arithmetic, exact halves | Good |
HashMap/HashSet views | By bucket ranges, sizes estimated | Usually fine |
TreeMap views | By tree structure, roughly balanced | Fair |
LinkedList | Must walk nodes to split | Poor |
Stream.iterate, Stream.generate | Inherently sequential; split by buffering batches | Poor |
BufferedReader.lines(), most I/O sources | Unknown size; buffered batches | Poor unless per-line work is heavy |
Two spliterator characteristics matter: SIZED and SUBSIZED. When both hold, the library knows exactly how many elements each leaf has, which lets operations like toArray and toList write results straight into the right position. A filter removes exact sizing downstream, which is one reason the same pipeline can be fast or slow depending on its first operation.
The rules that keep results correct
A parallel result equals the sequential result only if the operations obey three rules from the java.util.stream package documentation. The accumulator and combiner must be associative, so (a op b) op c equals a op (b op c). The identity passed to reduce must be a true identity, so identity op x equals x. And the behavioural parameters must be non-interfering and, ideally, stateless: they must not modify the source and should not depend on mutable state that changes during the run.
The demo below breaks the identity rule and the statelessness rule on purpose:
import java.util.*;
import java.util.concurrent.ForkJoinPool;
import java.util.stream.*;
public class ReduceDemo {
public static void main(String[] args) {
System.out.println("cores=" + Runtime.getRuntime().availableProcessors()
+ " commonPool=" + ForkJoinPool.commonPool().getParallelism());
// 10 is NOT an identity for +, so every leaf adds it once.
int seq = IntStream.rangeClosed(1, 100).reduce(10, Integer::sum);
int par = IntStream.rangeClosed(1, 100).parallel().reduce(10, Integer::sum);
System.out.println("reduce(10, sum) sequential=" + seq + " parallel=" + par);
// Shared mutable state: ArrayList is not thread-safe.
List<Integer> unsafe = new ArrayList<>();
try {
IntStream.range(0, 1_000_000).parallel().forEach(unsafe::add);
System.out.println("ArrayList size after parallel forEach=" + unsafe.size());
} catch (RuntimeException e) {
System.out.println("ArrayList parallel forEach threw " + e.getClass().getSimpleName());
}
// Let the stream build the list: each leaf fills its own, then they are joined.
List<Integer> safe = IntStream.range(0, 1_000_000).parallel().boxed().toList();
System.out.println("toList size=" + safe.size() + " ordered="
+ safe.equals(IntStream.range(0, 1_000_000).boxed().toList()));
}
}Running it with java 23 on a 32-core machine printed this, identically on three runs:
cores=32 commonPool=31
reduce(10, sum) sequential=5060 parallel=6050
ArrayList parallel forEach threw ArrayIndexOutOfBoundsException
toList size=1000000 ordered=trueSequentially, reduce(10, Integer::sum) adds 10 once to 5050. In parallel, every leaf starts from the identity, and on this machine the 100-element range was split into 100 single-element leaves, so 10 was added 100 times: 5050 + 1000 = 6050. The number of leaves depends on the pool's parallelism, so on another machine the wrong answer is a different wrong answer, which is exactly what makes this bug hard to reproduce. A non-associative operation such as subtraction fails the same way. If you need an offset, add it after the reduction.
Side effects and shared state
The second part of the demo adds to an ArrayList from forEach. ArrayList is not thread-safe; concurrent adds race on its internal array and size. On this run it threw ArrayIndexOutOfBoundsException; on another run it may silently lose elements, which is worse. Wrapping the list in Collections.synchronizedList makes it correct but serialises every add on one lock, and the list's order is now the order threads happened to run.
The fix is to let the stream build the result. A collector gives each leaf its own container and merges them in the combine step, so there is no shared state at all, and for ordered streams the merge preserves encounter order, as the last line shows. The same reasoning applies to counters (use count() or sum(), not an AtomicLong incremented in forEach) and to caches mutated inside map. If an operation must touch shared state, the stream is no longer a good fit for parallelism.
Ordering costs, and when to give it up
A stream from a List, an array or a range has an encounter order, and parallel execution must respect it for operations whose result depends on it. findFirst, limit, skip, ordered distinct and forEachOrdered all need to know what happened to the left before they can finish, which means buffering or extra coordination. If you do not need order, say so:
// Ordered: must return the first match in encounter order, so other leaves
// keep working until everything to the left is known.
Optional<Order> first = orders.parallelStream()
.filter(o -> o.amount() > 10_000).findFirst();
// Any match will do: the first leaf to find one wins and the rest are cancelled.
Optional<Order> any = orders.parallelStream()
.filter(o -> o.amount() > 10_000).findAny();
// Grouping in parallel: groupingBy merges per-leaf maps; groupingByConcurrent
// has all leaves insert into one ConcurrentMap and gives up encounter order.
Map<String, Long> byRegion = orders.parallelStream()
.unordered()
.collect(Collectors.groupingByConcurrent(Order::region, Collectors.counting()));findAny can return as soon as any leaf finds a match. unordered() lets limit and distinct take whichever elements arrive first. forEach runs in whatever order leaves finish, which is fine for independent output but not for writing a report.
For collecting into maps, groupingBy and toMap build one map per leaf and merge them, and merging large maps can cost more than the parallel work saved. groupingByConcurrent and toConcurrentMap have every leaf insert into one ConcurrentHashMap, which avoids the merge but gives up encounter order inside each group and adds contention when many elements share a key. Measure both. How the concurrent map behaves under contention is covered in the ConcurrentHashMap guide.
When parallelism pays
A commonly quoted rule of thumb says parallelism starts to pay when the number of elements N multiplied by the cost per element Q is large: on the order of ten thousand or more simple operations' worth of work. Treat it as a starting point for measurement, not a threshold. A million additions of primitive ints often gain little, because the loop is limited by memory bandwidth rather than arithmetic. A thousand elements that each run a 50-microsecond computation usually gain a lot.
Several things eat the gain: boxing (Stream<Integer> instead of IntStream), poorly splitting sources, ordered operations near the end of the pipeline, expensive merges, and a common pool that is already busy. On a busy server, per-request parallel streams compete for the same cores, so total throughput can drop while single-request latency looks better.
The only reliable test is measurement with JMH, the OpenJDK microbenchmark harness, which handles warm-up, JIT compilation and dead-code elimination that make hand-written timing loops misleading:
@State(Scope.Benchmark)
@BenchmarkMode(Mode.AverageTime)
@OutputTimeUnit(TimeUnit.MILLISECONDS)
@Warmup(iterations = 5) @Measurement(iterations = 10) @Fork(2)
public class ScoreBench {
@Param({"1000", "100000", "10000000"})
int n;
double[] data;
@Setup
public void setup() {
data = new Random(42).doubles(n).toArray();
}
@Benchmark
public double sequential() {
return Arrays.stream(data).map(ScoreBench::score).sum();
}
@Benchmark
public double parallel() {
return Arrays.stream(data).parallel().map(ScoreBench::score).sum();
}
static double score(double x) { // the per-element work, Q
return Math.log1p(x) * Math.sqrt(x + 1);
}
}Run it at the data sizes you actually have, on hardware like production, with the same container CPU limits. Expect parallel to lose at the smallest size and to win only above some crossover. If production runs many requests concurrently, also run a load test with parallelism switched on and off, because a microbenchmark on an idle machine hides contention.
Worked example: totals per merchant
A nightly job sums 40 million transaction records held in an ArrayList by merchant, excluding flagged ones. The original code loops sequentially and merges into a shared HashMap. Simply adding parallelStream() to that code would corrupt the map. Rewritten as a collector, it has no shared state:
record Txn(String account, String merchant, long cents, boolean flagged) {}
// Before: sequential, with a side effect into a shared HashMap.
Map<String, Long> totals = new HashMap<>();
txns.forEach(t -> { if (!t.flagged()) totals.merge(t.merchant(), t.cents(), Long::sum); });
// After: no shared state, safe to run in parallel, same result.
Map<String, Long> totals2 = txns.parallelStream() // txns is an ArrayList
.filter(t -> !t.flagged())
.collect(Collectors.toMap(Txn::merchant, Txn::cents, Long::sum));
// Riskier variant: one concurrent map, no ordering, less merging.
ConcurrentMap<String, Long> totals3 = txns.parallelStream()
.unordered()
.filter(t -> !t.flagged())
.collect(Collectors.toConcurrentMap(Txn::merchant, Txn::cents, Long::sum));The source is an ArrayList, so it splits evenly. The work per element is small, a filter and a hash lookup, so the gain depends on the number of merchants: with a few thousand merchants, per-leaf maps are small and merging is cheap; with millions of distinct merchants, merging large maps dominates and toConcurrentMap may win. Long::sum is associative and the map values start from the first element rather than a fake identity, so the result is exact. The team benchmarks both with JMH on production-like data and keeps a test comparing parallel and sequential results.
Failure modes
- Wrong totals from a fake identity or a non-associative operation, different on each machine.
- Lost elements or exceptions from adding to a non-thread-safe collection in forEach.
- Blocking I/O inside a parallel stream starving every other user of the common pool.
- No speed-up because the source is a LinkedList, an iterator or an I/O stream.
- Slower results from findFirst, limit or sorted on large parallel streams where order was not needed.
- Merging large per-leaf maps costing more than the parallel work saved.
- Per-request parallel streams lowering a server's total throughput under load.
- Container CPU limits or a pool parallelism setting different in production from the benchmark machine.
What to do next
- Search your code for parallel() and parallelStream() and list each use with its source type and data size.
- For every reduce and collect, confirm the identity is a true identity and the operations are associative.
- Remove side effects from parallel pipelines; replace shared collections and counters with collectors and terminal operations.
- Remove blocking calls from parallel streams, or move that work to virtual threads or a dedicated executor.
- Use findAny and unordered() wherever encounter order does not matter.
- Benchmark each parallel stream against its sequential version with JMH at production sizes, and delete the ones that do not win.
- Load-test services that use parallel streams per request, with the production container CPU limit.
- Add a test that compares parallel and sequential results on representative data.