Take a heap dump of almost any long-running Java service and sort by retained size. Strings and their backing arrays are usually at or near the top, often a quarter of the live heap or more. Many of those strings are the same value repeated: country codes, status names, HTTP header names, JSON field values, database column values, tenant IDs. A service that parses ten million records holding "US-EAST" keeps ten million separate copies of the same seven bytes, each in its own array with its own header.

String deduplication is a JVM feature that lets the garbage collector find these duplicates and make them share one backing array. It needs no code change: you turn it on with one flag. This article explains exactly what the feature shares and what it cannot, how candidates are chosen, what the dedup thread does, how to estimate the saving before you enable it, how to prove it afterwards, and when interning at the source is the better tool.

Advertisement

What a String costs, and what can be shared

Since JDK 9 (compact strings, JEP 254) a java.lang.String is a small object holding a reference to a byte[] called value, a coder byte that says whether the bytes are Latin-1 (one byte per character) or UTF-16 (two), and a cached hash. On a 64-bit JVM with compressed references the String object is typically 24 bytes. The array costs a 16-byte header plus the characters, rounded up to a multiple of 8. For a seven-character Latin-1 value that is 24 bytes for the String and 24 for the array.

Strings are immutable, so two Strings with equal contents can safely point at the same array. That is the whole idea. Deduplication changes the value field of a duplicate String so that it points at an existing equal array, and the duplicate array becomes garbage. It does not merge the String objects. a == b is still false after deduplication, a.equals(b) is still true, identity hash codes and locks are unaffected, and each String object still costs its own 24 bytes. At best, the feature removes the array half of the cost for each duplicate.

Compare that with String.intern(), which returns one canonical String object from a JVM-wide table. Interning removes both halves and makes == work, but it has to be called explicitly in your code, and every call does a hash table lookup on the calling thread.

How deduplication runs

The feature has two parts: candidate selection, done by the collector while it copies objects, and the deduplication itself, done later by a dedicated background thread.

String deduplication: the GC finds candidates, a concurrent thread shares their arraysGC evacuation pauseG1: String reaches age thresholdor is promoted earlyto an old regionenqueueCandidate queuereferences to String objectsdrainDedup threadconcurrent with the applicationhash array, look upDedup tableweak refs to value arrays, by content hashBeforeString Avalue -> [1]String Bvalue -> [2]byte[] #1 US-EASTbyte[] #2 US-EASTAfterString Avalue -> [1]String Bvalue -> [1]byte[] #1 US-EASTarray #2 is now garbage; A and B are still two objects (A != B)match: repoint valueno match: insert this array
Candidate selection happens during GC copying; the work of hashing, looking up and repointing happens on a concurrent thread. Only the array is shared, so the two String objects remain distinct.

Candidate selection (G1). Deduplication was introduced for G1 by JEP 192 in JDK 8u20, and G1 remains the most common way it is used. During an evacuation pause, when G1 copies a live String it checks two conditions: the object is being copied into a survivor region and its age has just reached StringDeduplicationAgeThreshold (default 3), or it is being promoted into an old region before reaching that age. Either way, the String has proved it is not short-lived, and a reference to it is added to a queue. Strings that die young are never considered, which is deliberate: most strings are temporary, and hashing them would waste CPU on objects about to disappear anyway.

The dedup thread. A background thread drains the queue while the application runs. For each candidate it hashes the contents of its array and looks the array up in the deduplication table. If an equal array is already there, it sets the candidate's value field to that array. If not, it inserts the candidate's array. The table holds weak references, so an array that nothing else uses can still be collected; since JDK 17 the table is built on the JVM's OopStorage weak-reference mechanism (JDK-8254598), and dead entries are removed concurrently rather than during pauses.

Collector support. G1 has supported it since JDK 8u20 and Shenandoah has its own implementation. JDK 18 added support to the Serial, Parallel and Z collectors (JDK-8272609, JDK-8267185, JDK-8267186), so on current JDKs the same flag works on all of them. How each collector chooses candidates differs in detail; the G1 rule above is the one to reason with, and on the others you should rely on the statistics rather than assumptions. For the collectors themselves see G1 in depth, ZGC and Shenandoah.

Advertisement

Turning it on and reading the statistics

The flag is -XX:+UseStringDeduplication. It is off by default. Pair it with unified logging on the stringdedup tag so you can see what it did; the old -XX:+PrintStringDeduplicationStatistics flag from JDK 8 no longer exists. The demo below creates five million strings with only four distinct values, the way a record parser does, then generates churn so they age and get promoted.

// DedupDemo.java: build many equal strings the way a parser does
import java.util.ArrayList;
import java.util.List;

public class DedupDemo {
    public static void main(String[] args) throws Exception {
        String[] regions = {"US-EAST", "US-WEST", "EU-CENTRAL", "AP-SOUTH"};
        List<String> keep = new ArrayList<>();
        for (int i = 0; i < 5_000_000; i++) {
            // new String(...) forces a fresh backing array, like decoding bytes off the wire
            keep.add(new String(regions[i % regions.length].toCharArray()));
        }
        System.out.println("built " + keep.size());
        // Allocate short-lived garbage so the long-lived strings age and get promoted
        for (int round = 0; round < 50; round++) {
            byte[][] churn = new byte[20_000][];
            for (int j = 0; j < churn.length; j++) churn[j] = new byte[512];
        }
        Thread.sleep(5_000);   // give the concurrent dedup thread time to drain
        System.out.println("done, still holding " + keep.size());
    }
}
# JDK 17+, G1 (the default collector)
java -Xmx2g -XX:+UseG1GC -XX:+UseStringDeduplication \
     -Xlog:gc,stringdedup*=debug:file=dedup.log:time,uptime \
     DedupDemo

# Compare the live heap with and without the flag
jcmd <pid> GC.class_histogram | head -15

At debug level the log reports, per dedup cycle and cumulatively, how many strings were inspected, how many were deduplicated, how many bytes of arrays were freed, and the size of the table. Look for three numbers: the fraction of inspected strings that were duplicates, the bytes deduplicated, and the table size. A high duplicate fraction and a small table is the ideal case, as in this demo: millions of candidates collapsing onto four arrays. A low duplicate fraction with a large, growing table means you are paying to hash strings that are mostly unique.

Then check the result where it matters, in the live heap. Run jcmd <pid> GC.class_histogram before and after (or on two otherwise identical instances). The count of java.lang.String instances will not change. The count and total size of [B (byte arrays) should drop. If they do not, the feature is not finding duplicates in your workload, whatever its log says about work performed.

A worked example: sizing the win before you enable it

Suppose a catalogue service keeps ten million product records in memory. Each record has a region, currency and category field, and a heap histogram shows about 30 million String instances. From a heap dump (Eclipse MAT's duplicate-strings report, or a query in your analyser) you learn that one of the fields has ten million values averaging 20 Latin-1 characters with only 50,000 distinct values.

def dedup_savings(n_strings, avg_len, distinct, latin1=True, header=16, align=8):
    # Bytes freed if every duplicate String shares one backing array.
    payload = avg_len if latin1 else 2 * avg_len
    array_bytes = -(-(header + payload) // align) * align    # round up to 8
    duplicates = max(n_strings - distinct, 0)
    return duplicates * array_bytes

# 10 million field values, average 20 Latin-1 characters, 50,000 distinct values
print(dedup_savings(10_000_000, 20, 50_000) / 2**20, "MiB")   # ~380 MiB

Each array costs 16 + 20 = 36 bytes, rounded to 40. With 9,950,000 duplicates, sharing frees about 398 million bytes, roughly 380 MiB, which on a 4 GiB heap is almost a tenth of it. The String objects themselves, 24 bytes times ten million or about 229 MiB, stay. If you interned the field at the source instead, you would free both, about 600 MiB, at the price of a code change.

Two conditions have to hold for the estimate to become real. The strings must live long enough to become candidates, which they do in a cache that holds records for hours. And they must be created as separate arrays in the first place; if your code already reuses constants or string literals, those are already shared and there is nothing to gain. That second point is why the only trustworthy input is a heap dump of production-like data, not an assumption about your domain.

What it costs

Deduplication is not free, and the costs land in different places.

  • CPU on a background thread. Every candidate's array is hashed and compared. On a service that promotes many unique strings, for example one that caches UUIDs or free text, this is pure overhead.
  • Memory for the table. The table grows with the number of distinct arrays it tracks. With few duplicates it can cost more than it saves.
  • A little extra pause work. Candidate selection adds a check and an enqueue during copying. It is small, but it is in the pause.
  • Delay. Nothing is shared until a string survives several collections and the thread gets to it, so dedup does nothing for a burst of strings that live for seconds, and a heap that is under pressure right now is not rescued by it.
  • No help for young garbage. Allocation rate and young collection frequency are unchanged; only the retained footprint of long-lived strings shrinks.

In practice the feature pays off for services with large, long-lived in-memory data containing repetitive text: caches, session stores, in-memory indexes, parsed configuration, Kafka or event consumers that hold windows of records. It does little for stateless request handlers whose strings die young.

Failure modes and surprises

SymptomLikely causeWhat to do
Flag set, no heap reductionDuplicates die young, or arrays were already sharedCheck the dedup log's duplicate fraction and the [B histogram; if both are flat, turn it off
Higher CPU after enablingMany long-lived unique strings being hashedMeasure the dedup thread's time in the log; disable, or dedupe only specific fields at the source
Old generation still fillsLeak or genuine data growth, not duplicatesDedup reduces a constant factor; a growing heap needs a leak investigation, see GC fundamentals
Code relying on ==Assuming dedup makes equal strings identicalIt never does; use equals, or intern explicitly
Different results on an old JDKCollector not supported before JDK 18On JDK 17 or earlier, only G1 and Shenandoah deduplicate

One more surprise: a raised StringDeduplicationAgeThreshold delays deduplication, and a lowered one spends CPU on strings that may still die. The default of 3 is sensible; change it only with before-and-after measurements.

Alternatives and how to choose

Three tools address duplicated strings, and they sit at different layers.

ApproachWhat is sharedCode changeBest for
-XX:+UseStringDeduplicationBacking arrays onlyNoneLarge long-lived heaps where you cannot or will not change code
String.intern()The whole StringYes, per call siteA known small set of values, when identity comparison is useful
Weak interner or enum at the sourceThe whole String, or no String at allYes, in the parserHot fields with a bounded vocabulary
// Targeted alternative: dedupe at the point of creation, where you know the data
import com.google.common.collect.Interner;
import com.google.common.collect.Interners;

final class Canon {
    private static final Interner<String> REGION = Interners.newWeakInterner();
    static String region(String raw) { return REGION.intern(raw); }
}

// In the parser: record.region = Canon.region(decodedField);

A weak interner (here Guava's) keeps canonical instances only while something references them, so it cannot leak the way an unbounded HashMap cache would. Better still, if a field has a truly fixed vocabulary, such as a currency or status, parse it into an enum and store no String at all. The JVM-level flag is the right choice when the duplicates are spread across code you do not own, such as libraries and frameworks, or when you need a quick win while a proper fix is scheduled. Where the JVM spends its memory more generally is covered in JVM memory and GC.

What to do next

  1. Take a heap dump from a production-like instance and run a duplicate-strings report; note the top fields by wasted bytes.
  2. Estimate the saving with the arithmetic above: duplicates times rounded array size. If it is under a few percent of the heap, stop here.
  3. Check your JDK and collector: G1 or Shenandoah on any supported JDK, or Serial, Parallel or ZGC on JDK 18 and later.
  4. Enable -XX:+UseStringDeduplication with -Xlog:stringdedup*=debug on one canary instance.
  5. Compare the [B histogram and old-generation occupancy after a full warm-up, and the dedup thread's CPU, against a control instance.
  6. Keep the flag if it frees meaningful memory at acceptable CPU; otherwise remove it.
  7. For the top one or two fields, fix duplication at the source with an enum or a weak interner, which frees the String objects too.
Key takeaway: String deduplication lets the garbage collector make equal, long-lived Strings share one backing array. It needs only a flag, works on G1 and Shenandoah and, from JDK 18, on Serial, Parallel and ZGC, and removes at most the array half of each duplicate's cost; the String objects stay distinct. Size the win from a heap dump first, prove it with a byte-array histogram and the stringdedup log, and fix the hottest fields at the source with enums or a weak interner.