Since Java 11 the JDK has shipped a modern HTTP client in the java.net.http module. It speaks HTTP/1.1 and HTTP/2, supports blocking and asynchronous calls, streams bodies with back-pressure, and needs no dependency. It replaced HttpURLConnection for most new code, and in many services it can replace OkHttp or Apache HttpClient as well.

The API is small, but some defaults are the opposite of what people expect, and most production incidents come from those defaults and lifecycle mistakes. This article explains the object model, then the settings that matter under load, a worked sizing example, failure modes and a checklist.

Advertisement

The object model

There are three main types. HttpClient holds configuration (protocol version, timeouts, redirect policy, proxy, TLS, authenticator, executor) and owns the connections. It is immutable once built and safe to share between threads. HttpRequest is an immutable description of one call: URI, method, headers, a per-request timeout and a BodyPublisher that produces the request body. Because it is immutable, you can build a request once and send it many times. HttpResponse<T> carries the status code, headers, the request that produced it, the protocol version used and a body of type T.

The body type is chosen by a BodyHandler<T>. When the status line and headers arrive, the client calls the handler with them, and the handler returns a BodySubscriber<T>, a reactive-streams subscriber that turns body bytes into the result. So you can choose how to read the body after seeing the status.

java.net.http: one long-lived HttpClient, many immutable requests, a body handler per responseYour codesend / sendAsyncHttpRequestimmutable, reusableHttpClientconfig + executorBodyPublisherrequest bytes outHTTP/1.1 poolidle keep-alive cacheHTTP/2 connstreams per originServerTLS, ALPNBodyHandlerstatus + headers inBodySubscriberbytes to type TrequestcreatesbytesThe handler sees the status and headers first and chooses a subscriber; the subscriber turns body bytes into the result type.Connections belong to the client, so creating a client per request throws the pool away every time.
The request path through HttpClient. Connections are owned by the client, so the client's lifetime decides whether connections are reused.

Building a client: the defaults that bite

The builder's defaults are documented in the javadoc and several are surprising. The default version is HTTP/2, and the client falls back to HTTP/1.1 when the server or a proxy cannot do HTTP/2. The default redirect policy is Redirect.NEVER: a 301 or 302 comes back to you as the response, not followed. There is no default connect timeout, so a connection to a black-holed address waits for the operating system to give up, which can take a long time. If you do not set an executor, the client uses an internal default one for its asynchronous work.

So the first rule is to always set a connect timeout and to decide the redirect policy on purpose. Redirect.NORMAL follows redirects except from HTTPS to HTTP; Redirect.ALWAYS follows those too and is rarely what you want. The client retries a redirected or failed request at most jdk.httpclient.redirects.retrylimit times, 5 by default.

import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.time.Duration;
import java.util.concurrent.Executors;

// One client for the whole application (or one per upstream with different settings).
HttpClient client = HttpClient.newBuilder()
        .version(HttpClient.Version.HTTP_2)          // the default; falls back to HTTP/1.1
        .connectTimeout(Duration.ofSeconds(2))       // default: none
        .followRedirects(HttpClient.Redirect.NORMAL) // default: NEVER
        .executor(Executors.newVirtualThreadPerTaskExecutor())
        .build();

HttpRequest request = HttpRequest.newBuilder(URI.create("https://prices.internal/v1/quote?sku=A12"))
        .timeout(Duration.ofMillis(800))             // per request
        .header("Accept", "application/json")
        .GET()
        .build();

HttpResponse<String> response = client.send(request, HttpResponse.BodyHandlers.ofString());
if (response.statusCode() != 200) {
    throw new IllegalStateException("quote failed: " + response.statusCode());
}
String json = response.body();

A client holds selector machinery and a connection pool. Create one per application, or per upstream with different settings, and inject it; a client per request opens and discards TCP and TLS connections every time.

Advertisement

Timeouts and deadlines

There are two built-in timeouts. The connect timeout on the builder bounds how long establishing a connection may take and raises HttpConnectTimeoutException. The request timeout on HttpRequest.Builder.timeout is defined as the time within which the response must be received, and raises HttpTimeoutException (the connect exception is a subclass of it, so catch the specific one first).

Neither is a complete end-to-end deadline for every case. With a streaming handler such as ofInputStream or ofLines, you get the response object as soon as the headers arrive and then read the body at your own pace, so do not rely on the request timeout to bound a slow download. For a hard deadline on the whole call, use sendAsync and put orTimeout on the future, and propagate a deadline budget from your own caller rather than using a fixed value per hop. A timeout also does not tell you whether the server processed the request, which matters for retries below.

Sync, async and virtual threads

send blocks the calling thread until the handler finishes and throws IOException or InterruptedException. sendAsync returns a CompletableFuture<HttpResponse<T>> immediately; dependent stages run on the client's executor. Before virtual threads, sendAsync was the way to make thousands of concurrent calls without thousands of platform threads. On Java 21 and later, blocking send on virtual threads gives the same scalability with plain sequential code, and exceptions and stack traces are easier to read.

Both styles need an explicit concurrency limit. Nothing in the client stops you from starting ten thousand requests at once; the upstream will then see a spike, the HTTP/1.1 pool will open a connection per concurrent request, and queued work will time out. Put a semaphore or a bounded executor in front of the client and size it from the upstream's capacity, as in this fan-out:

// Bounded fan-out with sendAsync: at most 50 requests in flight, each with its own deadline.
Semaphore inFlight = new Semaphore(50);

CompletableFuture<Quote> quote(String sku) {
    HttpRequest req = HttpRequest.newBuilder(URI.create(BASE + "/v1/quote?sku=" + sku))
            .timeout(Duration.ofMillis(800))
            .build();
    inFlight.acquireUninterruptibly();
    return client.sendAsync(req, HttpResponse.BodyHandlers.ofString())
            .orTimeout(1, TimeUnit.SECONDS)                 // whole-call deadline
            .thenApply(r -> {
                if (r.statusCode() != 200) throw new UpstreamException(r.statusCode());
                return Quote.parse(r.body());
            })
            .whenComplete((q, e) -> inFlight.release());
}

List<CompletableFuture<Quote>> all = skus.stream().map(this::quote).toList();
CompletableFuture.allOf(all.toArray(CompletableFuture[]::new)).join();

Request and response bodies

Publishers produce request bodies: BodyPublishers.ofString, ofByteArray, ofFile, ofInputStream (which takes a supplier so the body can be re-created on retry or redirect) and noBody. Set Content-Type yourself; the client does not infer it. For JSON, serialize with your library of choice to a byte array or string. There is no built-in JSON support, and a custom BodyHandler that parses JSON is easy to write with BodySubscribers.mapping.

Handlers decide how responses land in memory. ofString and ofByteArray buffer the whole body, which is fine for small documents and a memory risk for large ones. ofFile streams to disk. ofInputStream and ofLines hand you a stream to consume, and the connection is only released when you have read the body or closed the stream. discarding drains and drops the body.

// Large download: stream to disk, never into a String.
HttpResponse<Path> file = client.send(req,
        HttpResponse.BodyHandlers.ofFile(Path.of("/data/export.csv")));

// Line-by-line processing; the stream MUST be closed or the connection is never released.
HttpResponse<Stream<String>> lines = client.send(req, HttpResponse.BodyHandlers.ofLines());
try (Stream<String> s = lines.body()) {
    s.filter(l -> !l.isBlank()).forEach(this::ingest);
}

// Only the status matters: drain and drop the body so the connection can be reused.
HttpResponse<Void> ping = client.send(req, HttpResponse.BodyHandlers.discarding());

The javadoc warns that an unconsumed or unclosed streaming body can stop the request completing and the client being collected; an unclosed InputStream on an error path is the classic leak.

Connections, HTTP/2 and pooling

Over HTTP/1.1 a connection carries one request at a time. The client keeps idle connections in a keep-alive cache and reuses them for the same origin. According to the current JDK documentation, idle connections are kept for jdk.httpclient.keepalive.timeout seconds, 30 by default, and the cache size jdk.httpclient.connectionPoolSize defaults to 0, meaning unbounded. Concurrency therefore equals connection count: 200 concurrent calls to one host means up to 200 sockets.

Over HTTP/2 the client multiplexes many concurrent requests as streams on one connection per origin, up to the server's advertised concurrent-stream limit. That removes connection-count problems but concentrates everything on one TCP connection: one lost packet delays every stream on it, and a load balancer that spreads connections, not requests, will send all your traffic to one backend. If backends behind an HTTP/2 load balancer look unevenly loaded, that is usually why.

For plain http:// URIs the client may attempt an HTTP/2 upgrade and fall back to HTTP/1.1; for HTTPS it negotiates the protocol during the TLS handshake. response.version() tells you which one was used, which is worth logging. JEP 517 adds HttpClient.Version.HTTP_3 in JDK 26, over QUIC, with fallback to HTTP/2 or HTTP/1.1 when HTTP/3 is not available; treat it as opt-in and measure before relying on it.

Retries and idempotency

The client has limited automatic retry of its own. It retries connection failures unless jdk.httpclient.disableRetryConnect is true, and it does not automatically retry non-idempotent methods unless jdk.httpclient.enableAllMethodRetry is set, which defaults to false. Everything else is your job: status-code retries, backoff and deciding what is safe to repeat.

The rule is to retry only when a repeat cannot cause harm. A connect failure is always safe because nothing was sent. A timeout is not, because the server may have processed the request and only the response was lost. GET, HEAD, PUT and DELETE are idempotent by definition; a POST is safe only when the server de-duplicates it, typically with an idempotency key header. Retry 429 and 503 with backoff and honour Retry-After when the server sends it.

static final Set<Integer> RETRYABLE = Set.of(429, 502, 503, 504);

<T> HttpResponse<T> sendWithRetry(HttpRequest req, HttpResponse.BodyHandler<T> h)
        throws IOException, InterruptedException {
    boolean idempotent = Set.of("GET", "HEAD", "PUT", "DELETE").contains(req.method())
            || req.headers().firstValue("Idempotency-Key").isPresent();
    int maxAttempts = idempotent ? 3 : 1;
    for (int attempt = 1; ; attempt++) {
        try {
            HttpResponse<T> r = client.send(req, h);
            if (!RETRYABLE.contains(r.statusCode()) || attempt == maxAttempts) return r;
            sleep(backoff(attempt, r.headers().firstValue("Retry-After")));
        } catch (HttpConnectTimeoutException | ConnectException e) {
            if (attempt == maxAttempts) throw e;             // safe: nothing was sent
            sleep(backoff(attempt, Optional.empty()));
        } catch (HttpTimeoutException e) {
            if (!idempotent || attempt == maxAttempts) throw e; // may have been processed
            sleep(backoff(attempt, Optional.empty()));
        }
    }
}

Duration backoff(int attempt, Optional<String> retryAfter) {
    long base = 100L << attempt;                              // 200, 400, 800 ms
    long jitter = ThreadLocalRandom.current().nextLong(base);
    long hint = retryAfter.filter(v -> v.matches("\\d+"))    // seconds form only
            .map(v -> Long.parseLong(v) * 1000).orElse(0L);
    return Duration.ofMillis(Math.max(hint, base / 2 + jitter));
}

Keep a retry budget across the service (for example, retries may add at most 10% to request volume) so a struggling upstream is not hit harder.

Security, headers and lifecycle

TLS is configured with sslContext and sslParameters on the builder; the default uses the JVM's trust store and verifies host names. Load an internal CA into a trust store rather than disabling verification. The headers connection, content-length, expect, host and upgrade are restricted and managed by the client; jdk.httpclient.allowRestrictedHeaders lifts that, but the documentation says it is for testing only.

Since Java 21 HttpClient implements AutoCloseable: shutdown() stops accepting requests without waiting, awaitTermination waits, shutdownNow() interrupts active operations, and close() shuts down in order and waits. Close it after your server has drained. For debugging, -Djdk.httpclient.HttpClient.log=errors,requests,headers logs through the platform logging API; in production, record per-upstream latency, status, timeouts, retries and negotiated version in a thin wrapper.

Worked example: sizing a client for a pricing service

A checkout service calls an internal pricing API for every cart line. Peak load is 400 pricing calls per second, the API's p50 latency is 40 ms and its p99 is 250 ms, and the checkout request has a 1.5-second budget of which pricing may use 800 ms.

Little's law gives average concurrency as arrival rate times latency: 400 per second times 0.04 seconds is 16 calls in flight on average. At p99 latency the worst sustained case is 400 times 0.25, which is 100. Set the semaphore to about 120 and treat hitting it as a signal, not something to raise blindly. Over HTTP/2 those 100 calls fit within a typical server limit of concurrent streams on a single connection; over HTTP/1.1 they mean up to 100 sockets from each checkout instance, so with 30 instances the pricing tier must accept 3,000 connections.

Timeouts follow the budget: a 200 ms connect timeout (the network is local), a 350 ms request timeout per attempt, one retry for idempotent GETs on connect failure or 503, and an orTimeout of 800 ms on the whole call so two slow attempts cannot exceed the budget.

Failure modes

SymptomCauseFix
Threads stuck for minutes on connectNo connect timeout (default)Always set connectTimeout
Unexpected 301 or 302 responsesRedirect policy NEVER by defaultChoose NORMAL deliberately, or handle redirects
Socket count and latency climbNew client per request, or unclosed streamed bodiesOne shared client; close every stream in try-with-resources
OutOfMemoryError on big responsesofString or ofByteArray on large bodiesofFile or ofInputStream with bounded reads
Duplicate orders after timeoutsRetrying non-idempotent POSTsIdempotency keys, or no retry after send
One backend hot behind the load balancerHTTP/2 multiplexes onto one connectionRequest-level balancing, or several clients
Async callbacks slow under loadSmall executor shared with other workDefault or virtual-thread executor

Trade-offs against other clients

The JDK client has no dependencies and fits virtual threads naturally, but has no interceptors, retry policy or response cache; OkHttp and Apache HttpClient 5 offer those, and Netty gives lower-level control. For most service calls, the JDK client plus a thin wrapper is enough.

To go further, read CompletableFuture in depth for composing async calls, virtual threads for the blocking style, and async-profiler for finding where client time actually goes.

What to do next

  1. Search your code for HttpClient.newHttpClient() and newBuilder() calls on request paths and replace them with one injected client.
  2. Set a connect timeout, a per-request timeout and an overall deadline on every call.
  3. Decide the redirect policy explicitly.
  4. Put a concurrency limit in front of each upstream, sized with Little's law from p99 latency.
  5. Stream large bodies to files or streams and close them in try-with-resources, including on error paths.
  6. Retry only safe cases with capped, jittered backoff, and add idempotency keys to POSTs that need retries.
  7. Record latency, status, timeouts, retries and protocol version per upstream, and close the client on shutdown.
Key takeaway: The JDK HttpClient is one long-lived, thread-safe object that owns connections, plus immutable requests and body handlers that decide how each response is read. Its defaults need attention: no connect timeout, redirects not followed, and an unbounded HTTP/1.1 keep-alive cache. Share one client, set connect, request and overall deadlines, limit concurrency explicitly, stream and close large bodies, retry only what is safe, and close the client on shutdown. With virtual threads, blocking send is now the simplest way to use it at scale.