Every service that calls another service over HTTP has the same testing problem. The code you own builds a request, sends it, waits, and turns the response into a domain object or an error. The code you do not own sits on the other side of the network: a payment provider, an identity service, a partner API with a sandbox that is down every second Tuesday. Unit tests with an in-process mock never exercise the HTTP client, the serialisation, the timeouts or the retry policy. Tests against the real sandbox are slow, flaky and cannot produce the failures you most need to see, such as a connection reset halfway through a body.

WireMock sits in the gap. It is a real HTTP server, started inside your test JVM or as a separate process, that answers requests according to stub mappings you define and records every request it receives. Your client talks to it over a real socket, so the whole stack from URL building to JSON parsing runs exactly as in production. This article explains how WireMock is built, how to wire it into JUnit 5, how matching, verification, scenarios, fault injection and templating work, how to run it outside tests, and how to keep stubs honest as the real API changes. Examples use WireMock 3.x, the current stable line.

How WireMock works

WireMock has four moving parts. The server is an embedded Jetty instance listening on one port (or two, if you add HTTPS). The stub mappings are pairs of a request pattern and a response definition. The request journal is an in-memory log of every request that arrived, whether or not anything matched it. The admin API, served under /__admin on the same port, lets anything that speaks HTTP create mappings, query the journal, reset state and start recordings. The Java DSL you use in tests is a thin client of that admin API, or a direct call into it when the server runs in the same JVM.

Code under testHttpClient, retries, parserHTTPWireMock serverJetty, one port per instanceevery requestRequest journalkept for verify and debuggingmatchStub mappingspriority, then newest firstchosen stubResponse definitionbody, delay, fault, templateresponse or 404 with diffTest codestubFor, verifyadmin APIAdmin API under /__adminmappings, requests, scenarios, recordings, resetThe test configures stubs and reads the journal through the admin API; the client only ever sees HTTP.
WireMock as a server: requests hit Jetty, are logged to the journal, matched against stub mappings and answered from the chosen response definition, while the test drives everything through the admin API.

When a request arrives, WireMock evaluates it against every mapping. Each mapping that matches on all of its conditions is a candidate; the winner is the one with the highest priority (priority 1 beats priority 5), and among equal priorities the most recently added mapping wins. That last rule is the source of a classic surprise: a general stub registered in a setup method silently overrides nothing, but a general stub registered after a specific one shadows it. If nothing matches, WireMock returns 404 and, in its default configuration, a plain-text diff against the closest mapping, which is the single most useful debugging output it produces.

Because it is a real server, WireMock tests exercise things in-process mocks cannot: DNS-free but real TCP connections, the HTTP client's connection pool, header casing, content negotiation, gzip, chunked bodies and timeouts. That is also its limit. WireMock knows nothing about the provider's business rules. It returns what you told it to return, so a stub is a claim about the provider, and claims go stale.

Setting up WireMock 3 with JUnit 5

WireMock 3.x is published under the org.wiremock group id; the 2.x line used com.github.tomakehurst, and the Java package names still carry that older prefix. The 3.x line needs Java 11 or later, ships Jetty 11, and offers a separate wiremock-jetty12 artifact for projects that must align on Jetty 12. The project has announced that 3.x is in maintenance while 4.x, currently in beta, moves the baseline to Java 17 and Jetty 12 and splits the core from optional dependencies. For new work today, use the latest 3.x release and watch the 4.x notes.

There are two JUnit 5 entry points. The declarative @WireMockTest annotation starts one server per test class on a random port and injects a WireMockRuntimeInfo with its base URL. The programmatic WireMockExtension gives you full configuration, several servers per class and a handle to call methods on:

class PaymentClientTest {

    @RegisterExtension
    static WireMockExtension provider = WireMockExtension.newInstance()
        .options(wireMockConfig().dynamicPort().usingFilesUnderClasspath("wiremock"))
        .failOnUnmatchedRequests(true)
        .build();

    PaymentClient client;

    @BeforeEach
    void setUp() {
        client = new PaymentClient(
            URI.create(provider.baseUrl()),
            Duration.ofSeconds(2));          // the same timeout production uses
    }

    @Test
    void chargesTheCard() {
        provider.stubFor(post(urlEqualTo("/v1/charges"))
            .withHeader("Idempotency-Key", matching("[0-9a-f-]{36}"))
            .withRequestBody(matchingJsonPath("$.amount", equalTo("1250")))
            .willReturn(okJson("{\"id\":\"ch_1\",\"status\":\"succeeded\"}")));

        ChargeResult result = client.charge(new Charge(1250, "EUR"));

        assertEquals("succeeded", result.status());
        provider.verify(1, postRequestedFor(urlEqualTo("/v1/charges")));
    }
}

Three choices in that snippet matter more than they look. A dynamic port lets tests run in parallel and on shared CI agents without port clashes. Injecting the base URL into the client, rather than overriding a hostname, keeps the production code path intact. And failOnUnmatchedRequests(true) turns any request that no stub matched into a test failure, so a typo in a path cannot pass silently because the client happened to treat a 404 as an empty result.

Request matching without brittleness

A request pattern combines conditions on method, URL, headers, query parameters, cookies and body, and all of them must hold. The URL condition comes in four forms with different strictness: urlEqualTo compares path and query string exactly, urlPathEqualTo compares the path only so query parameters can be matched separately, and urlMatching and urlPathMatching take regular expressions. Prefer the path forms plus withQueryParam, because exact query strings break when a client reorders parameters.

Body matching is where stubs become either precise or brittle. equalToJson compares structurally and can be told to ignore array order and extra fields, which makes it tolerant of new optional fields. matchingJsonPath asserts only the fields the test cares about. equalTo on a raw body is almost always too strict for JSON. A useful rule: match on what changes the provider's answer, such as the amount and currency, and assert the rest with verify after the call, where a mismatch produces a readable failure instead of a 404.

Verification and stateful scenarios

Stubbing answers the question what does the client do with this response. Verification answers the other question: did the client send what the provider expects. verify(count, pattern) checks the journal after the action. Use it for side-effecting calls, where sending twice is a bug, and for contracts such as an idempotency key that must be identical across retries. findAll(pattern) returns the logged requests so you can assert on headers and bodies with your normal assertion library.

Retries are where a stub and the journal earn their keep together. A scenario is a small state machine attached to a set of stubs: each stub can require a state and move the scenario to a new one. The following pair makes the first call fail and the second succeed, which is exactly the sequence a retry policy needs to be tested against:

provider.stubFor(post("/v1/charges").inScenario("flaky-provider")
    .whenScenarioStateIs(Scenario.STARTED)
    .willReturn(aResponse().withStatus(503).withHeader("Retry-After", "1"))
    .willSetStateTo("recovered"));

provider.stubFor(post("/v1/charges").inScenario("flaky-provider")
    .whenScenarioStateIs("recovered")
    .willReturn(okJson("{\"id\":\"ch_1\",\"status\":\"succeeded\"}")));

ChargeResult result = client.charge(new Charge(1250, "EUR"));

List<LoggedRequest> sent = provider.findAll(postRequestedFor(urlEqualTo("/v1/charges")));
assertEquals(2, sent.size());
assertEquals(sent.get(0).getHeader("Idempotency-Key"),
             sent.get(1).getHeader("Idempotency-Key"));   // retry must reuse the key

Scenario state lives in the server, so it leaks between tests unless you reset it. The JUnit extension resets mappings, the journal and scenarios between tests by default; a server you manage yourself needs an explicit resetAll() or a call to POST /__admin/reset.

Injecting latency and faults

The failures that take services down are rarely clean 500 responses. They are slow responses, connections that die mid-body and payloads that are not what the content type promised. WireMock can produce each of them on demand, which is the main reason to prefer it over a hand-rolled fake:

Response optionWhat the client seesWhat it tests
withFixedDelay(3000)headers arrive after 3 srequest timeout, thread pool exhaustion
withChunkedDribbleDelay(5, 4000)body trickles in 5 chunks over 4 sread timeouts that only cover the first byte
withFault(Fault.CONNECTION_RESET_BY_PEER)TCP resetretry classification of I/O errors
withFault(Fault.EMPTY_RESPONSE)connection closed, no byteserror mapping for no response
withFault(Fault.MALFORMED_RESPONSE_CHUNK)200 then a broken chunkpartial body handling, no half-parsed objects
withFault(Fault.RANDOM_DATA_THEN_CLOSE)garbage, then closeparser robustness

The dribble delay deserves special attention. Many HTTP clients separate a connect timeout from a response timeout, and some response timeouts stop counting once headers arrive. A provider that sends headers promptly and then stalls the body hangs such a client far beyond its configured limit; one dribble test per client catches it. Pair these with assertions on what your code does next: which exception type surfaces, whether a retry happens, whether a circuit breaker opens, and whether the caller gets a clean error rather than a partially built object.

Templating, JSON mappings and recording

Static responses are enough for most tests. When a response must echo something from the request, such as an id in the path or a correlation header, WireMock offers Handlebars response templating. In 3.x, templating is enabled by default when the server is started programmatically, but it applies only to stubs that list the response-template transformer; globalTemplating(true) applies it to every stub and templatingEnabled(false) turns it off. Stubs can also live as JSON files under a mappings directory, with large bodies in a sibling __files directory:

{
  "request": {
    "method": "GET",
    "urlPathPattern": "/v1/customers/[a-z0-9]+"
  },
  "response": {
    "status": 200,
    "headers": { "Content-Type": "application/json" },
    "body": "{\"id\":\"{{request.path.[2]}}\",\"tier\":\"gold\"}",
    "transformers": ["response-template"]
  }
}

Templating is a sharp tool. Each template is logic in a test fixture, untested and invisible in the Java code. Use it to echo values, not to simulate provider behaviour; once a stub contains conditionals, you are maintaining a second implementation of someone else's service.

WireMock can also record. Started as a proxy in front of the real API, it forwards traffic and captures each exchange as a mapping through POST /__admin/recordings/start and /__admin/recordings/stop. Recording is the fastest way to bootstrap realistic stubs, especially for large or awkward payloads. Treat the output as raw material: strip credentials and personal data, replace exact body matches with JSON path matches, and delete the dozens of incidental calls you do not want to pin down.

Running WireMock outside tests

The same server runs outside tests. The standalone distribution and the official wiremock/wiremock Docker image read mappings from a mounted directory and expose the admin API, which makes WireMock useful as a fake dependency in a docker-compose stack for local development, for end-to-end tests of a service built as a container, or for a frontend team waiting on a backend that does not exist yet:

docker run --rm -p 8080:8080 \
  -v "$PWD/wiremock:/home/wiremock" \
  wiremock/wiremock:<pinned-3.x-tag>

curl -s localhost:8080/__admin/mappings | jq '.mappings | length'
curl -s localhost:8080/__admin/requests/unmatched

Pin the image tag so a major-version change never arrives unannounced. When WireMock runs in a container next to a service under test, the Testcontainers WireMock module or a generic container gives you the same lifecycle control as any other dependency.

Failure modes

Most WireMock pain comes from a handful of predictable mistakes.

  • Stub drift. The provider adds a required field or changes an error format; every test stays green and production breaks. Mitigate by recording fresh traffic periodically and diffing it against committed stubs, or by moving the agreement into consumer-driven contract tests that the provider runs too.
  • Over-specified matching. Exact body equality means any harmless client change breaks dozens of tests. Match on what selects the response; assert the rest with verify.
  • Shadowed stubs. A broad stub added late wins over an earlier specific one. Use priorities explicitly when stubs overlap, and read the 404 diff when a match is missed.
  • Leaked state. Scenarios and journal entries surviving across tests produce order-dependent failures. Rely on the extension reset, or reset in @BeforeEach.
  • Fixed ports. A hard-coded 8080 collides on shared agents and blocks parallel runs. Always use dynamic ports and inject the base URL.
  • Testing WireMock instead of your code. A test that stubs a response and asserts the client returned it proves little. Assert the decisions your code makes: mapping, retry, error classification, timeouts.

Trade-offs and where WireMock fits

WireMock is one of three tools for the same seam, and choosing well saves churn. An in-process mock built with Mockito replaces your own client interface and is right for testing business logic that only cares about the result of a call. WireMock replaces the remote service at the HTTP level and is right for testing the client itself: serialisation, headers, timeouts, retries and fault handling, written with the Java HttpClient or any other library. A real dependency in a container is right when you own the dependency or a faithful emulator exists, such as a database. For third-party APIs a container is rarely available, which is why WireMock is the usual choice there.

The cost is maintenance of stubs that encode assumptions about someone else's system. Keep them few, keep them close to the client they test, and add a separate, smaller suite that calls the real sandbox on a schedule rather than on every commit. The same layering appears in the agent CI lanes article, where recorded HTTP exchanges replay through an adapter in the fast lane and live calls run only in a gated lane.

What to do next

  1. Pick one outbound HTTP client in your service and add a WireMock test that exercises the real client class against a dynamic port, using the production timeouts.
  2. Enable failOnUnmatchedRequests(true) so unmatched calls fail loudly.
  3. Replace exact body matches with matchingJsonPath or equalToJson with ignored extras, and move incidental checks into verify.
  4. Add one scenario test for retry: a 503 then a 200, asserting the retry count and that the idempotency key is reused.
  5. Add fault tests: a fixed delay above your timeout, a dribbled body and a connection reset, asserting the exception type and the caller-visible error.
  6. Record a fresh exchange from the provider sandbox, scrub it, and diff it against your committed stubs; schedule that diff monthly.
  7. Pin the WireMock version (3.x today), and read the 4.x migration notes before the Java 17 baseline reaches your build.
  8. Write down which tests use Mockito, WireMock and containers, so the next person adds tests at the right layer; the JUnit 5 guide covers extension ordering if you combine them.
Key takeaway: WireMock is a real HTTP server that answers from stub mappings and logs every request, so it tests the parts of an HTTP client that in-process mocks skip: serialisation, timeouts, retries and fault handling. Use dynamic ports, fail on unmatched requests, match only on what selects a response and verify the rest. Use scenarios and faults to test the bad days, and refresh stubs from real traffic so they do not drift from the provider.