An integration test is only as honest as the dependencies it runs against. H2 pretending to be PostgreSQL accepts SQL that production rejects, an embedded broker skips the partitioning that breaks your consumer, and a mocked HTTP client never times out. Testcontainers is a Java library that starts real services in throwaway Docker containers from inside your test, hands you their address, and removes them when the test JVM ends.
This guide explains the mechanism first, then the decisions that make a suite fast and reliable: which lifecycle scope to use, how to know a container is ready, how containers talk to each other, how to wire Spring Boot, and what breaks in CI. Examples use Testcontainers for Java 2.0.5, the current release at the time of writing, with JUnit 5.
What happens when a container starts
Calling start() runs six steps. The library finds a Docker engine: the local socket, DOCKER_HOST, or a configured remote or cloud engine. It starts a small reaper container called Ryuk, unless disabled. It pulls the image if it is not cached, creates the container with a session label, and starts it. It runs the container's wait strategy until the service is ready. Finally it exposes the mapped ports, so your test reads a host and a random host port from the container object.
Two details explain most surprises. First, ports are remapped every time: the database listens on 5432 inside the container but on something like 55017 on the host, which is why you must call getJdbcUrl() or getMappedPort(5432) instead of hard-coding. Random ports are what let parallel builds on one machine coexist. Second, cleanup does not depend on your code. Ryuk keeps a connection to the test JVM; when the JVM exits, even through kill -9, Ryuk removes every container carrying that session's label.
Setting up 2.x: the renamed artifacts
Testcontainers 2.0 renamed every module with a testcontainers- prefix, so org.testcontainers:postgresql became org.testcontainers:testcontainers-postgresql. It also moved container classes into per-module packages such as org.testcontainers.postgresql instead of org.testcontainers.containers, and dropped JUnit 4 support. The module classes in the 2.0 documentation are constructed without a type parameter, as in new PostgreSQLContainer("postgres:16-alpine"). Import the bill of materials once so module versions stay aligned:
<!-- Maven: import the BOM once, then add modules without versions -->
<dependencyManagement>
<dependencies>
<dependency>
<groupId>org.testcontainers</groupId>
<artifactId>testcontainers-bom</artifactId>
<version>2.0.5</version>
<type>pom</type>
<scope>import</scope>
</dependency>
</dependencies>
</dependencyManagement>
<dependencies>
<dependency>
<groupId>org.testcontainers</groupId>
<artifactId>testcontainers-junit-jupiter</artifactId>
<scope>test</scope>
</dependency>
<dependency>
<groupId>org.testcontainers</groupId>
<artifactId>testcontainers-postgresql</artifactId>
<scope>test</scope>
</dependency>
</dependencies>
// Gradle equivalent
testImplementation platform("org.testcontainers:testcontainers-bom:2.0.5")
testImplementation "org.testcontainers:testcontainers-junit-jupiter"
testImplementation "org.testcontainers:testcontainers-postgresql"Migrating from 1.x is mostly mechanical: rename the artifacts, fix the imports, and move any remaining JUnit 4 @Rule containers to the Jupiter extension or to manual start() calls. OpenRewrite publishes recipes for the dependency rename if you have many modules.
A first test against real PostgreSQL
The JUnit Jupiter extension, activated by @Testcontainers, starts and stops fields marked @Container. The field's modifier sets the scope: a static field is started once before the first test in the class and stopped after the last, while an instance field is started and stopped around every test method.
import org.junit.jupiter.api.Test;
import org.testcontainers.junit.jupiter.Container;
import org.testcontainers.junit.jupiter.Testcontainers;
import org.testcontainers.postgresql.PostgreSQLContainer;
import java.sql.Connection;
import java.sql.DriverManager;
import static org.junit.jupiter.api.Assertions.assertEquals;
@Testcontainers
class OrderRepositoryIT {
// static: started once before the first test in this class, stopped after the last
@Container
static PostgreSQLContainer pg = new PostgreSQLContainer("postgres:16-alpine")
.withInitScript("schema.sql"); // classpath resource, runs once at start
@Test
void savesAndReadsBackAnOrder() throws Exception {
try (Connection conn = DriverManager.getConnection(
pg.getJdbcUrl(), pg.getUsername(), pg.getPassword())) {
var repo = new OrderRepository(conn);
long id = repo.insert("sku-42", 3);
assertEquals(3, repo.findQuantity(id));
}
}
}Pin the image tag. postgres:latest makes a test that passed yesterday fail today because a major version changed underneath it. Use the same major version you run in production; the point of the exercise is to test against what you ship. If you use Flyway or Liquibase, run your real migrations instead of an init script, so the schema in tests is the schema production gets.
Choosing a lifecycle scope
Container start time dominates suite time. A PostgreSQL container typically starts in a few seconds once the image is cached, a Kafka broker takes longer, and multiply that by every test and the suite becomes unusable. Choose the scope deliberately:
| Scope | How | Isolation | Cost |
|---|---|---|---|
| Per method | instance @Container field | Fresh service every test | Start time times test count |
| Per class | static @Container field | Shared within one class | One start per class |
| Per JVM (singleton) | static initializer calls start() | Shared by the whole run; tests clean up | One start per build |
| Across runs (reuse) | withReuse(true) plus opt-in | Survives between runs | Near zero locally; not for CI |
The singleton pattern is what most large suites settle on. You manage the lifecycle yourself and accept one consequence: tests share state, so each test must clean up or use unique keys.
// One container for the whole test JVM, shared by every class that extends this.
// No @Testcontainers / @Container here: we own the lifecycle, Ryuk cleans up at JVM exit.
public abstract class PostgresIT {
protected static final PostgreSQLContainer PG =
new PostgreSQLContainer("postgres:16-alpine");
static {
PG.start();
}
// Tests share one database, so each test must leave it clean or use unique keys.
@AfterEach
void truncate() throws Exception {
try (var c = DriverManager.getConnection(PG.getJdbcUrl(), PG.getUsername(), PG.getPassword());
var s = c.createStatement()) {
s.execute("TRUNCATE orders, order_lines RESTART IDENTITY CASCADE");
}
}
}Reuse goes further and keeps the container alive between runs on a developer machine. It is opt-in twice: set TESTCONTAINERS_REUSE_ENABLE=true or testcontainers.reuse.enable=true in ~/.testcontainers.properties, and call withReuse(true) on the container. The documentation marks it experimental and states that reusable containers are not suited for CI, because they are never removed automatically.
One more constraint: the Jupiter extension has only been tested with sequential test execution, and the documentation calls parallel use unsupported. If you enable JUnit parallelism, prefer the singleton pattern with tests that do not share rows.
Readiness: wait strategies
A container being "running" means its process started, not that the service accepts work. The database modules ship a sensible default; for GenericContainer you choose. Waiting for an open port is the weakest signal, because many servers bind the port before they finish initialising. An HTTP health probe or a specific log line is stronger:
GenericContainer<?> api = new GenericContainer<>("ghcr.io/acme/pricing-stub:1.4.2")
.withExposedPorts(8080)
// ready only when the app says so, not when the port is merely open
.waitingFor(Wait.forHttp("/health").forStatusCode(200)
.withStartupTimeout(Duration.ofSeconds(60)));
GenericContainer<?> legacy = new GenericContainer<>("acme/batch-engine:7")
.withExposedPorts(9000)
// this image prints a line once its caches are warm
.waitingFor(Wait.forLogMessage(".*engine ready.*\\n", 1));When a test fails intermittently with "connection refused" or a first query that times out, the wait strategy is the first suspect. Raise the startup timeout for CI machines, which are slower and pull images cold, rather than adding Thread.sleep to tests.
Several containers: networks and addresses
When containers must reach each other, for example a worker container that reads from Kafka and writes to PostgreSQL, put them on a shared network and give each an alias. There are now two addresses for every service, and confusing them is the most common multi-container bug:
Network net = Network.newNetwork();
KafkaContainer kafka = new KafkaContainer("apache/kafka:3.8.0")
.withNetwork(net);
PostgreSQLContainer pg = new PostgreSQLContainer("postgres:16-alpine")
.withNetwork(net)
.withNetworkAliases("db"); // other containers reach it as db:5432
GenericContainer<?> worker = new GenericContainer<>("acme/order-worker:latest")
.withNetwork(net)
.withEnv("DB_URL", "jdbc:postgresql://db:5432/test") // container-to-container: alias + internal port
.dependsOn(pg, kafka);
// The test JVM is OUTSIDE the network: it uses the mapped address instead.
String fromTest = pg.getJdbcUrl(); // jdbc:postgresql://localhost:<mapped>/testContainer-to-container traffic uses the alias and the internal port: db:5432. Traffic from the test JVM uses the mapped host and port from the container object. Kafka makes this sharper, because a broker advertises the addresses clients must reconnect to; the Kafka module configures its listeners so that the mapped address works from the test, but if you build a broker from a generic image you must configure an advertised listener for each side yourself.
Spring Boot: @ServiceConnection
Spring Boot can read connection details straight from a container. Annotate the container with @ServiceConnection and Boot creates the connection beans from the container's host, port and credentials, so there is no property wiring. For services without a supported connection type, @DynamicPropertySource registers properties from container getters instead. Spring Boot, in depth covers how those properties reach auto-configuration.
@Testcontainers
@SpringBootTest
class CheckoutIT {
@Container
@ServiceConnection // Boot reads host, port and credentials from the container
static PostgreSQLContainer pg = new PostgreSQLContainer("postgres:16-alpine");
@Autowired CheckoutService checkout;
@Test
void reservesStock() { /* ... */ }
}Watch the interaction with Spring's test-context cache. Contexts are cached by configuration, and a context that holds a connection to a stopped per-class container fails confusingly in the next class that reuses it. Combining a JVM-wide singleton container with a stable context configuration avoids this and keeps both caches effective.
Worked example: an order service
Take a service that accepts an order over HTTP, writes it to PostgreSQL in a transaction, and publishes an OrderPlaced event to Kafka through an outbox table. A useful test proves three things: the row is written, the event is published exactly once, and nothing is published if the transaction rolls back.
Set up a singleton PostgreSQL container and a singleton Kafka container in a base class. The test posts an order through the real HTTP layer, then polls a real Kafka consumer for up to ten seconds for the event keyed by the order ID, and queries the outbox to confirm the row was marked sent. A second test forces a constraint violation and asserts that no event arrives within the same window. Both tests use fresh order IDs, so they can share containers without truncation.
This suite catches bugs no mock can: a migration that works on H2 but not PostgreSQL, a serializer mismatch between producer and consumer, and an outbox poller that publishes before commit. The price is a few seconds of container start per build, paid once.
Running it in CI
Testcontainers needs a Docker-compatible engine where the tests run. On hosted runners with Docker this works unchanged. On Kubernetes-based runners there is usually no Docker socket, so you need a sidecar engine, a remote Docker host or a cloud container service; do this before you find out from a red build.
Image pulls are the other CI cost. Pin tags, pre-pull images in a cached layer of the runner image, and route pulls through an internal mirror to avoid public registry rate limits. The image name prefix setting rewrites Docker Hub image names onto your mirror:
# ~/.testcontainers.properties
# Reuse: developer laptops only, never in CI images
testcontainers.reuse.enable=true
# Mirror: CI and laptops; prefixes Docker Hub image names
# (env var equivalent: TESTCONTAINERS_HUB_IMAGE_NAME_PREFIX)
hub.image.name.prefix=registry.internal.example/mirror/Keep Ryuk enabled in CI. Disabling it with TESTCONTAINERS_RYUK_DISABLED=true is sometimes needed in locked-down environments, but then a cancelled job leaks containers until something else removes them.
Failure modes
- Hard-coded ports. Passes on one laptop, fails everywhere else. Always read the mapped port.
- Weak readiness. Port-open waits pass before the service works, giving intermittent first-query failures. Use HTTP or log waits.
- Shared state between tests. Singleton containers plus tests that assume an empty table produce order-dependent failures. Truncate or use unique keys.
- Floating tags.
latestsilently upgrades the dependency under test. Pin and upgrade deliberately. - Wrong address side. A container given
localhost:55017cannot reach its peer; inside the network use the alias and internal port. - No Docker in CI. Tests fail at discovery with a message about no valid Docker environment. Provide an engine or a remote host.
- Reuse in CI. Containers persist and accumulate across jobs. Keep reuse to developer machines.
Trade-offs
Testcontainers buys fidelity: the real database, the real broker, real network failure. You pay in start time, a Docker dependency on every build machine, and more complex fixtures. Unit tests with mocks, covered with Mockito, in depth, remain the right tool for logic; containers are for the boundaries where your code meets something that is not yours. A healthy pyramid has many fast unit tests, a moderate number of container-backed integration tests, and few end-to-end tests. The extension model these tests use is explained in JUnit 5, in depth, and the images you ship can be built the way Java containerization describes.
What to do next
- Add the 2.0.5 BOM and the JUnit Jupiter and PostgreSQL modules; on 1.x, rename artifacts and imports first.
- Replace one in-memory database test with a PostgreSQL container pinned to your production major version, running your real migrations.
- Measure suite time, then move shared containers into a JVM-wide singleton base class with per-test cleanup.
- Replace every port-open wait on generic containers with an HTTP or log-message wait.
- Confirm your CI runners have a Docker engine, pre-pull pinned images, and route pulls through a mirror.
- Enable reuse on your own machine only, and never in CI.