Jakarta Persistence (JPA, formerly the Java Persistence API) is a specification: annotations such as @Entity and @Id, the EntityManager interface, the JPQL query language and the rules for how entities move between memory and the database. Hibernate ORM is the most widely used implementation of it, and it adds its own native API and many settings beyond the specification. Hibernate 6 moved to the jakarta.persistence package namespace; Hibernate 7 implements Jakarta Persistence 3.2 and requires Java 17.
Most teams meet Hibernate through Spring Data repositories, which Spring Data covers. This article is about the layer underneath: what Hibernate keeps in memory during a transaction, when and in what order it sends SQL, and the handful of mechanics that explain most production surprises, from duplicate-key errors on a simple replace to imports that run one statement at a time. Each section ends in something you can check in your own code.
The moving parts
Two objects matter. The EntityManagerFactory (Hibernate's SessionFactory) is built once per database at start-up. It holds the mapping metamodel, pre-built SQL for each entity, identifier generators and the second-level cache, and it is thread-safe. The EntityManager (Hibernate's Session) is cheap, not thread-safe and normally lives for exactly one transaction. Inside it is the persistence context, a map from entity type and identifier to the managed Java object plus a snapshot of the values loaded from the database, and the action queue, the list of inserts, updates and deletes waiting to be sent.
Entity states and the persistence context
Every entity instance is in one of four states. Transient: created with new, unknown to Hibernate. Managed: associated with an open persistence context, either because you called persist or because Hibernate loaded it; changes are tracked automatically. Detached: was managed, but the context closed or you called detach or clear; changes are ignored. Removed: scheduled for deletion at flush.
@Transactional
public void lifecycle(EntityManager em) {
Customer c = new Customer("ana@example.com"); // transient: unknown to Hibernate
em.persist(c); // managed: tracked, INSERT queued
c.setName("Ana"); // no call needed: dirty checking at flush
Customer same = em.find(Customer.class, c.getId());
assert same == c; // one Java object per row per context
em.detach(c); // detached: changes are no longer tracked
c.setName("Ana Lima");
Customer managed = em.merge(c); // copies state onto a managed instance
assert managed != c; // keep using the returned object
em.remove(managed); // removed: DELETE queued for flush
}Two guarantees follow from the persistence context. Within one context, a given row is represented by one Java object, so find twice returns the same instance and does not query twice. And merge does not attach the object you pass; it copies its state onto a managed instance and returns that, so code that keeps using the argument after merge silently loses later changes. The persistence context is also a memory cost: every managed entity carries a snapshot, so a transaction that loads 200,000 rows holds two copies of each.
Dirty checking and flush
At flush, Hibernate compares every managed entity with its snapshot, queues an UPDATE for each one that changed, and then executes the action queue. With the default FlushModeType.AUTO, flush happens at commit and also before a query whose results could be affected by pending changes, so a JPQL query sees your unflushed inserts. FlushModeType.COMMIT flushes only at commit, which is faster for read-heavy transactions but lets queries miss your own changes.
The detail that surprises people is order. Hibernate does not replay your calls in sequence. In Hibernate 6 and 7 the action queue executes in a fixed order by type: orphan removals, entity inserts, entity updates, collection changes and finally entity deletes. That ordering keeps foreign keys satisfied in common cases, but it breaks a delete followed by an insert of a row with the same unique key:
// Unique constraint on (sku). Replace an item with a new row that has the same sku.
@Transactional
public void replaceItem(Long oldId, String sku) {
Item old = em.find(Item.class, oldId);
em.remove(old);
em.persist(new Item(sku));
// At commit Hibernate executes INSERT before DELETE: duplicate key on sku.
}
@Transactional
public void replaceItemFixed(Long oldId, String sku) {
Item old = em.find(Item.class, oldId);
em.remove(old);
em.flush(); // send the DELETE now, inside the same transaction
em.persist(new Item(sku));
}An explicit flush() fixes it, because it executes the queued DELETE inside the same transaction before the INSERT is queued. Updating the existing row instead of deleting and inserting is often better still. Flush order is internal behaviour and newer Hibernate development is reworking how actions are scheduled, so add a test for any code that depends on it.
Identifiers and JDBC batching
Hibernate can send many inserts as one JDBC batch, which is one of the largest performance wins available, but only if it knows each entity's identifier before executing the INSERT. With GenerationType.IDENTITY the database assigns the key during the insert, so Hibernate must execute each INSERT immediately to learn it, and insert batching is disabled for that entity. With GenerationType.SEQUENCE and an allocation size greater than 1, Hibernate fetches a block of identifiers with one sequence call and assigns them in memory; the JPA default allocation size is 50, and the database sequence's increment must match it.
@Entity
public class OrderLine {
@Id
@GeneratedValue(strategy = GenerationType.SEQUENCE, generator = "order_line_seq")
@SequenceGenerator(name = "order_line_seq", sequenceName = "order_line_seq",
allocationSize = 50) // sequence must INCREMENT BY 50 to match
private Long id;
@ManyToOne(fetch = FetchType.LAZY) // JPA default for ToOne is EAGER
private PurchaseOrder order;
@Version
private long version;
}Batching itself is off until you set a batch size. These properties, whose names were checked against the Hibernate source, are a sensible starting point:
hibernate.jdbc.batch_size=50
hibernate.order_inserts=true
hibernate.order_updates=true
hibernate.default_batch_fetch_size=32
hibernate.query.fail_on_pagination_over_collection_fetch=true
hibernate.generate_statistics=true # in staging and load testsorder_inserts and order_updates group statements by table so that alternating parent and child inserts can still batch. Client-generated UUIDs avoid the round trip too, but random UUIDs as primary keys fragment B-tree indexes; prefer time-ordered UUID versions if you choose them.
Lazy loading and proxies
JPA's defaults are a trap: @ManyToOne and @OneToOne are EAGER, while @OneToMany and @ManyToMany are LAZY. Eager to-one associations load whether or not you need them, often as extra queries, so set fetch = FetchType.LAZY on every to-one mapping and fetch explicitly where a use case needs the data.
A lazy to-one is a proxy: a generated subclass whose identifier is set and whose other fields load on first access. A lazy collection is a Hibernate collection wrapper that loads on first iteration. Both need an open persistence context, so touching one after the transaction ends throws LazyInitializationException. Spring Boot's spring.jpa.open-in-view defaults to true and keeps the context open through view rendering, which hides that exception by issuing queries from the web layer; set it to false and fetch what each endpoint needs.
To fetch on purpose, use a JOIN FETCH in JPQL or an entity graph for a specific query, and set hibernate.default_batch_fetch_size so that remaining lazy loads fetch up to that many proxies or collections per query using an IN list. One hazard: paginating a query that join-fetches a collection cannot be done in SQL, so Hibernate fetches all rows and paginates in memory with a warning. hibernate.query.fail_on_pagination_over_collection_fetch=true turns that warning into an exception you will see in tests.
Worked example: importing 100,000 order lines
A naive import that persists 100,000 entities in one transaction, with IDENTITY keys and default settings, sends 100,000 separate INSERT statements and holds 100,000 entities with snapshots in memory until commit; dirty checking at flush then walks all of them. The fixed version uses a pooled sequence, a batch size of 50, and flushes and clears every 50 entities:
@Transactional
public void importLines(Iterator<LineDto> rows, long orderId) {
PurchaseOrder order = em.getReference(PurchaseOrder.class, orderId); // proxy, no SELECT
int i = 0;
while (rows.hasNext()) {
LineDto r = rows.next();
em.persist(new OrderLine(order, r.sku(), r.qty()));
if (++i % 50 == 0) { // match hibernate.jdbc.batch_size
em.flush(); // send one batch of 50 INSERTs
em.clear(); // drop managed entities and snapshots
order = em.getReference(PurchaseOrder.class, orderId); // re-acquire after clear
}
}
}Now each flush sends one JDBC batch, clear keeps the persistence context at no more than 50 entities, and getReference gives a proxy for the parent without a SELECT. Clearing detaches everything, including the parent proxy, which is why it is re-acquired. Verify rather than assume: with hibernate.generate_statistics enabled, Hibernate's Statistics report JDBC statement and batch counts, and the database's statement log should show batched inserts. For pure write-through work with no need for dirty checking or cascades, Hibernate's StatelessSession skips the persistence context entirely. If throughput is still low, record the run with Java Flight Recorder to see whether time goes to JDBC waits, flush or garbage collection. Test the job against the real database with Testcontainers, because batching behaviour differs between drivers.
Second-level and query caches
The persistence context is a first-level cache scoped to one transaction. The optional second-level cache is shared across sessions and stores entity data by identifier, through a provider such as a JCache implementation. It is enabled with hibernate.cache.use_second_level_cache and opted into per entity with @Cacheable plus a concurrency strategy such as read-only or read-write.
It pays off for data that is read far more than it is written and looked up by identifier: reference tables, product catalogues, configuration. It does not help queries by other columns unless you also enable the query cache (hibernate.cache.use_query_cache), which stores result identifiers and is invalidated whenever any table in the query changes, so it rarely helps write-heavy tables. Bulk JPQL updates and native SQL bypass entity-level tracking, and changes made by other applications are invisible to the cache. In a multi-node deployment, a local cache per node serves stale data unless the provider replicates or invalidates across nodes.
Concurrency: optimistic and pessimistic locking
Add a @Version field to every entity that users edit concurrently. Hibernate includes the version in the WHERE clause of each update and increments it; if another transaction changed the row first, the update matches zero rows and Hibernate throws an optimistic lock exception. That turns lost updates into a visible conflict that you resolve by retrying the whole unit of work or by showing the user the newer data. For short, hot critical sections such as decrementing stock, LockModeType.PESSIMISTIC_WRITE issues a locking read; keep those transactions short and access rows in a consistent order to avoid deadlocks.
Bytecode enhancement
By default Hibernate detects changes by comparing snapshots and implements laziness with proxies. With build-time bytecode enhancement through the Hibernate Maven or Gradle plugin, entities record their own changes as setters run, so flush does not need to compare every field, and individual basic attributes such as a large text column can be lazy. The trade-off is a build step and behaviour that differs subtly from proxy-based laziness, so enable it deliberately and run your persistence tests against it.
Failure modes and trade-offs
- Delete then insert on a unique key fails at flush because deletes execute last.
- IDENTITY keys silently disable insert batching.
- Eager to-one defaults load object graphs nobody asked for.
- Open session in view moves queries into rendering and hides missing fetch plans.
- Using the argument after merge loses changes, because only the returned instance is managed.
- Large transactions without clear exhaust memory and make every flush slower.
- Entity equals and hashCode based on a generated identifier change after persist and break hash-based collections; use a natural key or a stable identity.
The overall trade-off: JPA and Hibernate remove a great deal of mapping and change-tracking code and handle object graphs well, in return for implicit SQL whose timing and shape you must learn to predict. For reporting queries, heavy bulk processing or SQL you want to control exactly, use JDBC, jOOQ or native queries alongside it; many healthy systems use Hibernate for transactional writes and plain SQL for reads that do not map to entities.
What to do next
- Enable SQL logging and
hibernate.generate_statisticsin a test environment and count the statements for your three busiest endpoints. - Set
fetch = FetchType.LAZYon every to-one association, then add JOIN FETCH or entity graphs where tests show they are needed. - Set
spring.jpa.open-in-view=falseif you use Spring Boot, and fix every LazyInitializationException it reveals. - Switch high-volume entities from IDENTITY to a pooled SEQUENCE, and set batch size, order_inserts and order_updates.
- Turn on
hibernate.query.fail_on_pagination_over_collection_fetch. - Add
@Versionto user-editable entities and decide the retry policy for conflicts. - Search for remove followed by persist on the same unique key and add a flush or rewrite as an update.