Spring Data lets you declare an interface such as interface OrderRepository extends JpaRepository<Order, Long> and get a working data access layer without writing an implementation. That convenience is real, and it is also where most Spring persistence performance problems start. The repository hides the queries it runs, the transactions it opens and the entity state it tracks. A page that runs one query in development runs forty in production, and nobody wrote any of them.
This article explains what Spring Data actually does, so you can keep the convenience and control the SQL. It covers Spring Data JPA, the most widely used module, and was checked against the Spring Data JPA reference documentation for the current 4.1 line; the APIs shown also exist in 3.x. How Spring Boot wires the DataSource and auto-configuration is covered in Spring Boot, in depth.
What Spring Data is
Spring Data is an umbrella project. Spring Data Commons defines the repository abstraction (Repository, CrudRepository, ListCrudRepository, PagingAndSortingRepository), query derivation from method names, projections, paging types and auditing. Store modules implement it for a specific store: JPA, JDBC, R2DBC, MongoDB, Redis, Cassandra and others. The interfaces look alike across modules, but the semantics underneath do not. JPA brings a persistence context with lazy loading and dirty checking, while Spring Data JDBC has neither. Code that works on one module is not portable to another just because the interface compiles.
How an interface becomes a bean
At startup, Spring Boot's auto-configuration (or @EnableJpaRepositories) scans for interfaces that extend Repository. For each one, a JpaRepositoryFactory builds a JDK dynamic proxy. Inherited CRUD methods are forwarded to a SimpleJpaRepository instance. Each other method is resolved once using the default lookup strategy, CREATE_IF_NOT_FOUND: use a declared @Query if present, otherwise a named query, otherwise derive a query from the method name. Transaction and exception-translation interceptors wrap the proxy, so JPA exceptions surface as Spring's DataAccessException hierarchy.
Two consequences follow. Derived query names are parsed at startup, so a method that refers to a property that does not exist fails application startup rather than the first request. And because everything goes through a proxy, a repository method called with no transaction open runs in its own short transaction, with the persistence context closed when it returns. Custom code that the interface cannot express goes in a fragment: declare an interface such as OrderRepositoryCustom, implement it in a class named OrderRepositoryCustomImpl, and extend both from the repository.
Query methods: derived, declared and native
public interface OrderRepository extends JpaRepository<Order, Long> {
// Derived: parsed from the name at startup.
List<Order> findByCustomerIdAndStatusOrderByCreatedAtDesc(long customerId, Status status);
boolean existsByExternalRef(String externalRef);
long countByStatus(Status status);
// Declared JPQL: use once the name stops being readable.
@Query("select o from Order o where o.customer.id = :cid and o.total > :min")
List<Order> bigOrders(@Param("cid") long customerId, @Param("min") BigDecimal min);
// Native SQL: database-specific features, no entity-level portability.
@Query(value = "select * from orders where created_at > now() - interval '1 day'", nativeQuery = true)
List<Order> lastDay();
// Dynamic result size without Pageable.
List<Order> findByStatus(Status status, Sort sort, Limit limit);
}Derived queries are excellent for one or two predicates and unreadable beyond three. Switch to @Query at that point, or to Specifications or Querydsl for filters assembled at runtime. Whatever the source, write a test that runs each query against a real database of the same kind as production, because H2 accepts things Postgres rejects and the reverse.
Projections: fetch what you show
Returning entities is the default, and it is wasteful for read paths. An entity load fetches every mapped column, puts a managed copy in the persistence context and keeps a snapshot for dirty checking. Spring Data projections return only what you need. A closed interface projection, whose getters match entity properties, lets Spring Data JPA select only those columns. An open projection, with a @Value SpEL expression, needs the whole entity, so it saves nothing on the query. DTO projections using a class or record return plain objects with no persistence context involvement; with @Query JPQL, write the constructor expression yourself.
public interface OrderSummary { // closed projection: selects four columns
Long getId();
Instant getCreatedAt();
Status getStatus();
BigDecimal getTotal();
}
public record OrderRow(Long id, Instant createdAt, BigDecimal total) {}
Window<OrderSummary> findFirst20ByCustomerIdOrderByCreatedAtDescIdDesc(long customerId, ScrollPosition position);
@Query("select new com.shop.OrderRow(o.id, o.createdAt, o.total) from Order o where o.status = :s")
List<OrderRow> rows(@Param("s") Status status);
<T> List<T> findByCustomerId(long customerId, Class<T> type); // caller picks the shape
Page, Slice and Window
Spring Data offers three ways to read a large result in pieces, and they cost very different amounts. A Page knows the total, so it runs an extra COUNT query on every request, and its offset makes the database walk and discard every earlier row. A Slice fetches page size plus one row to learn whether a next slice exists, with no count query, but still uses an offset. A Window with a ScrollPosition supports keyset scrolling: the next query filters on the sort keys of the last row returned, so page 500 costs the same as page 1.
ScrollPosition position = ScrollPosition.keyset(); // start at the beginning
Window<OrderSummary> window;
do {
window = orders.findFirst20ByCustomerIdOrderByCreatedAtDescIdDesc(customerId, position);
window.forEach(this::render);
if (!window.isEmpty()) {
position = window.positionAt(window.size() - 1); // keys of the last row
}
} while (window.hasNext());Keyset scrolling needs a sort that is unique and stable, which is why the example sorts by creation time and then id. The sort keys must also be returned in the results, which is why the projection includes them. A Window only moves forward, and for an HTTP API you serialise the last row's keys into an opaque cursor rather than exposing raw values. Use Page only when the UI truly shows a total, and consider caching that count.
save(), entity state and dirty checking
The save method persists or merges. According to the reference documentation, Spring Data JPA first looks for a version property of non-primitive type and treats the entity as new if it is null; without one, it treats the entity as new if the identifier is null. New entities go to persist, others to merge. A primitive version cannot signal newness, because JPA treats 0 as the first saved version.
This creates a common hidden cost. Entities with application-assigned identifiers, such as a UUID set in the constructor, have a non-null id from the start, so save calls merge, and merge first runs a SELECT to look for an existing row before the INSERT. Fix it with a nullable @Version Long version field, which also gives you optimistic locking, or by implementing Persistable with your own isNew().
Inside a transaction, entities you loaded are managed. Change a field and the change is written at flush or commit by dirty checking, with no save call needed. Calling save on a managed entity is harmless but misleading in code review. Two related methods differ sharply: deleteAll() loads every entity and deletes them one by one, running lifecycle callbacks, while deleteAllInBatch() issues a single DELETE statement. getReferenceById returns a lazy proxy without a query, which is ideal for setting a foreign key.
Transactions and bulk updates
Inherited CRUD methods carry their own transaction settings from SimpleJpaRepository: reads run with readOnly = true and other methods with a plain @Transactional. With Hibernate, Spring uses the read-only flag to skip flushing and dirty-checking snapshots, which saves real memory on large reads. Put @Transactional on service methods that span several repository calls, so they share one transaction and one persistence context; otherwise each call commits on its own and a failure halfway leaves partial writes. Isolation level choices are covered in Database isolation levels.
Bulk JPQL updates need @Modifying, and they bypass the persistence context. Entities already loaded keep their old values. Set @Modifying(flushAutomatically = true, clearAutomatically = true) so pending changes are written first and stale entities are evicted afterwards. Finally, Spring Boot enables open-session-in-view by default and logs a warning about it at startup. It keeps the persistence context open while the view renders, so lazy loads in controllers and serializers run extra queries outside your service transaction and hold a pooled connection longer. Set spring.jpa.open-in-view=false and fetch what each endpoint needs explicitly. Pool sizing is covered in Connection pooling architecture.
The N+1 problem and fetch planning
Load 20 orders, then touch order.getLines() on each, and a lazy association runs one query per order: 1 + 20. Spring Data offers several fixes. @EntityGraph(attributePaths = "lines") on a repository method fetches the association in the same query. A join fetch in @Query does the same explicitly. Hibernate's @BatchSize or hibernate.default_batch_fetch_size turns N lazy loads into a few IN queries.
Do not combine a collection fetch join with pagination. The database cannot apply LIMIT to parent rows when the result is multiplied by children, so Hibernate fetches everything and pages in memory, logging a warning. Set hibernate.query.fail_on_pagination_over_collection_fetch=true to turn that into an error, and page the parent ids first, then fetch their children in a second query.
Worked example: the order history endpoint
A customer order history page shows 20 orders with their line counts. The first version used Page<Order> findByCustomerId(long id, Pageable p) and a JSON serializer that walked lines. With SQL logging and hibernate.generate_statistics turned on in an integration test, one request ran 22 statements: a count, the page and 20 line queries. Deep pages got slower in proportion to the offset. These figures come from the illustrative test, not a benchmark.
The rewrite took four steps. The endpoint switched to the derived Window<OrderSummary> method shown earlier, removing the count and the offset. Line counts come from a second declared query keyed by the ids in the window, select o.id as id, size(o.lines) as lines from Order o where o.id in :ids, which replaces the 20 lazy loads. Open-in-view was disabled, so any remaining lazy access fails loudly in tests instead of silently querying. And the API returns an opaque cursor built from the last row's creation time and id. The same request now runs two statements whatever the page depth. A test asserts the statement count, so a future change that reintroduces N+1 fails the build.
Auditing and optimistic locking
Enable auditing with @EnableJpaAuditing, add @EntityListeners(AuditingEntityListener.class) to entities, and annotate fields with @CreatedDate and @LastModifiedDate (plus @CreatedBy with an AuditorAware bean). For concurrency, a @Version field makes every update include where version = ?; a lost race surfaces as Spring's ObjectOptimisticLockingFailureException, which the caller can retry after re-reading or map to an HTTP 409 or 412. What a transaction does and does not guarantee is covered in ACID transactions, in depth.
@Entity
@EntityListeners(AuditingEntityListener.class)
public class Order {
@Id private UUID id = UUID.randomUUID(); // assigned id ...
@Version private Long version; // ... so a nullable version marks it new
@CreatedDate private Instant createdAt;
@LastModifiedDate private Instant updatedAt;
// fields, getters, domain methods
}
Trade-offs and alternatives
Spring Data JPA pays off for aggregate-shaped domain models with moderate query complexity, where the persistence context and dirty checking save real code. It fits less well for reporting, heavy SQL and bulk data movement, where the persistence context only adds memory and surprises. Spring Data JDBC keeps the repository model but drops lazy loading, dirty checking and the session: an aggregate is loaded whole and saved whole, which is simpler to reason about and needs more deliberate aggregate boundaries. For SQL-heavy code, plain JdbcClient or a typed SQL library is often clearer than stretching derived queries. Mixing them is normal; use repositories where they fit and SQL where it is the clearer language.
Failure modes
- Hidden N+1. Lazy associations touched in loops or serializers. Assert statement counts in tests.
- Merge on every insert. Assigned ids without a nullable version. Add
@Version Longor implement Persistable. - Expensive counts. Page on large tables. Use Slice or Window unless a total is required.
- Stale entities after bulk updates. Use clearAutomatically or reload.
- Accidental per-call transactions. Several repository calls without a service transaction commit separately.
- Open-in-view. Lazy queries during rendering, connections held longer. Turn it off.
What to do next
- Enable SQL logging and Hibernate statistics in integration tests and record the statement count of your top ten endpoints.
- Set
spring.jpa.open-in-view=falseand fix the lazy-loading failures it reveals with entity graphs or projections. - Replace entity-returning read paths with closed interface or record projections.
- Move deep or infinite lists from Page to Window keyset scrolling with a unique sort.
- Add a nullable
@Versionto entities with assigned ids, and map optimistic-lock failures to retries or 409 responses. - Put
@Transactionalon service methods that combine repository calls, and addclearAutomaticallyto every@Modifyingquery.