A lock answers the question "may I be the one thread in here?" A java.util.concurrent.Semaphore answers a different one: "may I be one of at most N?" It holds a number of permits. acquire() takes one, blocking while none are free; release() gives one back and wakes a waiter. That makes it the natural tool for capping concurrency: at most 20 calls to a fragile downstream service, at most 4 simultaneous large file uploads, at most 200 requests in flight in a server before it starts shedding load.
Semaphores are older than Java (Dijkstra introduced them in the 1960s), and the Java version is deliberately minimal. That minimalism is where the bugs come from. A semaphore records no owner, so the runtime cannot tell a legitimate release from a double release, and it cannot detect a permit that leaked on an error path. This article explains how Semaphore works inside, walks through its API, shows the try/finally mistake that silently inflates the permit count, builds a worked example with virtual threads and load shedding, and ends with failure modes, monitoring and alternatives.
The mental model: tokens, not owners
Think of a semaphore as a bowl of tokens. To enter the guarded region a thread takes a token; on the way out it puts one back. If the bowl is empty, arriving threads wait in line. Three properties follow from that picture and from the Javadoc:
- No ownership. Any thread may call
release(), whether or not it acquired. That enables hand-off patterns where one thread acquires and another releases, and it means mistakes are not caught. - No upper bound. The initial count is not a cap. Each extra
release()adds a permit, so an unmatched release permanently raises the limit. The initial value may even be negative, in which case releases must happen before anyone can acquire. - Visibility. Actions in a thread before a
release()happen-before actions after a successful acquire in another thread. A semaphore is a real memory-model synchronisation point, not just a counter.
The lack of ownership is occasionally exactly what you want. A semaphore created with zero permits works as a one-way signal: a consumer calls acquire() and parks until a producer calls release(), and the producer's writes before the release are visible to the consumer afterwards. Each release lets exactly one waiter through, which distinguishes it from a CountDownLatch that opens for everyone and cannot be reset.
The mental model: tokens, not owners
Think of a semaphore as a bowl of tokens. To enter the guarded region a thread takes a token; on the way out it puts one back. If the bowl is empty, arriving threads wait in line. Three properties follow from that picture and from the Javadoc:
- No ownership. Any thread may call
release(), whether or not it acquired. That enables hand-off patterns where one thread acquires and another releases, and it means mistakes are not caught. - No upper bound. The initial count is not a cap. Each extra
release()adds a permit, so an unmatched release permanently raises the limit. The initial value may even be negative, in which case releases must happen before anyone can acquire. - Visibility. Actions in a thread before a
release()happen-before actions after a successful acquire in another thread. A semaphore is a real memory-model synchronisation point, not just a counter.
The lack of ownership is occasionally exactly what you want. A semaphore created with zero permits works as a one-way signal: a consumer calls acquire() and parks until a producer calls release(), and the producer's writes before the release are visible to the consumer afterwards. Each release lets exactly one waiter through, which distinguishes it from a CountDownLatch that opens for everyone and cannot be reset.
Inside Semaphore: AQS shared mode
Like ReentrantLock and CountDownLatch, Semaphore is built on AbstractQueuedSynchronizer (AQS). The AQS state integer holds the number of available permits, and acquisition uses the shared mode, because many threads can hold permits at once. The non-fair acquire path is a compare-and-swap loop, simplified here from the JDK source:
// Non-fair: try to take permits immediately
int nonfairTryAcquireShared(int acquires) {
for (;;) {
int available = getState();
int remaining = available - acquires;
if (remaining < 0 || compareAndSetState(available, remaining))
return remaining; // negative => caller must enqueue and park
}
}
// Fair: same, but yield to anyone already waiting
int fairTryAcquireShared(int acquires) {
for (;;) {
if (hasQueuedPredecessors()) return -1;
int available = getState();
int remaining = available - acquires;
if (remaining < 0 || compareAndSetState(available, remaining))
return remaining;
}
}
boolean tryReleaseShared(int releases) {
for (;;) {
int current = getState();
int next = current + releases;
if (next < current) throw new Error("Maximum permit count exceeded");
if (compareAndSetState(current, next)) return true; // AQS then unparks waiters
}
}When a permit is free, acquiring is a single successful CAS with no parking and no kernel involvement. When none is free, AQS appends the thread to its wait queue and parks it. A release adds to the state and unparks the head of the queue, which retries the acquire. Because the mode is shared, a release that frees several permits propagates so that more than one waiter can wake.
The API you will actually use
The table covers the methods you will actually use. Two details deserve emphasis. First, untimed tryAcquire() ignores fairness by design; if you want a non-blocking attempt that respects the queue, call tryAcquire(0, TimeUnit.SECONDS). Second, availablePermits() is a snapshot that may be stale by the time you read it, so never write if (sem.availablePermits() > 0) sem.acquire(); that is the same check-then-act race that tryAcquire exists to avoid.
| Method | Behaviour |
|---|---|
acquire() / acquire(n) | Block until permits are available; throws InterruptedException |
acquireUninterruptibly() | Block, ignoring interrupts (interrupt status is restored afterwards) |
tryAcquire() | Take a permit only if one is free right now; barges even on a fair semaphore |
tryAcquire(timeout, unit) | Wait up to the timeout; honours fairness; returns false on timeout |
release() / release(n) | Add permits and wake waiters; never blocks, never checks ownership |
availablePermits() | Current count, for monitoring only |
drainPermits() | Take all currently available permits and return how many |
getQueueLength() | Estimate of how many threads are waiting |
reducePermits(n) | Protected: shrink the count, for subclasses that resize a limit |
Acquire, try, finally: the placement bug
The canonical usage pairs every acquire with exactly one release in a finally block. The position of the acquire relative to the try matters more than it looks:
// WRONG: if acquire() is interrupted, finally still releases a permit we never took
try {
sem.acquire();
callDownstream();
} finally {
sem.release(); // inflates the count by one on every interrupted wait
}
// RIGHT: acquire outside, release inside
sem.acquire();
try {
callDownstream();
} finally {
sem.release();
}The wrong version works in every test that does not interrupt threads, then quietly adds a permit each time a waiting thread is cancelled, for example when an executor is shut down or a request times out. After enough cancellations your limit of 20 has become 35 and the downstream service is overloaded with nothing in the logs to say why. A small wrapper makes the pairing structural and idempotent:
final class Permit implements AutoCloseable {
private final Semaphore sem;
private final AtomicBoolean closed = new AtomicBoolean();
private Permit(Semaphore sem) { this.sem = sem; }
static Permit acquire(Semaphore sem) throws InterruptedException {
sem.acquire(); // throws before a Permit exists: nothing to release
return new Permit(sem);
}
@Override public void close() {
if (closed.compareAndSet(false, true)) sem.release(); // double close cannot inflate
}
}
try (Permit p = Permit.acquire(downstreamLimit)) {
callDownstream();
}
Worked example: protecting a dependency with virtual threads
A worked example. An order service fans out to an inventory API that its owners say can handle about 200 requests per second, with a typical latency of 50 ms. By Little's law, concurrency equals throughput times latency, so 200 x 0.05 = 10 requests in flight on average. A semaphore with around 10 to 15 permits caps the pressure on that dependency no matter how many requests arrive, and a timed acquire turns overload into a fast, explicit rejection rather than an ever-growing queue.
With virtual threads this pattern becomes the standard one. JEP 444 explicitly advises against pooling virtual threads to limit concurrency and recommends constructs such as semaphores instead: create a virtual thread per task, and guard the scarce resource rather than the threads.
private final Semaphore inventoryLimit = new Semaphore(12);
List<StockLevel> checkStock(List<String> skus) throws Exception {
try (var exec = Executors.newVirtualThreadPerTaskExecutor()) {
List<Future<StockLevel>> futures = skus.stream()
.map(sku -> exec.submit(() -> {
if (!inventoryLimit.tryAcquire(200, TimeUnit.MILLISECONDS)) {
throw new OverloadedException("inventory saturated"); // map to 503 / fallback
}
try {
return inventoryClient.get(sku);
} finally {
inventoryLimit.release();
}
}))
.toList();
List<StockLevel> out = new ArrayList<>();
for (Future<StockLevel> f : futures) out.add(f.get());
return out;
}
}Thousands of virtual threads may be created, but only 12 talk to the inventory API at once. The 200 ms timeout is the load-shedding knob: callers that cannot get a slot quickly fail fast, and the caller decides whether to retry later, degrade, or return an error. Choose the timeout from your own latency budget, not from the dependency's.
Remember what a semaphore does not do: it limits concurrency, not rate. Twelve permits with 50 ms calls allow about 240 calls per second; if latency drops to 5 ms, the same twelve permits allow about 2,400. If the contract is "requests per second", use a rate limiter, possibly together with the semaphore.
Fairness, starvation and weighted permits
The default semaphore is non-fair: a thread that arrives just as a permit is released may take it ahead of threads that have been waiting. That is faster, because a running thread can proceed without the hand-off cost of waking a parked one, and for pools of interchangeable work it is usually fine. new Semaphore(n, true) makes acquisition first-in, first-out, at the cost of throughput.
Fairness matters most with multi-permit acquires. If tasks take weighted shares, say acquire(8) for a large export and acquire(1) for small ones out of 10 permits, a non-fair semaphore can let a steady stream of small tasks starve the large one indefinitely, which the Javadoc warns about. A fair semaphore prevents that starvation, but now the large task at the head of the queue blocks the small tasks behind it until eight permits are free. Choose which behaviour your workload can live with, and if neither is acceptable, use separate semaphores for separate classes of work.
Failure modes and monitoring
| Failure | Symptom | Fix |
|---|---|---|
| Leaked permit on an exception or early return | Throughput slowly falls; eventually all callers block | try/finally or the AutoCloseable wrapper |
| Release in finally after an interrupted acquire | Effective limit creeps upward; downstream overloaded | Acquire before the try block |
| Double release | Limit higher than configured | Idempotent release wrapper |
| Acquiring permits one at a time for a multi-permit task | Two tasks each hold part of what they need: deadlock | acquire(n) atomically |
| Two semaphores taken in different orders | Classic lock-order deadlock | Fixed acquisition order, or one semaphore |
| Untimed acquire on a request path | Requests pile up, latency explodes under overload | tryAcquire with a timeout and a rejection path |
Monitor the semaphore like any other capacity. Export availablePermits() and getQueueLength() as gauges, and alert when the queue is persistently non-empty or when available permits sit at zero with low traffic, which is the signature of a leak. In a thread dump, waiting threads show as parked on java.util.concurrent.Semaphore$NonfairSync (or FairSync), which tells you immediately which limit they are queued behind.
Test the unhappy paths explicitly. Write a test that starts more tasks than permits, interrupts or cancels some of them while they wait, lets others throw from inside the guarded region, and then asserts that availablePermits() is back to its initial value once everything has finished. That one assertion catches leaks, double releases and the finally placement bug together.
When to use something else
A semaphore is the right tool when you need to cap how many things happen at once and the "thing" is not an object you hand out. If it is an object, such as a pool of connections or buffers, a BlockingQueue of those objects already blocks when empty, and adding a semaphore on top only creates two counts that can drift apart. For plain mutual exclusion, use ReentrantLock or synchronized: a Semaphore(1) is not reentrant and has no owner, so a thread that re-enters deadlocks itself, and a stray release from another thread is not detected. For one-shot or phased coordination, use CountDownLatch or Phaser. For limits across a fleet, a JVM semaphore is per process; with 20 instances and 12 permits each you have allowed 240 concurrent calls, so divide the budget or use a distributed limiter. Resilience libraries such as Resilience4j package the same idea as a semaphore-based bulkhead with metrics included.
What to do next
- Find every
acquire()in your code and check that it sits before itstryand that the matchingrelease()is in thefinally. - Wrap permits in an AutoCloseable so leaks and double releases become impossible by construction.
- Size each limit with Little's law from the dependency's capacity and latency, then load-test it.
- Replace untimed acquires on request paths with
tryAcquire(timeout, unit)and a clear rejection path. - Export available permits and queue length as metrics and alert on a stuck-at-zero pattern.
- If you run virtual threads, limit concurrency with semaphores around resources, not with thread pools.
- Read on: virtual threads, CountDownLatch and Java concurrency fundamentals.