Every asynchronous API eventually hands you an object that stands for a value that does not exist yet. Java calls it a CompletableFuture, C++ a std::future, JavaScript a Promise, Python has two different Future classes, and Rust has a trait. They look alike and behave differently in exactly the places that cause production bugs: which thread runs your callback, whether work starts before anyone asks for it, what cancellation really stops, and what happens to an error nobody looks at.
This article builds a future and promise from scratch in about sixty lines, so the mechanism is concrete, then walks the five design decisions every implementation makes and shows how each major language answers them, with code. A worked fan-out with deadlines and a catalogue of failure modes finish it. The conceptual architecture of the write-once cell is covered in futures and promises architecture; this page is about building, choosing and debugging.
The model in one picture
Build one in sixty lines
A future needs four things: a state that moves exactly once from pending to settled, a lock so concurrent settles and registrations cannot interleave, a list of callbacks waiting for the result, and a way to block for the value when you must. Here is a complete, thread-safe version in Python using only the standard library. One object plays both roles; in a real library you would hand callers a read-only view.
import threading
PENDING, DONE, FAILED = "pending", "done", "failed"
class Promise:
def __init__(self):
self._cv = threading.Condition()
self._state, self._value, self._callbacks = PENDING, None, []
def _settle(self, state, value):
with self._cv:
if self._state is not PENDING:
return False # write-once: later settles lose
self._state, self._value = state, value
callbacks, self._callbacks = self._callbacks, []
self._cv.notify_all() # wake blocked result() callers
for cb in callbacks: # never run user code under the lock
cb(self)
return True
def resolve(self, value): return self._settle(DONE, value)
def reject(self, exc): return self._settle(FAILED, exc)
def on_complete(self, cb):
with self._cv:
if self._state is PENDING:
self._callbacks.append(cb)
return
cb(self) # already settled: run now, on this thread
def result(self, timeout=None):
with self._cv:
if not self._cv.wait_for(lambda: self._state is not PENDING, timeout):
raise TimeoutError("still pending")
if self._state is FAILED:
raise self._value
return self._value
def then(self, fn):
out = Promise()
def step(src):
if src._state is FAILED:
out.reject(src._value) # errors skip transformations
return
try:
v = fn(src._value)
except Exception as e:
out.reject(e)
return
if isinstance(v, Promise): # fn was itself async: flatten
v.on_complete(lambda q: out._settle(q._state, q._value))
else:
out.resolve(v)
self.on_complete(step)
return outThree details carry the design. The check-and-append in on_complete happens under the same lock as the transition in _settle, so a callback registered at the instant of completion is either queued and run by the settler or run immediately by the registrar, never lost and never run twice. Callbacks run outside the lock, so a callback that registers another callback cannot deadlock. And then flattens a returned promise, which is the difference between Java's thenApply and thenCompose collapsed into one method, as JavaScript does.
Composition is more callbacks
Combinators are just more callbacks. Waiting for all of a list is a countdown that fails fast on the first error:
def all_of(promises):
out, results = Promise(), [None] * len(promises)
remaining, lock = [len(promises)], threading.Lock()
if not promises:
out.resolve([])
for i, pr in enumerate(promises):
def done(src, i=i):
if src._state is FAILED:
out.reject(src._value) # first failure wins; the rest are ignored
return
results[i] = src._value
with lock:
remaining[0] -= 1
last = remaining[0] == 0
if last:
out.resolve(results)
pr.on_complete(done)
return outNotice what fail-fast does not do: the other inputs keep running. Nothing in a future can stop the work that will complete it unless the producer checks for cancellation. That single fact explains most cancellation surprises below.
Five decisions every implementation makes
Every implementation answers the same five questions, and the answers are what you must know about the one in front of you.
| Question | Java CompletableFuture | C++ std::future | JavaScript Promise | Python | Rust Future |
|---|---|---|---|---|---|
| Does work start before anyone waits? | Yes, eager | Yes with std::async(launch::async) | Yes, executor runs at construction | Submitted work and asyncio tasks yes; a bare coroutine not until awaited | No, lazy until polled |
| How many readers? | Many | One; shared_future for many | Many | Many | One owner |
| Where do callbacks run? | Completing thread, or an executor for *Async | No callbacks in the standard | Always later, as microtasks | concurrent: completing thread; asyncio: the loop | No callbacks; the executor polls |
| What does cancel stop? | Only the future, not the task | No cancel | No cancel in the type | concurrent: only if not started | Dropping stops the work |
| Unobserved error? | Silently dropped | Rethrown by get() | Unhandled rejection event | asyncio logs at garbage collection | Must be handled to compile cleanly |
Java: eager, and cancel does not stop the task
Java's CompletableFuture is eager and multi-reader. Its non-async methods run the callback on whichever thread completed the previous stage, or on the caller if the stage is already done; the *Async variants dispatch to the common pool or an executor you pass. Its cancel(mayInterruptIfRunning) ignores the flag: the Javadoc states that interrupts are not used, so cancel completes the future with a CancellationException and the supplier keeps running. The same holds for orTimeout, which completes the future exceptionally after the deadline but leaves the task occupying its thread. If work must actually stop, the task has to check a flag or the future's state itself. The full API is in CompletableFuture in Java.
C++: one reader, broken promises, blocking destructors
C++ splits the two ends into separate types and keeps them minimal:
#include <future>
#include <thread>
std::promise<int> p;
std::future<int> f = p.get_future();
std::thread producer([pr = std::move(p)]() mutable {
try {
pr.set_value(compute());
} catch (...) {
pr.set_exception(std::current_exception()); // errors travel as values
}
});
int v = f.get(); // blocks; rethrows a stored exception; afterwards f.valid() == false
producer.join();
auto shared = std::async(std::launch::async, compute).share(); // many readersThree rules cause most C++ bugs. get() may be called once; a second call is undefined behaviour unless you use std::shared_future. If a promise is destroyed before it is satisfied, waiting readers receive a std::future_error with the broken_promise code, so an early return in a producer surfaces as an error rather than a hang. And the future returned by std::async blocks in its destructor until the task finishes, so discarding it turns intended parallelism into sequential code. The standard has no then; continuation libraries and executors fill that gap.
JavaScript: always asynchronous
JavaScript guarantees that reactions never run synchronously, even on a promise that is already settled. They are queued as microtasks and run after the current script finishes, before timers:
console.log("1 sync");
Promise.resolve().then(() => console.log("3 microtask"));
setTimeout(() => console.log("4 timer"), 0);
console.log("2 sync");
// 1, 2, 3, 4 - always
// fail fast versus collect everything
const all = await Promise.all([a(), b(), c()]); // rejects on the first failure
const settled = await Promise.allSettled([a(), b(), c()]);
const ok = settled.filter(r => r.status === "fulfilled").map(r => r.value);That rule removes a whole class of reentrancy bugs that Java's same-thread callbacks allow, at the cost of a scheduling hop per step. A rejection with no handler raises an unhandled-rejection event; since Node.js 15 the default is to treat it as an uncaught error and exit, so a forgotten catch in a server can take the process down. Cancellation is a separate mechanism, AbortController, passed into the API that does the work. The async/await patterns article covers the syntax built on top.
Python: two futures that must not be mixed
Python has two incompatible future types. concurrent.futures.Future is thread-safe; add_done_callback runs on the thread that completes it, or immediately on the caller if it is already done, and cancel() succeeds only while the task is still waiting in the queue. asyncio.Future belongs to one event loop and is not thread-safe; calling its methods from another thread corrupts the loop's state quietly. Bridge explicitly:
import asyncio, concurrent.futures
pool = concurrent.futures.ThreadPoolExecutor(max_workers=8)
async def handler(key):
loop = asyncio.get_running_loop()
return await loop.run_in_executor(pool, legacy_blocking_lookup, key)
# from an ordinary thread: schedule onto the loop, get a thread-safe future back
fut = asyncio.run_coroutine_threadsafe(handler("k1"), loop)
print(fut.result(timeout=2))
Rust: lazy futures, cancelled by drop
Rust inverts the model. A Future is a state machine that does nothing until an executor polls it, so creating two futures and awaiting them one after the other runs them sequentially; to overlap them you poll them together, for example with tokio::join!. Cancellation is dropping the future: its state machine is destroyed at its current await point and never resumes. That makes timeouts real, because tokio::time::timeout drops the inner future when the deadline passes, but it also means any code after an await may never run, so cleanup belongs in destructors. See the Rust async runtime picture.
Worked example: a deadline fan-out and the threads it leaks
A page needs a profile, which is mandatory, recent orders, which can degrade to an empty list, and recommendations, which are optional, all within about 150 ms. In Java:
ExecutorService io = Executors.newFixedThreadPool(32); // blocking clients: never the common pool
CompletableFuture<Profile> profile = CompletableFuture
.supplyAsync(() -> profiles.fetch(userId), io)
.orTimeout(150, TimeUnit.MILLISECONDS); // mandatory: fail the page
CompletableFuture<List<Order>> orders = CompletableFuture
.supplyAsync(() -> orderService.recent(userId), io)
.completeOnTimeout(List.of(), 150, TimeUnit.MILLISECONDS); // degrade
CompletableFuture<Recs> recs = CompletableFuture
.supplyAsync(() -> recommender.forUser(userId), io)
.completeOnTimeout(Recs.EMPTY, 120, TimeUnit.MILLISECONDS)
.exceptionally(ex -> Recs.EMPTY); // optional: swallow, but log upstream
CompletableFuture<Page> page = profile
.thenCombine(orders, Page::new)
.thenCombine(recs, Page::withRecs)
.whenComplete((pg, ex) -> { if (ex != null) log.warn("page failed", ex); });The three calls start immediately, so latency is the slowest mandatory branch, not the sum. Now count threads. If the recommender hangs, each request leaves one io thread blocked long after the page returned, because the timeout completed the future, not the call. At 300 requests per second and a 5-second client timeout, that is demand for 1,500 blocked calls against a pool of 32 threads: all 32 block, the rest queue, profile fetches queue behind dead recommendation calls, and the mandatory branch starts timing out too. The fix is to give the client its own timeout shorter than the page budget, and to isolate optional dependencies on their own small pool so a sick dependency can only exhaust itself. Sizing those pools is covered in thread pools.
Failure modes
- Lost errors. A Java chain with no terminal handler fails silently. End every chain with
whenComplete,handleor a join that someone observes. - Starvation deadlock. A task running on a fixed pool blocks on
get()of another task queued on the same pool; with every thread doing this, nothing progresses. Compose instead of blocking, or use separate pools. - Hijacked threads. A slow non-async callback runs on an I/O or event-loop thread that completed the stage. Move heavy work with the async variant and an explicit executor.
- Phantom cancellation. Cancel or timeout completes the future while the work continues and holds resources. Pass cancellation into the work itself.
- Never-completed promises. A code path forgets to resolve; readers wait forever and callbacks leak. Use timeouts on every wait and complete in a finally block.
- Accidental sequencing. Discarded
std::asyncfutures, or awaiting Rust futures one by one, removes the parallelism you designed. - Cross-thread asyncio calls. Resolving an
asyncio.Futurefrom a worker thread corrupts the loop. Usecall_soon_threadsafeorrun_coroutine_threadsafe.
What to do next
- Type out the Python promise above and write a test that registers callbacks from ten threads while another thread settles it; confirm each runs exactly once.
- For the language you ship in, fill in the five-question table from its documentation and keep it next to the code.
- Grep for blocking
get()orjoin()calls inside tasks running on a shared pool and replace them with composition. - Give every outbound call its own client-level timeout shorter than the caller's budget, and do not rely on future timeouts to free threads.
- Put optional dependencies on separate small executors so a hung dependency exhausts only its own pool.
- Ensure every chain ends in an observed handler and that unhandled-rejection events are logged and alerted on.
- Load-test with one dependency hanging and watch thread counts, not just latency.