Every asynchronous API eventually hands you an object that stands for a value that does not exist yet. Java calls it a CompletableFuture, C++ a std::future, JavaScript a Promise, Python has two different Future classes, and Rust has a trait. They look alike and behave differently in exactly the places that cause production bugs: which thread runs your callback, whether work starts before anyone asks for it, what cancellation really stops, and what happens to an error nobody looks at.

This article builds a future and promise from scratch in about sixty lines, so the mechanism is concrete, then walks the five design decisions every implementation makes and shows how each major language answers them, with code. A worked fan-out with deadlines and a catalogue of failure modes finish it. The conceptual architecture of the write-once cell is covered in futures and promises architecture; this page is about building, choosing and debugging.

The model in one picture

One asynchronous result: states and hand-offsProducerholds the promisePENDINGcallbacks queuedFULFILLEDvalue, set onceREJECTEDerror, set onceCANCELLEDa kind of rejectionresolve / rejectcancelConsumershold the futurethen / getExecutorruns callbacksdispatchresult or errorFirst settle wins; later resolve calls are ignored or rejected. Callbacks run once.
A producer holds the promise and settles it exactly once; consumers hold the future and either block for the result or register callbacks, which an executor runs once the state leaves PENDING.

Build one in sixty lines

A future needs four things: a state that moves exactly once from pending to settled, a lock so concurrent settles and registrations cannot interleave, a list of callbacks waiting for the result, and a way to block for the value when you must. Here is a complete, thread-safe version in Python using only the standard library. One object plays both roles; in a real library you would hand callers a read-only view.

import threading

PENDING, DONE, FAILED = "pending", "done", "failed"

class Promise:
    def __init__(self):
        self._cv = threading.Condition()
        self._state, self._value, self._callbacks = PENDING, None, []

    def _settle(self, state, value):
        with self._cv:
            if self._state is not PENDING:
                return False                      # write-once: later settles lose
            self._state, self._value = state, value
            callbacks, self._callbacks = self._callbacks, []
            self._cv.notify_all()                 # wake blocked result() callers
        for cb in callbacks:                      # never run user code under the lock
            cb(self)
        return True

    def resolve(self, value): return self._settle(DONE, value)
    def reject(self, exc):    return self._settle(FAILED, exc)

    def on_complete(self, cb):
        with self._cv:
            if self._state is PENDING:
                self._callbacks.append(cb)
                return
        cb(self)                                  # already settled: run now, on this thread

    def result(self, timeout=None):
        with self._cv:
            if not self._cv.wait_for(lambda: self._state is not PENDING, timeout):
                raise TimeoutError("still pending")
            if self._state is FAILED:
                raise self._value
            return self._value

    def then(self, fn):
        out = Promise()
        def step(src):
            if src._state is FAILED:
                out.reject(src._value)            # errors skip transformations
                return
            try:
                v = fn(src._value)
            except Exception as e:
                out.reject(e)
                return
            if isinstance(v, Promise):            # fn was itself async: flatten
                v.on_complete(lambda q: out._settle(q._state, q._value))
            else:
                out.resolve(v)
        self.on_complete(step)
        return out

Three details carry the design. The check-and-append in on_complete happens under the same lock as the transition in _settle, so a callback registered at the instant of completion is either queued and run by the settler or run immediately by the registrar, never lost and never run twice. Callbacks run outside the lock, so a callback that registers another callback cannot deadlock. And then flattens a returned promise, which is the difference between Java's thenApply and thenCompose collapsed into one method, as JavaScript does.

Composition is more callbacks

Combinators are just more callbacks. Waiting for all of a list is a countdown that fails fast on the first error:

def all_of(promises):
    out, results = Promise(), [None] * len(promises)
    remaining, lock = [len(promises)], threading.Lock()
    if not promises:
        out.resolve([])
    for i, pr in enumerate(promises):
        def done(src, i=i):
            if src._state is FAILED:
                out.reject(src._value)            # first failure wins; the rest are ignored
                return
            results[i] = src._value
            with lock:
                remaining[0] -= 1
                last = remaining[0] == 0
            if last:
                out.resolve(results)
        pr.on_complete(done)
    return out

Notice what fail-fast does not do: the other inputs keep running. Nothing in a future can stop the work that will complete it unless the producer checks for cancellation. That single fact explains most cancellation surprises below.

Five decisions every implementation makes

Every implementation answers the same five questions, and the answers are what you must know about the one in front of you.

QuestionJava CompletableFutureC++ std::futureJavaScript PromisePythonRust Future
Does work start before anyone waits?Yes, eagerYes with std::async(launch::async)Yes, executor runs at constructionSubmitted work and asyncio tasks yes; a bare coroutine not until awaitedNo, lazy until polled
How many readers?ManyOne; shared_future for manyManyManyOne owner
Where do callbacks run?Completing thread, or an executor for *AsyncNo callbacks in the standardAlways later, as microtasksconcurrent: completing thread; asyncio: the loopNo callbacks; the executor polls
What does cancel stop?Only the future, not the taskNo cancelNo cancel in the typeconcurrent: only if not startedDropping stops the work
Unobserved error?Silently droppedRethrown by get()Unhandled rejection eventasyncio logs at garbage collectionMust be handled to compile cleanly

Java: eager, and cancel does not stop the task

Java's CompletableFuture is eager and multi-reader. Its non-async methods run the callback on whichever thread completed the previous stage, or on the caller if the stage is already done; the *Async variants dispatch to the common pool or an executor you pass. Its cancel(mayInterruptIfRunning) ignores the flag: the Javadoc states that interrupts are not used, so cancel completes the future with a CancellationException and the supplier keeps running. The same holds for orTimeout, which completes the future exceptionally after the deadline but leaves the task occupying its thread. If work must actually stop, the task has to check a flag or the future's state itself. The full API is in CompletableFuture in Java.

C++: one reader, broken promises, blocking destructors

C++ splits the two ends into separate types and keeps them minimal:

#include <future>
#include <thread>

std::promise<int> p;
std::future<int> f = p.get_future();

std::thread producer([pr = std::move(p)]() mutable {
    try {
        pr.set_value(compute());
    } catch (...) {
        pr.set_exception(std::current_exception());   // errors travel as values
    }
});

int v = f.get();      // blocks; rethrows a stored exception; afterwards f.valid() == false
producer.join();

auto shared = std::async(std::launch::async, compute).share();  // many readers

Three rules cause most C++ bugs. get() may be called once; a second call is undefined behaviour unless you use std::shared_future. If a promise is destroyed before it is satisfied, waiting readers receive a std::future_error with the broken_promise code, so an early return in a producer surfaces as an error rather than a hang. And the future returned by std::async blocks in its destructor until the task finishes, so discarding it turns intended parallelism into sequential code. The standard has no then; continuation libraries and executors fill that gap.

JavaScript: always asynchronous

JavaScript guarantees that reactions never run synchronously, even on a promise that is already settled. They are queued as microtasks and run after the current script finishes, before timers:

console.log("1 sync");
Promise.resolve().then(() => console.log("3 microtask"));
setTimeout(() => console.log("4 timer"), 0);
console.log("2 sync");
// 1, 2, 3, 4 - always

// fail fast versus collect everything
const all = await Promise.all([a(), b(), c()]);          // rejects on the first failure
const settled = await Promise.allSettled([a(), b(), c()]);
const ok = settled.filter(r => r.status === "fulfilled").map(r => r.value);

That rule removes a whole class of reentrancy bugs that Java's same-thread callbacks allow, at the cost of a scheduling hop per step. A rejection with no handler raises an unhandled-rejection event; since Node.js 15 the default is to treat it as an uncaught error and exit, so a forgotten catch in a server can take the process down. Cancellation is a separate mechanism, AbortController, passed into the API that does the work. The async/await patterns article covers the syntax built on top.

Python: two futures that must not be mixed

Python has two incompatible future types. concurrent.futures.Future is thread-safe; add_done_callback runs on the thread that completes it, or immediately on the caller if it is already done, and cancel() succeeds only while the task is still waiting in the queue. asyncio.Future belongs to one event loop and is not thread-safe; calling its methods from another thread corrupts the loop's state quietly. Bridge explicitly:

import asyncio, concurrent.futures

pool = concurrent.futures.ThreadPoolExecutor(max_workers=8)

async def handler(key):
    loop = asyncio.get_running_loop()
    return await loop.run_in_executor(pool, legacy_blocking_lookup, key)

# from an ordinary thread: schedule onto the loop, get a thread-safe future back
fut = asyncio.run_coroutine_threadsafe(handler("k1"), loop)
print(fut.result(timeout=2))

Rust: lazy futures, cancelled by drop

Rust inverts the model. A Future is a state machine that does nothing until an executor polls it, so creating two futures and awaiting them one after the other runs them sequentially; to overlap them you poll them together, for example with tokio::join!. Cancellation is dropping the future: its state machine is destroyed at its current await point and never resumes. That makes timeouts real, because tokio::time::timeout drops the inner future when the deadline passes, but it also means any code after an await may never run, so cleanup belongs in destructors. See the Rust async runtime picture.

Worked example: a deadline fan-out and the threads it leaks

A page needs a profile, which is mandatory, recent orders, which can degrade to an empty list, and recommendations, which are optional, all within about 150 ms. In Java:

ExecutorService io = Executors.newFixedThreadPool(32);   // blocking clients: never the common pool

CompletableFuture<Profile> profile = CompletableFuture
    .supplyAsync(() -> profiles.fetch(userId), io)
    .orTimeout(150, TimeUnit.MILLISECONDS);                    // mandatory: fail the page

CompletableFuture<List<Order>> orders = CompletableFuture
    .supplyAsync(() -> orderService.recent(userId), io)
    .completeOnTimeout(List.of(), 150, TimeUnit.MILLISECONDS); // degrade

CompletableFuture<Recs> recs = CompletableFuture
    .supplyAsync(() -> recommender.forUser(userId), io)
    .completeOnTimeout(Recs.EMPTY, 120, TimeUnit.MILLISECONDS)
    .exceptionally(ex -> Recs.EMPTY);                         // optional: swallow, but log upstream

CompletableFuture<Page> page = profile
    .thenCombine(orders, Page::new)
    .thenCombine(recs, Page::withRecs)
    .whenComplete((pg, ex) -> { if (ex != null) log.warn("page failed", ex); });

The three calls start immediately, so latency is the slowest mandatory branch, not the sum. Now count threads. If the recommender hangs, each request leaves one io thread blocked long after the page returned, because the timeout completed the future, not the call. At 300 requests per second and a 5-second client timeout, that is demand for 1,500 blocked calls against a pool of 32 threads: all 32 block, the rest queue, profile fetches queue behind dead recommendation calls, and the mandatory branch starts timing out too. The fix is to give the client its own timeout shorter than the page budget, and to isolate optional dependencies on their own small pool so a sick dependency can only exhaust itself. Sizing those pools is covered in thread pools.

Failure modes

  • Lost errors. A Java chain with no terminal handler fails silently. End every chain with whenComplete, handle or a join that someone observes.
  • Starvation deadlock. A task running on a fixed pool blocks on get() of another task queued on the same pool; with every thread doing this, nothing progresses. Compose instead of blocking, or use separate pools.
  • Hijacked threads. A slow non-async callback runs on an I/O or event-loop thread that completed the stage. Move heavy work with the async variant and an explicit executor.
  • Phantom cancellation. Cancel or timeout completes the future while the work continues and holds resources. Pass cancellation into the work itself.
  • Never-completed promises. A code path forgets to resolve; readers wait forever and callbacks leak. Use timeouts on every wait and complete in a finally block.
  • Accidental sequencing. Discarded std::async futures, or awaiting Rust futures one by one, removes the parallelism you designed.
  • Cross-thread asyncio calls. Resolving an asyncio.Future from a worker thread corrupts the loop. Use call_soon_threadsafe or run_coroutine_threadsafe.

What to do next

  1. Type out the Python promise above and write a test that registers callbacks from ten threads while another thread settles it; confirm each runs exactly once.
  2. For the language you ship in, fill in the five-question table from its documentation and keep it next to the code.
  3. Grep for blocking get() or join() calls inside tasks running on a shared pool and replace them with composition.
  4. Give every outbound call its own client-level timeout shorter than the caller's budget, and do not rely on future timeouts to free threads.
  5. Put optional dependencies on separate small executors so a hung dependency exhausts only its own pool.
  6. Ensure every chain ends in an observed handler and that unhandled-rejection events are logged and alerted on.
  7. Load-test with one dependency hanging and watch thread counts, not just latency.
Key takeaway: A future is a write-once cell plus a callback list guarded by one lock; everything else is policy. Before relying on one, know whether its work starts eagerly, how many readers it allows, which thread runs callbacks, what cancel actually stops and what happens to unobserved errors. Java and Python futures do not stop work on cancel, C++ futures are single-shot and can block in destructors, JavaScript reactions are always asynchronous, and Rust futures are lazy and cancelled by drop. Compose instead of blocking, and put timeouts on the work, not just the future.