Choosing a log library looks like a five-minute decision: pick whatever the tutorial used. Then the service goes to production and the library starts making decisions for you. It decides whether a debug statement in a hot loop costs nothing or a megabyte of garbage per second, whether a slow disk stalls request threads or silently drops lines, whether log records carry the trace ID your incident responders need, and, as Log4Shell demonstrated, whether a string from an attacker can make your process execute code.
This guide gives you a method rather than a winner. It explains what a log call actually does, separates the API you write against from the backend that does the work, shows where cost and loss happen, and ends with a worked selection for two services and a benchmark harness you run yourself. What to put in each event, levels and redaction are covered in structured logging best practices; this page is about the machinery that carries those events.
What happens inside a log call
Every mainstream library runs the same pipeline. The call site checks whether the level is enabled; if not, the call should return almost immediately. If it is enabled, the library builds a record: timestamp, level, logger name, message template, arguments, and context such as a request ID pulled from thread-local or async-local storage. An encoder turns the record into bytes, usually JSON in production. Finally a writer, called an appender, handler, sink or transport depending on the ecosystem, puts those bytes somewhere: stdout for a container log collector, a file, a socket, or an OpenTelemetry exporter.
Each stage is a selection criterion. The level check decides the cost of disabled logging. Record building decides allocation. The encoder decides whether output is structured and safe against injection. The writer decides latency, durability and what happens when the destination cannot keep up.
Facade versus backend
The most important architectural choice is to separate the logging API your code calls from the implementation that formats and writes. Libraries you publish should depend only on the API, so the application that embeds them chooses the backend once. Each ecosystem has settled on its own version of this split.
| Ecosystem | API or facade | Common backends | Notes |
|---|---|---|---|
| Java | SLF4J (2.x adds a fluent API) | Logback, Log4j 2 | Log4j 2 also has its own API; bridges route java.util.logging and others |
| Python | stdlib logging | stdlib handlers, structlog on top | structlog can render through stdlib or replace it |
| Go | log/slog (Go 1.21+) | slog handlers, zap, zerolog | slog's Handler interface lets zap and others plug in |
| Rust | log or tracing | env_logger, tracing-subscriber | tracing adds spans and structured fields |
| .NET | Microsoft.Extensions.Logging | Built-in providers, Serilog, NLog | Serilog can act as a provider behind ILogger |
| Node.js | No standard | pino, winston | Pick one per codebase and wrap it |
The practical rule: application code may use the backend's richer features through the facade, but reusable packages must never force a backend on their consumers. In Java, a library that depends on Logback directly will fight the application's Log4j 2 configuration; in Node, a package that writes with its own winston instance ignores your redaction rules.
The cost model: what a disabled debug call costs
A disabled call should cost a level comparison. It often costs much more because the arguments are evaluated before the library sees them. String concatenation, f-strings and calls such as toString() on a large object run even when the level is off. Every ecosystem has an idiom that defers that work, and your chosen library must support it.
// Java (SLF4J): placeholders defer formatting; the fluent API defers argument computation
log.debug("cart {} has {} items", cartId, items.size()); // cheap when disabled
log.atDebug().addKeyValue("cart", cartId)
.addArgument(() -> expensiveSummary(cart)).log("cart state {}");
# Python: pass args, never pre-format; guard truly expensive work explicitly
log.debug("cart %s has %d items", cart_id, len(items)) # formatted only if enabled
if log.isEnabledFor(logging.DEBUG):
log.debug("cart state %s", expensive_summary(cart))
// Go (log/slog): LogAttrs avoids the variadic any-slice allocation
logger.LogAttrs(ctx, slog.LevelDebug, "cart",
slog.String("cart", cartID), slog.Int("items", len(items)))
// .NET: compile-time generated logging methods avoid boxing and parsing templates
[LoggerMessage(Level = LogLevel.Debug, Message = "Cart {CartId} has {Count} items")]
static partial void CartItems(ILogger logger, string cartId, int count);Enabled calls cost allocation and encoding. Libraries designed for low allocation, such as zap and zerolog in Go or pino in Node, encode directly into reusable buffers rather than building an intermediate map. That matters in services handling tens of thousands of requests per second with several log lines each; it rarely matters in a batch job. Measure before paying for complexity, using the harness later in this article.
Synchronous, asynchronous and what gets dropped
A synchronous writer performs I/O on the calling thread. It is simple and loses nothing that was acknowledged, but a slow disk, a full pipe to a log collector, or a blocked network sink stalls requests. An asynchronous writer puts records on a queue and lets a background thread write them, which removes I/O from the request path but forces a decision you must make explicitly: what happens when the queue is full.
- Block: the caller waits. Latency spikes, but nothing is lost. Log4j 2's async loggers, built on the LMAX Disruptor, wait for space by default when their ring buffer is full.
- Drop: the record is discarded, ideally counted. Python's
QueueHandlerenqueues without waiting, so with a bounded queue a full queue becomes a handler error and a lost record. Logback'sAsyncAppenderby default discards TRACE, DEBUG and INFO events once only a fifth of its queue remains, keeping WARN and ERROR. That default surprises people investigating a gap in INFO logs during an incident. - Grow: an unbounded queue never blocks or drops until the process runs out of memory, which converts a logging slowdown into an outage.
Asynchronous writers also lose whatever is still queued when the process crashes, which is precisely the moment you want the last lines. Make sure shutdown hooks flush the queue, write fatal errors synchronously, and in containers prefer writing to stdout and letting a node agent handle shipping, so the application never blocks on the network. pino takes this further by moving transports into worker threads. Treat the pipeline beyond the process with the same rigour described in OpenTelemetry pipeline design.
Structure, context and OpenTelemetry
Selection criteria for the record itself: first-class key-value fields rather than string interpolation; a JSON encoder whose field names you control; and context propagation that works with your concurrency model. Thread-local MDC in Java breaks across executor hops unless something copies it; Python needs contextvars-based context for asyncio; Go passes context.Context explicitly, and slog handlers can extract values from it.
OpenTelemetry defines a Logs Bridge API intended for library authors rather than application code: existing libraries feed their records into OpenTelemetry through appenders or handlers, which attach the active trace and span IDs and export over OTLP. Before choosing, check that a maintained bridge exists for your backend. The Java side, including the agent's automatic log correlation, is covered in OpenTelemetry Java in depth.
Security and supply-chain history
Log4Shell, CVE-2021-44228, disclosed in December 2021, let attackers trigger JNDI lookups by placing a string such as ${jndi:ldap://...} anywhere that reached a log message, which led to remote code execution. The fixes came in waves: 2.15.0, then 2.16.0 for CVE-2021-45046, 2.17.0 for CVE-2021-45105 and 2.17.1 for CVE-2021-44832. The general lessons apply to every library you evaluate:
- Prefer libraries that treat message arguments as data and never interpret them. Feature lists that include lookups, templating or expression evaluation in messages are a risk surface.
- Use structured encoders so newlines and quotes in user input cannot forge extra log lines, which is the log injection attack.
- Check the project's release cadence, security advisories and how quickly past issues were patched. A library you cannot upgrade quickly is a liability.
- Keep the dependency tree small; a logging library is loaded into every service you own.
Worked example: choosing for two services
Consider two services. A Go payment API handles a steady high request rate with strict latency targets. A Python worker processes documents in batches. Score candidates on weighted criteria, where the weights come from each service's needs:
| Criterion | Weight (Go API) | Weight (Python worker) |
|---|---|---|
| Structured output and context propagation | 3 | 3 |
| Disabled and enabled call cost | 3 | 1 |
| Explicit overload behaviour | 2 | 1 |
| OpenTelemetry bridge maintained | 2 | 2 |
| Standard API or facade | 2 | 2 |
| Security record and dependency size | 2 | 2 |
For the Go API, log/slog scores well on standardisation and bridging and wins outright unless measurement shows its default handlers cost too much; in that case zap or zerolog can sit behind the slog API, keeping application code portable. For the Python worker, the stdlib logging module plus structlog for structured output and context is enough, with a QueueHandler added only if profiling shows handler I/O in the hot path. Neither choice was made on a published benchmark; both were made on criteria, then confirmed with a measurement like this one:
import io, json, logging, timeit
def make_logger(level):
log = logging.getLogger("bench")
log.handlers.clear()
h = logging.StreamHandler(io.StringIO()) # isolate encoding cost from disk I/O
h.setFormatter(logging.Formatter('{"t":"%(asctime)s","l":"%(levelname)s","m":"%(message)s"}'))
log.addHandler(h)
log.setLevel(level)
log.propagate = False
return log
payload = {"cart": "c-123", "items": list(range(50))}
for level in (logging.INFO, logging.DEBUG):
log = make_logger(level)
n = 200_000
secs = timeit.timeit(lambda: log.debug("cart %s", json.dumps(payload)), number=n)
print(logging.getLevelName(level), f"{secs / n * 1e6:.2f} us per call")
# The INFO run shows the cost of a disabled call whose argument is still evaluated (json.dumps).
# Rewrite it to pass payload and serialise lazily, rerun, and compare before choosing anything.
Failure modes
- Two backends fighting. Transitive dependencies bring in a second logging implementation; output is duplicated or configuration silently ignored. Use bridges and exclude extra bindings.
- Silent drops under load. An async appender discards INFO during the incident you are investigating. Export a drop counter and alert on it.
- Lost last words. Queued records vanish on crash. Flush on shutdown and write fatal paths synchronously.
- Hot-path allocation. Pre-formatted strings in debug calls create garbage even when disabled. Lint for concatenation and f-strings in log calls.
- Context lost across async boundaries. Request IDs disappear after a thread pool or event-loop hop. Test correlation explicitly.
- Interpreting user input. Lookup or template features in messages turn logging into code execution or injection.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Standard API (slog, stdlib logging, SLF4J) | Portability, ecosystem bridges | Sometimes slower defaults |
| Low-allocation backend (zap, zerolog, pino) | Lower CPU and GC on hot paths | Less familiar APIs, more configuration |
| Asynchronous writer | I/O off the request path | Drop or block policy, crash loss |
| stdout plus node agent | Simple app, central shipping | Depends on platform log pipeline |
| Direct OTLP export | Trace correlation end to end | Network dependency inside the app |
Picking a log library is a smaller version of the same requirement-by-requirement method used in how to pick a message broker: name the requirements, weight them, and test the top two candidates on your own workload.
What to do next
- Inventory which logging APIs and backends are already on your classpath or dependency tree, including transitive ones.
- Choose one facade per language and require reusable packages to depend only on it.
- Decide and document the overload policy for every async writer, and export a dropped-record counter.
- Adapt the benchmark harness to your language and measure disabled and enabled call cost with realistic arguments.
- Verify that trace and span IDs appear in log records after thread pool and async hops.
- Subscribe to security advisories for the chosen library and confirm you can ship an upgrade to every service within a day.