Most software spends a large share of its code handling failure and almost none of its design effort on what it says when failure happens. The result is familiar: "Something went wrong", "Invalid input", a stack trace in a dialog box, or a 500 with an HTML error page returned to a JSON client. Each one turns a recoverable situation into a support ticket.
This guide treats error messages as an engineering artifact rather than copywriting. It covers the anatomy of a message a person can act on, how one failure should be rendered differently for an end user, an API client and an operator, how to back messages with a stable catalog in code, conventions for command-line tools and web forms, what must never appear in an error, and how to test and measure messages so they stay good. The wire format for API errors is covered in how to design an API; this page is about the words and the machinery behind them.
What an error message is for
An error message exists to get the reader to the next useful action. Jakob Nielsen's usability heuristics put it as helping users recognize, diagnose and recover from errors, and that ordering is a good test. Recognize: the reader can tell that something failed and what it was. Diagnose: they understand why, in their terms, not the system's. Recover: they know what to do now, even if the answer is "nothing, it is our fault, try later".
That gives three questions every message must answer: what happened, why, and what next. The third question is the one most often missing and the one that matters most, because it is the only part that changes what the reader does.
A corollary: if you cannot write the "what next" line, the failure is probably being surfaced at the wrong layer: the code that knows what to do should handle it, or translate it into something the caller can act on.
Anatomy, by rewriting
The fastest way to learn the anatomy is to rewrite real messages. Each rewrite below keeps the facts and adds what was missing.
| Before | Problem | After |
|---|---|---|
| Invalid input. | No location, no reason, no fix | The start date 31/02/2026 is not a real date. Enter a date as DD/MM/YYYY. |
| Error 0x80070005 | Code without meaning | You do not have permission to write to C:\Reports. Choose another folder or ask your administrator for write access. |
| Payment failed. | No cause, no next step | Your card ending 4242 expired in 09/2026. Update the card in Billing, then retry. You have not been charged. |
| Something went wrong. | Hides whether retry helps | We could not save your draft because our servers are busy. Your text is still here. Try again in a minute. |
| NullPointerException at OrderService.java:212 | Internals shown to a user | We could not load this order. Our team has been notified. Reference: 7f3a9c1e2b44. |
The patterns in the right-hand column generalise. Name the specific object (which field, which file, which card). State the cause in the reader's vocabulary. Give exactly one primary next step, and say what state they are in, especially whether anything was lost or charged. Keep the headline short and put detail after it, because people read the first few words and skip the rest. When the system is at fault, say so plainly and give a reference the reader can quote to support.
One error, three audiences
The core engineering decision is to stop formatting strings at the point of failure. Raise a typed error that carries a stable code, the parameters that explain it, the underlying cause and a unique id, and let separate renderers produce text for each audience. The end user gets a localised sentence with no internals. An API client gets a machine-readable body with the code and a field it can branch on, such as whether retry is safe. The operator gets a structured log record with the full cause chain and context.
The thread that ties them together is a correlation id. Show it to the user, return it in the API body and put it on the log record. When a customer pastes "ref 7f3a9c1e2b44" into a ticket, support can find the exact log line in seconds instead of searching by timestamp and email address. The error id can be your trace id.
A stable error catalog in code
Messages scattered across hundreds of raise statements drift: the same failure is described five ways, codes collide, and nobody can review the wording. A catalog fixes that. Each entry has a stable code that is never reused, an HTTP status, a user template, a required action line and a retryable flag. Code raises by code and parameters; the catalog owns the words.
from dataclasses import dataclass, field
import uuid
@dataclass(frozen=True)
class ErrorSpec:
code: str # stable, never reused: "PAY-CARD-EXPIRED"
status: int # HTTP status for API callers
title: str # fixed summary of the problem type
user: str # template shown to people
action: str # the next step, always present
retryable: bool = False
CATALOG = {s.code: s for s in [
ErrorSpec("PAY-CARD-EXPIRED", 402, "Card expired",
"Your card ending {last4} expired in {expiry}.",
"Update the card in Billing, then retry the payment."),
ErrorSpec("PAY-GATEWAY-TIMEOUT", 503, "Payment provider unavailable",
"We could not reach the payment provider. You have not been charged.",
"Try again in a minute. If it keeps failing, contact support.",
retryable=True),
]}
@dataclass
class AppError(Exception):
code: str
params: dict = field(default_factory=dict)
cause: Exception | None = None
error_id: str = field(default_factory=lambda: uuid.uuid4().hex[:12])
def render_user(self):
spec = CATALOG[self.code]
return f"{spec.user.format(**self.params)} {spec.action} (ref {self.error_id})"
def render_api(self):
spec = CATALOG[self.code]
return {"type": f"https://errors.example.com/{self.code}",
"title": spec.title,
"status": spec.status,
"detail": spec.user.format(**self.params),
"code": self.code, "retryable": spec.retryable,
"instance": f"/errors/{self.error_id}"}
def log_fields(self):
return {"error.code": self.code, "error.id": self.error_id,
"error.cause": repr(self.cause), **self.params}Several properties of this design matter more than the details. The action field is mandatory, so a missing next step fails at definition time rather than in production. The code is namespaced by domain, so teams can own their slice without collisions, and because codes are never reused, dashboards and support macros keyed on them stay valid for years. Templates take named parameters rather than pre-formatted strings, which is what makes translation possible. The cause is kept on the object for logs but never interpolated into user text. For API callers the render follows the RFC 9457 problem details shape (type, title, status, detail, instance) with extension members for the code and retryability; the API style guide covers enforcing that shape across services.
Command-line tools
CLI errors have their own conventions, and scripts depend on them. Write errors to standard error, not standard output, so pipelines do not swallow them as data. Exit with a non-zero status that means something. Shells reserve 126 (found but not executable), 127 (command not found) and 128 plus a signal number (killed by a signal), so avoid those. The BSD sysexits.h header defines a widely copied set: 64 for usage errors, 65 for bad input data, 66 for a missing input file, 69 for an unavailable service, 70 for an internal software error, 75 for a temporary failure where retry may work and 78 for configuration errors. Document whichever set you choose, because a deploy script that retries on 75 and stops on 78 is only possible if the tool is consistent.
The best compiler diagnostics show what good looks like. The Rust compiler prints a code such as E0382, points at the exact span in the source, explains the cause and suggests a fix, and rustc --explain E0382 prints a long-form explanation. The same shape works for any tool:
$ deploy --env prod
error[DEP-012]: config file not found: ./deploy.yaml
looked in: ./deploy.yaml, ~/.config/deploy/deploy.yaml
hint: run `deploy init` to create one, or pass --config PATH
more: deploy explain DEP-012
$ echo $?
78The first line is greppable and carries the code; the detail lines say where the tool looked. The hint gives the next action, and an explain subcommand holds the long version so the default output stays short. Add a --verbose or debug flag that prints the cause chain for operators, and keep colour optional, disabled when standard error is not a terminal or when the NO_COLOR environment variable is set.
Forms and accessibility
In a web form, the best error is the one prevented: constrain inputs (a date picker, an input mode for numbers), accept common formats instead of rejecting them, and validate as the user leaves a field rather than on every keystroke. When an error does occur, show it next to the field it concerns, in text, not only as a red border, because colour alone fails people with colour-blindness. Never clear what the user typed. On submit, summarise all errors at the top with links to each field and move focus to the summary.
WCAG makes some of this a requirement. Success criterion 3.3.1, Error Identification, requires that a detected input error identifies the item in error and describes it in text, and 3.3.3, Error Suggestion, requires suggesting a correction when one is known, unless that would compromise security. In markup, set aria-invalid on the field and connect the message with aria-describedby so a screen reader announces it with the field; use a live region or role=alert for the submit summary so it is announced without the user hunting for it.
What an error must never say
Error text is an attack surface. Stack traces, SQL fragments, file paths, hostnames, library versions and internal ids in user-visible errors give an attacker a map of your system, so they belong in logs only. Login and password-reset flows must not reveal whether an account exists: "The email or password is incorrect" is deliberately less helpful than "No account with that email", and the reset page should say the same thing whether or not the address is registered, with the same response time. Logs have the opposite risk: the parameters you attach for operators must not include passwords, tokens or full card numbers, so redact at the logging layer rather than trusting every call site.
Testing and measuring messages
Messages rot like any other code, so test them like code. Lint the catalog for missing actions, vague phrases and length; snapshot-test the rendered output so a wording change shows up in review; and keep the set of retired codes so a test can stop anyone reusing one.
import pytest
BANNED = ("an error occurred", "something went wrong", "invalid input", "null")
@pytest.mark.parametrize("spec", CATALOG.values(), ids=lambda s: s.code)
def test_every_message_is_actionable(spec):
text = (spec.user + " " + spec.action).lower()
assert spec.action.strip(), "every error needs a next step"
assert not any(b in text for b in BANNED), "vague wording"
assert len(spec.user) <= 160, "keep the headline short"
RETIRED_CODES = {"PAY-CARD-DECLINED-OLD"} # append-only list kept in the repo
def test_retired_codes_are_never_reused():
assert not RETIRED_CODES & CATALOG.keys()
def test_render_snapshot(snapshot):
err = AppError("PAY-CARD-EXPIRED", {"last4": "4242", "expiry": "09/2026"})
err.error_id = "abc123def456"
assert err.render_user() == snapshotThen measure in production. Count errors by code, not by message text, and look at the top ten every week. A code that is frequent and user-facing is a design problem, not a support problem: the best fix is often to prevent it, as with an expired card that could have been flagged a week earlier. Tie the high-volume codes to a runbook entry, as described in how to write a runbook. For agent and tool interfaces the reader is a model, but the rules are the same: MCP error handling shows how a tool error the model can read lets it correct its own call.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Catch-all handler | Every failure reads "Something went wrong" | Map exceptions to catalog codes at the boundary; unknown errors get a generic code plus an id |
| Message built at raise site | Same failure worded five ways, untranslatable | Raise by code and parameters; render centrally |
| Wrong retry advice | Users hammer a permanent failure, or give up on a transient one | Make retryable an explicit catalog field and render it |
| Internals leak | Stack traces or SQL in UI or API detail | User and API renderers never read the cause; scan responses in tests |
| Codes reused or renamed | Dashboards and support macros silently wrong | Codes are append-only; retire, never recycle |
| Error replaces input | Form cleared after a failed submit | Preserve state; show errors inline next to fields |
Trade-offs
A catalog costs ceremony: every new failure needs an entry, a code and reviewed wording, which slows teams that are used to raising with a string. The payoff is consistency, translation and measurability, and it grows with the number of services and locales; a small internal tool may reasonably skip it. Specific messages help users but can help attackers too, so authentication paths deliberately trade helpfulness for safety. Typed error results, as in functional error handling, make failure explicit in signatures and pair naturally with a catalog, at the cost of more verbose code paths.
What to do next
- Pull the ten most frequent user-facing errors from your logs and rewrite each to answer what happened, why and what next.
- Introduce a catalog with stable, namespaced codes, a mandatory action field and a retryable flag; raise by code and parameters.
- Render separately for users, API clients and logs, and put one correlation id on all three.
- For CLIs, write errors to stderr, document your exit codes and add an explain subcommand or URL for long-form help.
- Audit forms for inline, text-based errors with aria-invalid and aria-describedby, and confirm input is never cleared.
- Remove stack traces, SQL and account-existence hints from every user-visible path; redact secrets at the logging layer.
- Add catalog lint and snapshot tests, and review error counts by code every week.