Every name in a codebase is a tiny piece of documentation that is read far more often than it is written, and unlike a comment it travels with the value everywhere it goes: into call sites, stack traces, log lines, database schemas, dashboards and other teams' code. A good name lets a reader predict what a thing is and what it does without opening it. A bad one forces them to read the implementation, or worse, lets them predict something false.
This guide treats naming as an engineering skill with checkable rules: what a name must carry, how long it should be, conventions per kind of identifier, the boundary-crossing names that cost most to change, common anti-patterns, and safe renaming, ending with a checklist for code review.
Why names carry so much weight
A reader building a mental model of unfamiliar code works mostly from names. They scan a function, see retry_count and fetch_invoice, and form expectations: an integer, a network call that may fail. If those expectations hold, they can skip the bodies and reason at the level of the names. If they do not hold, every name becomes suspect and the reader has to verify everything, which is slow and error-prone.
Names also decide what is findable. Engineers search code and logs by name far more than they navigate by structure. A concept that is called account in one module, customer in another and tenant in a third cannot be traced with one search, and bugs hide in the seams between the three. Consistent names are what make grep, code search and log queries work.
Names are also read by tools: code search and AI coding assistants infer intent from identifiers, so a validate that also writes to the database misleads them too.
What a name must carry
A useful name answers the questions a reader will have at the point of use, and nothing more. Four kinds of information come up again and again:
- Role: what the thing is for in this context, not what type it is.
deadlinebeatsdate2;retry_budgetbeatsnonce the scope is more than a few lines. - Unit: for any quantity with a unit, put the unit in the name unless the type system enforces it.
timeout_msversustimeout_sis the difference between a working retry loop and one that waits sixteen minutes. - Kind: whether it is a single value, a collection, a mapping or a predicate. Plural nouns for collections,
x_by_yfor maps,is_orhas_for booleans. - State or lifecycle when it matters:
raw_payloadversusvalidated_payload,draft_orderversusplaced_order. Naming the stage stops someone passing unvalidated data where validated data is required.
Length should be proportional to scope: the further a name is read from where it is defined, the more context it must carry by itself. A loop index used on the next line can be i; a public API field must be unambiguous to someone who has never seen your code. Long names in tiny scopes add noise that hides the logic.
Naming by kind
Each kind of identifier has a grammar that readers expect. Following it lets the reader infer behaviour from shape alone.
| Kind | Convention | Good | Avoid |
|---|---|---|---|
| Boolean | Predicate, positive form | is_active, has_access, can_retry | not_disabled, flag |
| Query function | Noun or get_, no side effects | total_price(), get_user(id) | process() |
| Command function | Imperative verb | send_receipt(), cancel_order() | receipt() |
| Expensive call | Verb that signals cost | fetch_, load_, compute_ | get_ for a network call |
| Collection | Plural noun | pending_jobs | job_list, data |
| Mapping | values_by_key | users_by_email | user_map |
| Type or class | Singular domain noun | Invoice, RetryPolicy | InvoiceManager, Utils |
| Event | Past tense fact | OrderPlaced, PaymentFailed | PlaceOrder |
Two rules do most of the work. First, a function name should describe its effect at the level the caller cares about: if calling it can change state, send a message or cost a network round trip, the name should say so. Second, pairs should be symmetric. If you have start, its partner is stop, not end or halt; open goes with close, begin with end, min with max, first with last. Readers guess the second name from the first, so make the guess right.
Generic suffixes such as Manager, Helper and Util usually mean a type has no single responsibility yet; asking what it actually does often yields a better name.
Worked example: renaming a function until it explains itself
Here is a function of the kind found in many codebases. It works, it is short, and it is almost impossible to use correctly without reading it:
def process(d, f=False):
res = []
for x in d:
t = x["ts"] - x["st"]
if t > 30 and not (f and x["rt"]):
res.append(x)
return resA caller has to guess everything. process says nothing about the effect. f is a boolean whose meaning is invisible at the call site. ts and st could be timestamps or states, and 30 has no unit.
Suppose the call sites show that f is set to True to skip retried jobs, and that the fields are epoch seconds. The rename then follows mechanically:
SLOW_JOB_THRESHOLD_SECONDS = 30
def find_slow_jobs(jobs: list[dict], skip_retries: bool = False) -> list[dict]:
"""Return jobs whose wall-clock run time exceeded the threshold."""
slow_jobs = []
for job in jobs:
# input keys belong to the job record format, so bind them to clear locals
started_epoch_s, finished_epoch_s, is_retry = job["st"], job["ts"], job["rt"]
run_seconds = finished_epoch_s - started_epoch_s
if run_seconds > SLOW_JOB_THRESHOLD_SECONDS and not (skip_retries and is_retry):
slow_jobs.append(job)
return slow_jobsEvery change answers one of the four questions. The function name states its effect and result. The flag keeps its polarity and default, so call sites change only by name. The threshold is a named constant with its unit. The record keys stay as they are, because they belong to a data format other code shares; well-named locals carry role and unit instead. The behaviour is unchanged, but the next reader can use it from the signature alone.
One word per concept, one concept per word
In large systems the bigger problem is vocabulary drift: one idea under several names, or one name for several ideas. Synonyms break search. Homonyms are worse: if account means a login identity in the auth service and a billing ledger in the payments service, someone will eventually join the two on the wrong key.
The fix is a glossary that lives in the repository, next to the code, and is changed in pull requests like code. Each entry gives the term, a one-sentence definition, the identifiers that implement it, and the terms it must not be confused with. Domain-driven design calls this a ubiquitous language: the words the business uses, used identically in conversations, tickets, code and schemas. Where two parts of the system genuinely mean different things by one word, qualify both names (auth_account, billing_account) rather than letting context disambiguate.
Abbreviations need a policy too: industry-standard ones (id, url) are fine, while team-invented ones (cust_ord_proc) are a private dialect every new hire must learn.
Names that cross boundaries
Inside one function, a bad name costs a few seconds to fix with an editor rename. Once a name crosses a boundary it acquires readers you do not control, and the cost of changing it jumps. Spend your naming effort where it is hardest to undo:
- API fields and endpoints. Clients hard-code them. Pick one casing convention per API and keep it; prefer full words; include units in numeric fields (
ttl_seconds); avoid booleans that will grow a third state, and use an enum such asstatusinstead. - Database tables and columns. Queries, reports, ETL jobs and BI dashboards reference them by name, often outside your repository. Name columns for the fact they hold (
placed_at,amount_cents), not for how they were first used. - Metrics and labels. Dashboards and alert rules bind to the exact string. Follow your metrics system's conventions; for Prometheus that means snake_case, base units such as seconds and bytes, and the
_totalsuffix on counters. - Events and messages. Consumers you have never met subscribe by name. Never reuse an event name with new meaning; version the schema.
- Configuration keys and feature flags. Name a flag for the behaviour when on, and record an owner and expiry.
# Counter: monotonically increasing, base unit, _total suffix
http_requests_total{method="POST", route="/orders", status_class="5xx"}
# Histogram in base units (seconds, not milliseconds)
http_request_duration_seconds_bucket{route="/orders", le="0.25"}
# Gauge: current value, unit in the name
queue_depth_messages{queue="billing"}
process_resident_memory_bytes
# Avoid
request_time_ms # use base units: seconds
errors{user_id="83141"} # unbounded label: one series per user
Anti-patterns and how to spot them in review
| Anti-pattern | Example | Why it hurts | Fix |
|---|---|---|---|
| Lying name | validate() that also saves | Readers skip the side effect | Name the effect or split the function |
| Negated boolean | disable_cache = False | Double negatives at every call site | use_cache = True |
| Meaningless noun | data, info, obj | Carries no role | Name the domain thing |
| Numbered siblings | user1, user2 | Hides the difference | sender, recipient |
| Type encoding | strName, lstJobs | Duplicates the type system and goes stale | Drop the prefix |
| Missing unit | timeout = 30 | Seconds or milliseconds? | timeout_s = 30 |
| Stale name | legacy_new_flow | Describes history, not behaviour | Rename to the current behaviour |
In review, the cheapest check is to read only the signatures and call sites of a change and predict the behaviour. Wherever your prediction is wrong or uncertain, a name is the problem. Also search for the new identifier's obvious synonyms; if the concept already exists under another name, reuse it.
Linters can enforce casing, boolean prefixes and banned words like Util. They cannot judge whether a name tells the truth, so the prediction test stays a human job, and it matters most for generated code, whose plausible names may describe intent rather than behaviour.
Renaming safely
Inside one codebase, use the language-server rename, which understands scope, then search for the old name as a string: refactoring tools miss reflection, serialised field names, templates, SQL strings, config files and dashboards. Ship a pure rename as its own commit.
Across a boundary, a rename is a migration and needs the expand and contract pattern: add the new name alongside the old, move every reader and writer, verify, then remove the old one. For an API field that means returning both fields for a deprecation window and announcing the removal date. For a metric it means emitting both series until dashboards and alerts have moved. For a database column it looks like this:
-- Renaming a column without downtime: expand, migrate, contract.
-- Release 1 (expand): add the new name, keep the old one working.
ALTER TABLE orders ADD COLUMN placed_at timestamptz;
-- application writes BOTH created_ts and placed_at; reads prefer placed_at
-- Backfill in batches so locks stay short.
UPDATE orders SET placed_at = created_ts
WHERE placed_at IS NULL AND id BETWEEN :lo AND :hi;
-- Release 2: every reader uses placed_at; writes still go to both.
-- Verify: SELECT count(*) FROM orders WHERE placed_at IS DISTINCT FROM created_ts; -- expect 0
-- Release 3 (contract): stop writing created_ts, then drop it.
ALTER TABLE orders DROP COLUMN created_ts;Each release is independently deployable and reversible: if release 2 misbehaves, roll it back and the old column is still being written. A direct RENAME COLUMN is safe only when every reader deploys atomically with the schema change, which with replicas, analytics jobs and several services is rarely true.
Some names should never be renamed. A public metric with years of history may be cheaper to keep than to migrate; record the mismatch in the glossary so the next reader is warned.
Trade-offs
Precise names get long, and long names make expressions hard to read. Restructure rather than abbreviate: extract a well-named intermediate variable so each line holds fewer names. Consistency with a codebase's existing convention usually beats local correctness; change a convention deliberately and everywhere, not one file at a time.
Renaming churns history and breaks open branches. Rename when a name is actively misleading or about to spread, and leave merely imperfect names alone. If reviewers disagree about a local name and neither misleads, the author's choice wins; save debate for names that cross a boundary.
What to do next
- Pick one recent pull request and apply the prediction test: read only signatures and call sites, write down what you expect each function to do, and check.
- Grep for
data,info,Manager,Utilandflagin one service and rename the three worst offenders, one commit each. - Add units to every numeric configuration key and duration parameter that lacks one, starting with timeouts and intervals.
- Start a
GLOSSARY.mdwith the ten most-used domain terms, their definitions and the identifiers that implement them. - Audit your metric names against your metrics system's conventions and fix new metrics before dashboards bind to them.
- For any rename that crosses a boundary, write the expand and contract plan before the first commit.
- Add the mechanical rules (casing, boolean prefixes, banned suffixes) to your linter so review time goes on meaning.
Related reading: How to design an API for naming at the interface level, running a deprecation for the communication side of a public rename, working with legacy code for renaming inside code with weak tests, documentation for where a glossary fits, and reviewing AI-generated code for checking names an assistant chose.