Every name in a codebase is a tiny piece of documentation that is read far more often than it is written, and unlike a comment it travels with the value everywhere it goes: into call sites, stack traces, log lines, database schemas, dashboards and other teams' code. A good name lets a reader predict what a thing is and what it does without opening it. A bad one forces them to read the implementation, or worse, lets them predict something false.

This guide treats naming as an engineering skill with checkable rules: what a name must carry, how long it should be, conventions per kind of identifier, the boundary-crossing names that cost most to change, common anti-patterns, and safe renaming, ending with a checklist for code review.

Advertisement

Why names carry so much weight

A reader building a mental model of unfamiliar code works mostly from names. They scan a function, see retry_count and fetch_invoice, and form expectations: an integer, a network call that may fail. If those expectations hold, they can skip the bodies and reason at the level of the names. If they do not hold, every name becomes suspect and the reader has to verify everything, which is slow and error-prone.

Names also decide what is findable. Engineers search code and logs by name far more than they navigate by structure. A concept that is called account in one module, customer in another and tenant in a third cannot be traced with one search, and bugs hide in the seams between the three. Consistent names are what make grep, code search and log queries work.

Names are also read by tools: code search and AI coding assistants infer intent from identifiers, so a validate that also writes to the database misleads them too.

What a name must carry

A useful name answers the questions a reader will have at the point of use, and nothing more. Four kinds of information come up again and again:

  • Role: what the thing is for in this context, not what type it is. deadline beats date2; retry_budget beats n once the scope is more than a few lines.
  • Unit: for any quantity with a unit, put the unit in the name unless the type system enforces it. timeout_ms versus timeout_s is the difference between a working retry loop and one that waits sixteen minutes.
  • Kind: whether it is a single value, a collection, a mapping or a predicate. Plural nouns for collections, x_by_y for maps, is_ or has_ for booleans.
  • State or lifecycle when it matters: raw_payload versus validated_payload, draft_order versus placed_order. Naming the stage stops someone passing unvalidated data where validated data is required.

Length should be proportional to scope: the further a name is read from where it is defined, the more context it must carry by itself. A loop index used on the next line can be i; a public API field must be unambiguous to someone who has never seen your code. Long names in tiny scopes add noise that hides the logic.

Advertisement

Naming by kind

Each kind of identifier has a grammar that readers expect. Following it lets the reader infer behaviour from shape alone.

KindConventionGoodAvoid
BooleanPredicate, positive formis_active, has_access, can_retrynot_disabled, flag
Query functionNoun or get_, no side effectstotal_price(), get_user(id)process()
Command functionImperative verbsend_receipt(), cancel_order()receipt()
Expensive callVerb that signals costfetch_, load_, compute_get_ for a network call
CollectionPlural nounpending_jobsjob_list, data
Mappingvalues_by_keyusers_by_emailuser_map
Type or classSingular domain nounInvoice, RetryPolicyInvoiceManager, Utils
EventPast tense factOrderPlaced, PaymentFailedPlaceOrder

Two rules do most of the work. First, a function name should describe its effect at the level the caller cares about: if calling it can change state, send a message or cost a network round trip, the name should say so. Second, pairs should be symmetric. If you have start, its partner is stop, not end or halt; open goes with close, begin with end, min with max, first with last. Readers guess the second name from the first, so make the guess right.

Generic suffixes such as Manager, Helper and Util usually mean a type has no single responsibility yet; asking what it actually does often yields a better name.

Worked example: renaming a function until it explains itself

Here is a function of the kind found in many codebases. It works, it is short, and it is almost impossible to use correctly without reading it:

def process(d, f=False):
    res = []
    for x in d:
        t = x["ts"] - x["st"]
        if t > 30 and not (f and x["rt"]):
            res.append(x)
    return res

A caller has to guess everything. process says nothing about the effect. f is a boolean whose meaning is invisible at the call site. ts and st could be timestamps or states, and 30 has no unit.

Suppose the call sites show that f is set to True to skip retried jobs, and that the fields are epoch seconds. The rename then follows mechanically:

SLOW_JOB_THRESHOLD_SECONDS = 30

def find_slow_jobs(jobs: list[dict], skip_retries: bool = False) -> list[dict]:
    """Return jobs whose wall-clock run time exceeded the threshold."""
    slow_jobs = []
    for job in jobs:
        # input keys belong to the job record format, so bind them to clear locals
        started_epoch_s, finished_epoch_s, is_retry = job["st"], job["ts"], job["rt"]
        run_seconds = finished_epoch_s - started_epoch_s
        if run_seconds > SLOW_JOB_THRESHOLD_SECONDS and not (skip_retries and is_retry):
            slow_jobs.append(job)
    return slow_jobs

Every change answers one of the four questions. The function name states its effect and result. The flag keeps its polarity and default, so call sites change only by name. The threshold is a named constant with its unit. The record keys stay as they are, because they belong to a data format other code shares; well-named locals carry role and unit instead. The behaviour is unchanged, but the next reader can use it from the signature alone.

One word per concept, one concept per word

In large systems the bigger problem is vocabulary drift: one idea under several names, or one name for several ideas. Synonyms break search. Homonyms are worse: if account means a login identity in the auth service and a billing ledger in the payments service, someone will eventually join the two on the wrong key.

The fix is a glossary that lives in the repository, next to the code, and is changed in pull requests like code. Each entry gives the term, a one-sentence definition, the identifiers that implement it, and the terms it must not be confused with. Domain-driven design calls this a ubiquitous language: the words the business uses, used identically in conversations, tickets, code and schemas. Where two parts of the system genuinely mean different things by one word, qualify both names (auth_account, billing_account) rather than letting context disambiguate.

Abbreviations need a policy too: industry-standard ones (id, url) are fine, while team-invented ones (cust_ord_proc) are a private dialect every new hire must learn.

Names that cross boundaries

The cost of a bad name grows with every boundary it crossesLocal variableone functionPrivate helperone modulePublic APIcallers in repoWire and storagefields, columnsPublishedmetrics, eventsIDE renameIDE rename + grepdeprecate, aliasexpand / contractoften neverRename cost: seconds > minutes > a release > a migration > a dashboard, alert and partner rewriteSpend naming effort herenames that cross a boundary, names in a glossary,anything persisted or publishedKeep it cheap hereloop indices, lambdas, short local scopes:rename freely during reviewName quality is a budget decision: the further a name travels, the more it is worth getting right the first time.
Rename cost rises at each boundary a name crosses. Local names can be fixed in review; published names are close to permanent.

Inside one function, a bad name costs a few seconds to fix with an editor rename. Once a name crosses a boundary it acquires readers you do not control, and the cost of changing it jumps. Spend your naming effort where it is hardest to undo:

  • API fields and endpoints. Clients hard-code them. Pick one casing convention per API and keep it; prefer full words; include units in numeric fields (ttl_seconds); avoid booleans that will grow a third state, and use an enum such as status instead.
  • Database tables and columns. Queries, reports, ETL jobs and BI dashboards reference them by name, often outside your repository. Name columns for the fact they hold (placed_at, amount_cents), not for how they were first used.
  • Metrics and labels. Dashboards and alert rules bind to the exact string. Follow your metrics system's conventions; for Prometheus that means snake_case, base units such as seconds and bytes, and the _total suffix on counters.
  • Events and messages. Consumers you have never met subscribe by name. Never reuse an event name with new meaning; version the schema.
  • Configuration keys and feature flags. Name a flag for the behaviour when on, and record an owner and expiry.
# Counter: monotonically increasing, base unit, _total suffix
http_requests_total{method="POST", route="/orders", status_class="5xx"}

# Histogram in base units (seconds, not milliseconds)
http_request_duration_seconds_bucket{route="/orders", le="0.25"}

# Gauge: current value, unit in the name
queue_depth_messages{queue="billing"}
process_resident_memory_bytes

# Avoid
request_time_ms              # use base units: seconds
errors{user_id="83141"}      # unbounded label: one series per user

Anti-patterns and how to spot them in review

Anti-patternExampleWhy it hurtsFix
Lying namevalidate() that also savesReaders skip the side effectName the effect or split the function
Negated booleandisable_cache = FalseDouble negatives at every call siteuse_cache = True
Meaningless noundata, info, objCarries no roleName the domain thing
Numbered siblingsuser1, user2Hides the differencesender, recipient
Type encodingstrName, lstJobsDuplicates the type system and goes staleDrop the prefix
Missing unittimeout = 30Seconds or milliseconds?timeout_s = 30
Stale namelegacy_new_flowDescribes history, not behaviourRename to the current behaviour

In review, the cheapest check is to read only the signatures and call sites of a change and predict the behaviour. Wherever your prediction is wrong or uncertain, a name is the problem. Also search for the new identifier's obvious synonyms; if the concept already exists under another name, reuse it.

Linters can enforce casing, boolean prefixes and banned words like Util. They cannot judge whether a name tells the truth, so the prediction test stays a human job, and it matters most for generated code, whose plausible names may describe intent rather than behaviour.

Renaming safely

Inside one codebase, use the language-server rename, which understands scope, then search for the old name as a string: refactoring tools miss reflection, serialised field names, templates, SQL strings, config files and dashboards. Ship a pure rename as its own commit.

Across a boundary, a rename is a migration and needs the expand and contract pattern: add the new name alongside the old, move every reader and writer, verify, then remove the old one. For an API field that means returning both fields for a deprecation window and announcing the removal date. For a metric it means emitting both series until dashboards and alerts have moved. For a database column it looks like this:

-- Renaming a column without downtime: expand, migrate, contract.
-- Release 1 (expand): add the new name, keep the old one working.
ALTER TABLE orders ADD COLUMN placed_at timestamptz;
-- application writes BOTH created_ts and placed_at; reads prefer placed_at

-- Backfill in batches so locks stay short.
UPDATE orders SET placed_at = created_ts
WHERE placed_at IS NULL AND id BETWEEN :lo AND :hi;

-- Release 2: every reader uses placed_at; writes still go to both.
-- Verify: SELECT count(*) FROM orders WHERE placed_at IS DISTINCT FROM created_ts;  -- expect 0

-- Release 3 (contract): stop writing created_ts, then drop it.
ALTER TABLE orders DROP COLUMN created_ts;

Each release is independently deployable and reversible: if release 2 misbehaves, roll it back and the old column is still being written. A direct RENAME COLUMN is safe only when every reader deploys atomically with the schema change, which with replicas, analytics jobs and several services is rarely true.

Some names should never be renamed. A public metric with years of history may be cheaper to keep than to migrate; record the mismatch in the glossary so the next reader is warned.

Trade-offs

Precise names get long, and long names make expressions hard to read. Restructure rather than abbreviate: extract a well-named intermediate variable so each line holds fewer names. Consistency with a codebase's existing convention usually beats local correctness; change a convention deliberately and everywhere, not one file at a time.

Renaming churns history and breaks open branches. Rename when a name is actively misleading or about to spread, and leave merely imperfect names alone. If reviewers disagree about a local name and neither misleads, the author's choice wins; save debate for names that cross a boundary.

What to do next

  1. Pick one recent pull request and apply the prediction test: read only signatures and call sites, write down what you expect each function to do, and check.
  2. Grep for data, info, Manager, Util and flag in one service and rename the three worst offenders, one commit each.
  3. Add units to every numeric configuration key and duration parameter that lacks one, starting with timeouts and intervals.
  4. Start a GLOSSARY.md with the ten most-used domain terms, their definitions and the identifiers that implement them.
  5. Audit your metric names against your metrics system's conventions and fix new metrics before dashboards bind to them.
  6. For any rename that crosses a boundary, write the expand and contract plan before the first commit.
  7. Add the mechanical rules (casing, boolean prefixes, banned suffixes) to your linter so review time goes on meaning.

Related reading: How to design an API for naming at the interface level, running a deprecation for the communication side of a public rename, working with legacy code for renaming inside code with weak tests, documentation for where a glossary fits, and reviewing AI-generated code for checking names an assistant chose.

Key takeaway: A name is documentation that travels with the value. It should carry role, unit, kind and lifecycle stage where they matter, with length proportional to how far from its definition it is read. Follow the grammar readers expect for each kind of identifier, keep one word per concept in a shared glossary, and spend the most effort on names that cross boundaries, because API fields, columns, metrics and events become close to permanent. When a name has to change across a boundary, treat it as a migration and use expand and contract. Test names in review by predicting behaviour from signatures alone.