A schema registry is easy to install and easy to misuse. The schema registry architecture article explains the machinery: every record carries a magic byte and a four-byte schema id, the registry stores versions under subjects, and a compatibility checker decides whether a new version may be added. This article is about the decisions that machinery leaves to you, the ones that decide whether schema evolution stays boring for years or becomes a recurring incident.

There are six: subject naming, who registers schemas, compatibility per subject, safe evolution, genuinely breaking changes, and running the registry so its outage is a non-event. A worked evolution of a real event and a checklist follow.

Advertisement

The one idea behind every pattern

A topic is a contract between teams that deploy independently. A producer written today will emit records that a consumer written two years ago must read, and a consumer deployed tomorrow may replay records written two years ago. The registry's job is to make that contract machine-checkable: before a schema can be used, someone proves it can coexist with the versions already in the log.

Every pattern below follows from one rule: the compatibility check must run before a schema reaches production, against the full set of versions that readers might meet. Most registry incidents break one half of that rule: a producer registered at runtime, so the check failed in production, or the check compared too few versions and a replaying consumer failed later.

Pattern 1: choose a subject naming strategy deliberately

The serializer maps each record to a subject, and the subject is the unit of versioning and compatibility. Confluent serializers support three strategies, set with key.subject.name.strategy and value.subject.name.strategy.

StrategySubject nameUse whenCost
TopicNameStrategy (default)<topic>-valueOne event type per topicCannot mix record types on a topic without a union
RecordNameStrategyfully qualified record nameSame event type shared across many topicsNo per-topic compatibility; any topic can carry any type
TopicRecordNameStrategy<topic>-<record name>Several event types on one topic, versioned per typeTopic no longer has one schema; consumers must dispatch on type

The default is right for most topics, and it gives you the strongest guarantee: one subject per topic, so a consumer knows exactly which family of schemas it can meet. Reach for something else only when you need several event types on one topic, which is common for event-sourced aggregates where ordering across types matters: an order's OrderPlaced, OrderPaid and OrderShipped must share a partition key and a topic to stay ordered.

For that case there are two workable designs. TopicRecordNameStrategy versions each type independently and is simple to evolve, but nothing stops a producer from putting an unrelated type on the topic. The alternative keeps TopicNameStrategy and makes the topic's schema a union whose branches are schema references to the individual event schemas. That keeps one checked contract per topic and makes a new event type a reviewed change. Pick one design per organisation and write it down.

Advertisement

Pattern 2: CI registers schemas, producers never do

Serializers register schemas automatically by default, which moves the compatibility check to the moment a new producer build first sends a record, so a failure is an exception in a running service.

The pattern is to treat schemas as code. They live in a repository next to the service that owns the topic, a pull request runs the compatibility check against the registry, and only the merge to the main branch registers the new version. Producers run with auto.register.schemas=false, so a schema that was not registered through the pipeline makes the producer fail at startup, where a deploy health check catches it, rather than corrupting the topic.

Schemas change through a pipeline; producers only look them upSchema in git.avsc / .protoPRCI: compatibilityPOST /compatibility/...mergeCI: registerPOST /subjects/.../versionsRegistryid = 42A failed compatibility check blocks the merge, not a production deploy.Producerauto.register = falselookup idRegistrycached after 1st callmagic byte + id 42 + bytesKafka topicordersConsumerreader schema v3fetch writer schema 42Producers fail fast if their schema is not already registered: an unreviewed schema never reaches the topic.Consumers resolve every record's writer schema by id and project it onto the reader schema they were compiled with.Registry outage: cached ids keep working; new producer instances and unseen ids fail until it returns.
CI-owned registration. The compatibility check happens on the pull request; production producers only look up ids, and the registry's cache makes it a startup dependency rather than a per-record one.
#!/usr/bin/env bash
# CI step: fail the build if the schema would break the subject's compatibility rule.
set -euo pipefail
REG=https://schema-registry.internal:8081
SUBJECT=orders-value
SCHEMA_JSON=$(jq -Rs '{schemaType: "AVRO", schema: .}' < schemas/order_placed.avsc)

curl -sf -X POST -H "Content-Type: application/vnd.schemaregistry.v1+json" \
  --data "$SCHEMA_JSON" \
  "$REG/compatibility/subjects/$SUBJECT/versions/latest?verbose=true" \
  | tee /tmp/compat.json
jq -e '.is_compatible == true' /tmp/compat.json

# On merge to main only: register, which returns the global id.
if [ "${CI_BRANCH:-}" = "main" ]; then
  curl -sf -X POST -H "Content-Type: application/vnd.schemaregistry.v1+json" \
    --data "$SCHEMA_JSON" "$REG/subjects/$SUBJECT/versions"
fi

On the producer, leave use.latest.version off so each build serializes with the schema it was compiled against and fails fast if that schema is unknown.

from confluent_kafka import SerializingProducer
from confluent_kafka.schema_registry import SchemaRegistryClient
from confluent_kafka.schema_registry.avro import AvroSerializer

registry = SchemaRegistryClient({"url": "https://schema-registry.internal:8081"})
serializer = AvroSerializer(
    registry,
    open("schemas/order_placed.avsc").read(),
    conf={
        "auto.register.schemas": False,   # CI registers; producers never do
        "use.latest.version": False,      # serialize with the schema compiled into this build
    },
)
producer = SerializingProducer({
    "bootstrap.servers": "kafka:9092",
    "value.serializer": serializer,
    "enable.idempotence": True,
})

Pattern 3: pick compatibility per subject, and know when transitive matters

The compatibility level answers one question: which side may upgrade first. The default, BACKWARD, means a consumer using the new schema can read data written with the previous one, so you upgrade consumers first and producers after.

LevelGuaranteeUpgrade orderTypical safe changes (Avro)
BACKWARDNew readers read previous dataConsumers firstAdd field with default, delete field
FORWARDPrevious readers read new dataProducers firstDelete field with default, add field
FULLBoth directions, previous versionEither orderAdd or delete only fields with defaults
*_TRANSITIVESame, against all versions, not just the lastAs aboveAs above, checked against history
NONENo checkCoordinatedAnything; for scratch subjects only

The non-transitive levels check the new version against the latest one only. That is enough when readers only ever see recent data, but Kafka topics are not always short-lived. Consider a subject at BACKWARD. Version 2 deletes a field region that had no default, which is backward compatible because a v2 reader simply ignores it. Version 3 adds region back with a different type and a default. Against v2 that is fine. Against v1 data it is not: a v3 reader meeting a v1 record finds a region field of the wrong type. A consumer replaying from the beginning of the topic crashes on the oldest records.

So use the transitive level whenever readers can meet old data: compacted topics, which keep the latest value for every key forever, as described in the log compaction article; topics with long or infinite retention; and topics replayed to rebuild state, such as Kafka Streams changelogs. Use FULL_TRANSITIVE for topics consumed by many teams whose deploy order you cannot control. Set the level per subject with PUT /config/<subject> and keep the global default at BACKWARD or stricter.

Worked example: evolving OrderPlaced

An orders service publishes OrderPlaced to the orders topic, which is compacted by order id and consumed by billing, analytics and a fraud model. Because it is compacted, the subject orders-value runs at FULL_TRANSITIVE.

// v1
{"type": "record", "name": "OrderPlaced", "namespace": "com.shop.orders",
 "fields": [
   {"name": "order_id", "type": "string"},
   {"name": "amount_cents", "type": "int"},
   {"name": "currency", "type": "string"}]}

// v2: add an optional field WITH a default -> backward and forward compatible
   {"name": "coupon_code", "type": ["null", "string"], "default": null}

// v3: widen int -> long (Avro promotes int to long when reading old data)
   {"name": "amount_cents", "type": "long"}

// v4 (rejected under BACKWARD): make currency an enum.
//    Old records hold arbitrary strings the enum cannot represent.

Version 2 adds an optional coupon code. The field is a union with null and a default of null, so old readers ignore it and new readers fill the default when reading v1 records. It passes in both directions, and producers and consumers can deploy in any order.

Version 3 widens the amount from int to long because a large wholesale order overflowed. Avro's schema resolution promotes an int written by v1 or v2 into a long when a v3 reader reads it, so v3 readers are fine. The reverse is not true: a v2 reader cannot read a long. Under FULL_TRANSITIVE the registry rejects this change, which is correct, because billing still runs v2. The team either relaxes the subject to BACKWARD_TRANSITIVE and upgrades every consumer first, or treats it as breaking.

Version 4 tries to turn currency into an enum. Old records contain strings the enum may not include, so the change is rejected under every level except NONE. It is breaking, and it needs the migration pattern below.

Pattern 4: migrate breaking changes through a new topic

When a change cannot be made compatible, do not lower the subject to NONE. Make the break explicit:

  1. Create a new topic, for example orders.v2, with its own subject and the new schema.
  2. Dual-write from the producer, or run a small stream processor that converts old records into the new topic and can backfill history.
  3. Move consumers one at a time to the new topic, each at a known offset or timestamp, and verify outputs match.
  4. Stop writing the old topic once its last consumer has moved, keep it for its retention period, then delete it.

It costs a period of double writing, but every consumer always reads one schema family and the migration can pause or roll back at any step. For change-data-capture topics the same pattern applies, but the break often originates in a database migration; the Debezium CDC article covers how schema changes surface in change events.

Pattern 5: share types through schema references

Types such as Money, Address or a tracing header appear in dozens of events. Copying their definition into every schema guarantees drift. The registry supports references for Avro, Protobuf and JSON Schema: a schema declares that it depends on another subject at a specific version, and the registry resolves the reference when checking and serving it.

References pin a version, which is a feature. Updating Money registers a new version of the com.shop.Money subject but changes no event until each event schema is updated to point at it, and each of those updates passes its own compatibility check. Keep shared types small, strictly checked and owned in one repository.

Pattern 6: know your format&#x27;s traps

Avro needs the writer's schema to decode anything, which is why it pairs so naturally with a registry; its traps are fields added without defaults and renames, which must be done with reader-side aliases rather than by changing the name. Protobuf decodes by field number, so the rules are different: never reuse or renumber a field, mark deleted numbers and names as reserved, and remember that for a plain, non-optional proto3 scalar, unset is indistinguishable from its zero value, so adding a field whose default meaning is not zero is a semantic break even when the registry accepts it.

JSON Schema is the hardest to check because its content model is a choice. With an open model, where additional properties are allowed, adding a property is harmless to old readers. With a closed model, "additionalProperties": false, an old reader rejects any new property, so additions become forward-incompatible. Decide the model before the first topic ships. In every format, the registry checks structure, not meaning: changing amount from cents to euros passes every check.

Running the registry

The registry is on the read path only on a cache miss. Serializers cache schema-to-id lookups and deserializers cache id-to-schema lookups for the life of the process, so steady-state traffic barely touches it. Confluent's implementation stores its state in a compacted Kafka topic, _schemas by default, and elects one primary instance to handle writes while the others serve reads and forward writes. Back that topic up and protect it from deletion like any database.

An outage therefore has a precise blast radius: running producers and consumers continue, while new processes that need a lookup, consumers meeting an id they have not cached and every CI registration fail. Run at least three instances and alert on client-side lookup errors, not only the health endpoint. Across regions, schema ids must mean the same thing on every cluster that holds the same records. Replicate one registry's state to the others, or run one primary registry for all regions; never run independent registries that assign ids for topics that are mirrored between clusters.

Failure modes

SymptomCauseFix
Producer throws on first send after deployRuntime registration failed the compatibility checkRegister in CI; auto.register.schemas=false
Replaying consumer crashes on old recordsNon-transitive level on a compacted or long-retention topicUse a *_TRANSITIVE level
Subject set to NONE to unblock a releaseBreaking change forced throughNew topic migration; restrict config writes to CI
Unknown magic byte errorsA producer wrote without the registry serializerEnforce serializer use; validate on ingest
Same id decodes differently in two regionsIndependent registries behind mirrored topicsOne source of ids, replicated
Values silently wrong after a clean evolutionSemantic change the checker cannot seeReview units and meaning in schema PRs; new field name for new meaning

What to do next

  1. List every subject and its compatibility level with GET /subjects and GET /config/<subject>; flag compacted or long-retention topics that are not transitive.
  2. Turn off auto.register.schemas in production producers and move registration into CI using the script above.
  3. Restrict registration and configuration writes to the CI identity.
  4. Write down your subject naming strategy and how multi-event topics are modelled.
  5. Pick a breaking-change runbook based on the new-topic migration, and rehearse it once on a test topic.
  6. Add client-side alerts on schema lookup failures and a registry outage to your next disaster-recovery drill.
  7. If you are new to the platform underneath, read the Kafka overview to see how topics, partitions and retention interact with these choices.
Key takeaway: A schema registry turns a topic into a checked contract, but only if the check runs early and against the right history. Keep the default topic naming unless you need multi-event topics, register schemas from CI and never from producers, use transitive compatibility wherever old data can be replayed, ship breaking changes as new topics instead of disabling checks, share types through references, learn your format's traps, and run the registry as the small but critical database it is.