A gRPC system has two layers that are easy to blur together. The transport layer, HTTP/2 streams, deadlines, load balancing and keepalives, is covered in gRPC Architecture in Depth, and streaming semantics in gRPC bidi streaming architecture. This article is about the other layer: the Protobuf contract. In practice it decides whether fifty services can evolve independently for years or whether every schema change becomes a coordinated release.

We will encode a message by hand so the compatibility rules stop being folklore. Then we will look at field presence and why it matters for updates, enforce evolution in CI with buf, lay out packages for versioning, implement partial updates with FieldMask, return structured errors, and walk a field addition across a running fleet. Facts about the encoding and compatibility rules come from the proto3 language guide and the encoding documentation on protobuf.dev; buf behaviour comes from the buf documentation.

Two compilations of one contract

Build time.proto filesone repo or modulebuf lintstyle, namingbuf breakingvs main branchcodegenpluginsRun timeclient codebuilds messageserializestubtag-value bytes5-byte prefixHTTP/2 streamDATA framesservergenerated skeletontrailersgrpc-status + detailsrich error back to clientThe schema is checked before code exists; at run time only field numbers and wire types travel.
The contract layer. Schemas are linted and checked for breaking changes before code generation; on the wire only field numbers, wire types and values travel, and errors come back in trailers.

The key architectural idea is that a .proto file is compiled twice. At build time it becomes generated classes in every language you use. At run time, the bytes that cross the network carry no names at all, only field numbers, wire types and values. Every gRPC message is a Protobuf payload preceded by a 5-byte prefix: one byte saying whether it is compressed and four bytes of length. Everything about compatibility follows from the fact that a reader with an old schema and a writer with a new one only need to agree on numbers and types, never on names or code.

Encoding a message by hand

Each field on the wire is a tag followed by a value. The tag is a varint holding (field_number << 3) | wire_type. Varints store 7 bits per byte, least significant group first, with the high bit set on every byte except the last. There are six wire types, and only four matter in proto3:

Wire typeIdUsed for
VARINT0int32, int64, uint32, uint64, sint32, sint64, bool, enum
I641fixed64, sfixed64, double
LEN2string, bytes, embedded messages, packed repeated fields
I325fixed32, sfixed32, float

Take this message and encode one instance by hand:

syntax = "proto3";
package acme.orders.v1;

message Order {
  int64 id = 1;
  string sku = 2;
  int32 quantity = 3;
  repeated int32 tags = 4;
}

// instance: id = 150, sku = "ab-7", quantity = 0, tags = [1, 2, 300]
FieldTag byteValue bytesWhy
id = 1500896 01(1 << 3) | 0 = 8; 150 = 0x96 with continuation, then 0x01
sku = "ab-7"1204 61 62 2d 37(2 << 3) | 2 = 18; length 4, then UTF-8
quantity = 0nonenoneimplicit presence: a zero is not written
tags2204 01 02 ac 02packed by default: one LEN record; 300 = ac 02

The whole message is 15 bytes: 08 96 01 12 04 61 62 2d 37 22 04 01 02 ac 02. Several design rules fall out of this. Field numbers 1 to 15 fit in a one-byte tag, so give them to the fields set most often. Negative int32 values are sign-extended to 64 bits and always take ten bytes, while sint32 uses zigzag encoding and stores -1 in one byte, so use sint types for fields that are often negative. And because the reader only sees numbers, renaming a field is invisible on the binary wire, while reusing a number is catastrophic. A decoder small enough to read makes this tangible:

def read_varint(buf: bytes, i: int) -> tuple[int, int]:
    shift = result = 0
    while True:
        b = buf[i]; i += 1
        result |= (b & 0x7F) << shift
        if not b & 0x80:
            return result, i
        shift += 7

def walk(buf: bytes):
    i = 0
    while i < len(buf):
        tag, i = read_varint(buf, i)
        field, wtype = tag >> 3, tag & 7
        if wtype == 0:
            value, i = read_varint(buf, i)
        elif wtype == 2:
            n, i = read_varint(buf, i)
            value, i = buf[i:i + n], i + n
        elif wtype == 1:
            value, i = buf[i:i + 8], i + 8
        elif wtype == 5:
            value, i = buf[i:i + 4], i + 4
        else:
            raise ValueError(f"unsupported wire type {wtype}")
        yield field, wtype, value

msg = bytes.fromhex("08 96 01 12 04 61 62 2d 37 22 04 01 02 ac 02")
for f in walk(msg):
    print(f)   # (1, 0, 150), (2, 2, b'ab-7'), (4, 2, b'\x01\x02\xac\x02')

Notice that the decoder never needs the schema to skip a field. That is why an old reader can step over a field it does not know, and why proto3 runtimes can keep such unknown fields and write them back out on re-serialisation.

Presence: zero versus absent

In the example, quantity = 0 produced no bytes, so a reader cannot tell whether the writer set it to zero or never set it. That is implicit presence, the proto3 default for scalar fields. For many fields it is fine. For fields where zero is meaningful and different from absent, such as a discount percentage, a retry count or a boolean feature flag, it causes bugs, especially in updates where absent means leave unchanged. Mark those fields optional, which gives them explicit presence and generated has-methods. Wrapping scalars in google.protobuf.Int32Value worked because message fields always have presence; proto3 optional makes that workaround unnecessary.

Evolution rules enforced in CI

The compatibility rules are documented in the proto3 guide. int32, uint32, int64, uint64 and bool are wire-compatible with each other, though values can be truncated. sint32 and sint64 are compatible only with each other. fixed32 pairs with sfixed32 and fixed64 with sfixed64. string and bytes are compatible if the bytes are valid UTF-8. Moving fields into an existing oneof is not safe, and field numbers must never be reused. The useful architectural move is to stop relying on reviewers to remember these rules and make a tool reject violations. buf does this by comparing the current schemas against a previous version:

# buf.yaml
version: v2
modules:
  - path: proto
lint:
  use:
    - STANDARD
breaking:
  use:
    - FILE
# CI step on every pull request
buf lint
buf breaking --against '.git#branch=main'

buf groups its breaking rules into four categories, from strictest to most lenient: FILE and PACKAGE catch changes that break generated source code, per file or per package; WIRE_JSON catches changes that break the binary or JSON encoding; WIRE catches only binary breakage. Choose deliberately. If anything consumes your messages as JSON, through gRPC transcoding, gRPC-Web tooling, logs parsed downstream or a REST gateway, WIRE is not enough, because renaming a field is wire-safe but changes the JSON key. If other teams compile against your generated code, use FILE or PACKAGE so moving a message between files is caught too. Removing a field the right way, by deleting it and adding reserved 3; reserved "quantity";, passes these checks and also stops anyone reusing the number or name later.

Package layout and versioning

Put the major version in the package name, as in acme.orders.v1, and keep the directory structure the same as the package. A breaking redesign becomes a new package, acme.orders.v2, served alongside v1 from the same process until clients move. That keeps breaking changes explicit and lets the breaking-change check stay strict inside a version. The trade-offs between version carriers are discussed in API Versioning Strategies.

Give every RPC its own request and response messages, even if they start with a single field. rpc GetOrder(GetOrderRequest) returns (Order) can later gain a read mask without changing its signature, whereas a method that takes a bare shared message is frozen.

Partial updates with FieldMask

A classic update API takes a whole Order and overwrites the stored one, which silently wipes any field the client did not know about or did not send. The Protobuf answer is google.protobuf.FieldMask, a list of field paths the client intends to change:

import "google/protobuf/field_mask.proto";

message UpdateOrderRequest {
  Order order = 1;
  google.protobuf.FieldMask update_mask = 2;   // e.g. paths: ["quantity", "tags"]
}
import grpc

UPDATABLE = {"sku", "quantity", "tags"}

def UpdateOrder(self, request, context):
    stored = self.repo.get(request.order.id)
    paths = list(request.update_mask.paths)
    if not paths:
        context.abort(grpc.StatusCode.INVALID_ARGUMENT, "update_mask is required")
    for path in paths:
        if path not in UPDATABLE:
            context.abort(grpc.StatusCode.INVALID_ARGUMENT, f"field {path} is not updatable")
        if path == "tags":
            del stored.tags[:]
            stored.tags.extend(request.order.tags)
        else:
            setattr(stored, path, getattr(request.order, path))
    self.repo.put(stored)
    return stored

This handles flat fields only, which is often all you need. Rejecting unknown paths matters: silently ignoring a path is how a client believes it changed something it did not. With a mask, a zero value is unambiguous, because the mask says quantity was meant to change, so the field mask and optional presence solve the same problem from two sides. Read masks work the same way in reverse and let list calls return only the fields a client needs.

Errors that carry structure

A gRPC status is a code plus a string. For a validation failure on three fields, a string is a poor contract: clients end up parsing English. The richer model is google.rpc.Status, which adds a list of typed details such as BadRequest with per-field violations, ErrorInfo with a machine-readable reason and domain, and RetryInfo with a suggested delay. It travels in the grpc-status-details-bin trailer alongside the normal status, so clients that do not understand it still see the code and message. In Python, the grpcio-status package converts between the two:

import grpc
from google.rpc import code_pb2, error_details_pb2, status_pb2
from google.protobuf import any_pb2
from grpc_status import rpc_status

def invalid(context, violations):
    detail = any_pb2.Any()
    detail.Pack(error_details_pb2.BadRequest(field_violations=[
        error_details_pb2.BadRequest.FieldViolation(field=f, description=d) for f, d in violations
    ]))
    context.abort_with_status(rpc_status.to_status(status_pb2.Status(
        code=code_pb2.INVALID_ARGUMENT, message="order failed validation", details=[detail])))

# client side
try:
    stub.UpdateOrder(req)
except grpc.RpcError as call:
    status = rpc_status.from_call(call)
    for d in (status.details if status else []):
        if d.Is(error_details_pb2.BadRequest.DESCRIPTOR):
            br = error_details_pb2.BadRequest(); d.Unpack(br)

Agree on the code choice as part of the contract: INVALID_ARGUMENT for bad input that will never succeed, FAILED_PRECONDITION for state conflicts, UNAVAILABLE for transient failures. Client retry policies key off these codes, so a server that returns UNAVAILABLE for a validation error will be hammered with retries.

The codegen pipeline

Code generation is where organisations accumulate accidental complexity. protoc and buf both run plugins that turn a parsed schema into code, one plugin for messages and another for gRPC stubs in each language. With buf the plugin list lives in a config file, so every team generates with the same versions:

# buf.gen.yaml
version: v2
plugins:
  - remote: buf.build/protocolbuffers/go
    out: gen/go
    opt: paths=source_relative
  - remote: buf.build/grpc/go
    out: gen/go
    opt: paths=source_relative

Decide once whether generated code is committed or produced at build time, and either way pin plugin versions, because a plugin upgrade can change generated APIs even when no schema changed. Also produce a descriptor set, the compiled schema as bytes. It is what server reflection, JSON transcoding gateways and generic debugging tools consume, and keeping one per release lets you decode archived payloads years later.

Worked example: adding a field across a fleet

Suppose the orders team adds optional string gift_note = 5; to Order, which flows client to edge service to orders service to a fulfilment consumer. The safe order of operations is:

  1. Merge the schema change. buf breaking passes because a new field number is additive.
  2. Deploy the orders service and fulfilment consumer first, coded to treat a missing gift_note as no note. Old clients keep working because they never send field 5.
  3. Check every hop in between. The edge service, built against the old schema, receives field 5 as an unknown field. If it forwards the parsed message, the runtime keeps the unknown field and re-emits it. If it copies data into its own struct or round-trips through JSON, field 5 is silently dropped. That is the most common real-world way an additive change loses data.
  4. Deploy clients that set the field.
  5. To remove a field later, reverse it: stop writing, wait until no reader depends on it, delete it and reserve its number and name.

Failure modes

  • Number reuse. Deleting field 3 and later adding a different field 3 makes old data decode into the new field. Always reserve.
  • Changing a type across groups. int32 to sint32 is not compatible even though both are integers, because zigzag changes the meaning of the bits.
  • Presence-blind updates. Without optional or a field mask, set to zero and leave unchanged look the same.
  • Lossy middle hops that convert messages to structs or JSON and drop unknown fields.
  • JSON breakage under a WIRE-only check, from a rename that the binary wire never notices.

Trade-offs

Protobuf with gRPC buys compact payloads, generated clients and machine-checked evolution. It costs readability on the wire, a build toolchain and friction at the browser edge. For public HTTP APIs consumed by many unknown clients, a resource-oriented JSON design, as in REST API Design, in depth, is often the better external face, with Protobuf behind it. Inside a mesh, where sidecars already understand HTTP/2 (see Service Mesh Architecture), the contract-first model pays for itself quickly.

What to do next

  1. Put all .proto files in one module with a buf.yaml, and run buf lint and buf breaking against main on every pull request, with the breaking category chosen on purpose.
  2. Audit existing schemas for deleted fields without a reserved line, and add them.
  3. Mark scalar fields where zero differs from absent as optional.
  4. Give every RPC its own request and response messages and move the major version into the package name.
  5. Add an update_mask to every update RPC and reject unknown paths.
  6. Return google.rpc.Status details for validation errors and agree on which codes clients may retry.
  7. Trace one message through every hop and confirm no service drops unknown fields.
  8. Pin codegen plugin versions and archive a descriptor set with each release.
Key takeaway: On the wire a Protobuf message is only field numbers, wire types and values, so compatibility means never reusing numbers and only changing types within compatible groups. Make a tool enforce that on every pull request, use optional and field masks so zero and absent stay distinct, return typed error details, and check that every hop in a call chain preserves unknown fields.