WebSocket and gRPC bidirectional streaming both give you a long-lived channel where either side can send a message at any time. That shared shape is why teams argue about them, and also why the argument is usually framed wrongly. They sit at different layers. WebSocket is a framing protocol over a single TCP connection: it moves opaque messages and promises nothing about what they mean. gRPC bidi is a remote procedure call whose request and response are both streams of typed messages, carried on one HTTP/2 stream, with a contract, deadlines, status codes and flow control already decided for you.

This article compares them the way an implementer meets them: the bytes on the wire, slow receivers, what a browser can open, proxies, and what you build when a connection drops. It then writes one small collaborative-editing service both ways, shows the common hybrid, and ends with a decision guide and checklist.

Advertisement

Two stacks for one job

Same job, two stacks: a long-lived, two-way message channelWebSocket (RFC 6455)gRPC bidirectional streamingYour messagesJSON / binary, your own schemaWebSocket framesopcode, FIN, mask, lengthOne upgraded TCP connectionHTTP/1.1 Upgrade (or RFC 8441/9220)TLS + TCPone socket per channelProtobuf messagesschema in a .proto contractLength-prefixed messages1-byte flag + 4-byte lengthOne HTTP/2 streamheaders, DATA, trailers, windowsTLS + TCPmany streams share one socketYou buildacks, resume, status, heartbeatsYou getdeadlines, status codes, flow controlNeither keeps a replay log: after a disconnect, resuming is always your protocol.Browsers speak WebSocket natively; they cannot open a gRPC bidi stream today.
WebSocket gives you frames and leaves meaning to you; gRPC gives you a typed call on an HTTP/2 stream with status, deadlines and per-stream windows.

A WebSocket starts as an HTTP request. The client sends GET with Upgrade: websocket and a random Sec-WebSocket-Key; the server answers 101 Switching Protocols with a derived Sec-WebSocket-Accept, and from then on the TCP connection carries WebSocket frames instead of HTTP. Each frame has a FIN bit, an opcode (text, binary, continuation, close, ping, pong), a mask bit and a payload length; client-to-server frames are masked, which protects intermediaries rather than encrypting anything. RFC 8441 and RFC 9220 bootstrap WebSockets over HTTP/2 and HTTP/3 streams, but most deployments still upgrade over HTTP/1.1: one TCP connection per WebSocket.

A gRPC bidi call starts as an HTTP/2 request: a HEADERS frame with :path /collab.v1.DocSync/Sync, content-type: application/grpc and optional grpc-timeout. Both directions then carry DATA frames containing length-prefixed messages: one byte saying whether the message is compressed, four bytes of big-endian length, then the protobuf bytes. The call ends with a trailing HEADERS frame carrying grpc-status and grpc-message. Because it is one HTTP/2 stream, many calls share a single TCP connection, each with its own flow-control window.

What each gives you, and what you build

ConcernWebSocketgRPC bidi
Message schemaNone; you pick JSON, MessagePack, protobufProtobuf contract, generated stubs
Message boundariesYes, frames reassemble to messagesYes, length prefix
MultiplexingOne channel per connection (HTTP/1.1)Many streams per connection
Flow controlTCP only; no per-message signalHTTP/2 stream and connection windows
End-of-call statusClose code (e.g. 1000, 1011) and reasongrpc-status code plus message in trailers
Deadlines and cancelBuild itgrpc-timeout, context cancellation
MetadataOnly on the upgrade requestHeaders and trailers per call
Browser supportNative everywhereNo bidi from browsers (see below)
Half-closeNo; close is for both directionsYes; client can end its side and keep reading

WebSocket is small and universal, which is why you rebuild message types, acknowledgements, error codes and timeouts on top of it. gRPC hands you those as shared conventions, at the cost of HTTP/2 on the whole path and a code generator in your build.

Advertisement

Flow control and the slow consumer

Every streaming system eventually meets a receiver that reads more slowly than the sender writes: a phone on a bad network, a background tab, a service stuck in garbage collection. What matters is where the excess goes.

In gRPC, HTTP/2 gives each stream a receive window, a budget of bytes the sender may transmit before the receiver grants more with WINDOW_UPDATE. When the reader stops consuming, the window drains, Send blocks, and pressure propagates back to the producer. A connection-level window is shared by all streams, so one stalled stream can starve its neighbours. The HTTP/2 deep dive has the window arithmetic.

WebSocket has no such signal. The server's send writes into a user-space buffer; once the kernel socket buffer is full that buffer just grows. In both the browser API and the Node ws package the only visible symptom is bufferedAmount, which you must poll and act on. On the receiving side, the browser's onmessage fires for every message whether or not your code keeps up. So a WebSocket service needs an explicit policy: drop, coalesce, or disconnect with a resume point. The Node example later disconnects with close code 1013 (try again later) once a per-connection high-water mark is crossed. The backpressure article compares those policies in depth.

The browser reality

If any endpoint is a web page, this section usually decides the question. Every browser ships the WebSocket API. No browser lets JavaScript open a gRPC bidi stream. The grpc-web project's README is explicit that client-side and bi-directional streaming are not currently supported; it handles unary calls and server streaming through a proxy that translates to real gRPC. The deeper reason is that the browser's fetch cannot run a full-duplex request and response at the same time: where request-body streaming exists it is half-duplex, so the whole request must be sent before the response is read. That rules out a true bidi stream regardless of library.

WebTransport (streams over HTTP/3) may change this, but support is still uneven. For browsers today the practical choices are WebSocket, or server-sent events down plus ordinary requests up.

Proxies, load balancers and keepalives

Both protocols produce long-lived connections, which break load-balancer assumptions. A WebSocket pins one client to one backend for the life of the socket; an L7 proxy must understand the upgrade and not apply its normal request timeout. gRPC needs HTTP/2 end to end; an L4 balancer works but balances connections, not calls, so every stream on one channel lands on the same backend. That is fine for bidi, where the stream is the unit of affinity anyway, and a trap for unary traffic sharing the channel.

Idle timeouts sit on every hop: cloud load balancers, NAT gateways, corporate proxies, and a quiet stream is cut silently unless something sends bytes. WebSocket has ping and pong frames, but the browser API does not expose them, so the server pings and clients that need liveness send application heartbeats. gRPC has HTTP/2 PING keepalive; servers enforce a minimum interval and close over-eager clients with GOAWAY, so agree settings on both sides. For deploys, both need draining: gRPC servers send GOAWAY and can cap connection age; WebSocket servers close with 1001 (going away) and clients reconnect with jitter. Load balancing WebSockets covers draining and reconnect storms in depth.

The same service, written both ways

The worked example is a collaborative document. Each client sends edits, each with a client sequence number; the server appends them to a per-document log, acknowledges each with the server sequence it was assigned, and streams every other client's edits. On reconnect the client says which server sequence it last applied, and the server replays from there. That resume rule is the important part: it is what makes a dropped connection harmless, and neither protocol provides it.

In gRPC the contract is a proto file. The oneof gives you typed message kinds, and adding a field later is compatible by default:

syntax = "proto3";
package collab.v1;

service DocSync {
  // Client streams edits up; server streams acks and everyone else's edits down.
  rpc Sync(stream ClientFrame) returns (stream ServerFrame);
}

message ClientFrame {
  oneof kind {
    Hello hello = 1;  // must be first: which document, resume point
    Edit  edit  = 2;
  }
}
message Hello      { string doc_id = 1; uint64 resume_after = 2; }
message Edit       { uint64 client_seq = 1; bytes op = 2; }
message ServerFrame {
  oneof kind {
    Ack        ack    = 1;
    RemoteEdit remote = 2;
  }
}
message Ack        { uint64 client_seq = 1; uint64 server_seq = 2; }
message RemoteEdit { uint64 server_seq = 1; bytes op = 2; }

The server handler runs the two directions on separate goroutines because a stream allows one concurrent sender and one concurrent receiver; doing both from one loop deadlocks the first time the client stops reading. gRPC bidi streaming architecture walks through half-close and cancellation in detail.

// Go, abridged. One goroutine owns Send, one owns Recv.
func (s *server) Sync(stream pb.DocSync_SyncServer) error {
    first, err := stream.Recv()
    if err != nil {
        return err
    }
    hello := first.GetHello()
    if hello == nil {
        return status.Error(codes.InvalidArgument, "first frame must be Hello")
    }
    sub := s.log.Subscribe(hello.DocId, hello.ResumeAfter) // replay, then tail
    defer sub.Close()

    errc := make(chan error, 2)
    go func() { // sender: Send blocks when the HTTP/2 window is used up
        for frame := range sub.C {
            if err := stream.Send(frame); err != nil {
                errc <- err
                return
            }
        }
        errc <- stream.Context().Err()
    }()
    go func() { // receiver
        for {
            f, err := stream.Recv()
            if err == io.EOF { // client half-closed: finish cleanly
                errc <- nil
                return
            }
            if err != nil {
                errc <- err
                return
            }
            seq := s.log.Append(hello.DocId, f.GetEdit())
            sub.Ack(f.GetEdit().ClientSeq, seq) // ack travels via sub.C
        }
    }()
    return <-errc // returning ends the RPC and cancels the other goroutine
}

The WebSocket version implements the same protocol by hand. Notice what had to be added: a type field and validation instead of a schema, a hello-first rule enforced with close code 1008 (policy violation), an explicit high-water mark because nothing pushes back, and a ping loop because silence is indistinguishable from death.

// Node.js with the "ws" package. Same protocol, rebuilt by hand.
import { WebSocketServer } from "ws";

const wss = new WebSocketServer({ port: 8080, maxPayload: 1 << 20 });
const HIGH_WATER = 4 << 20; // bytes queued in user space before we act

wss.on("connection", (ws) => {
  let hello = null;
  ws.isAlive = true;
  ws.on("pong", () => { ws.isAlive = true; });

  ws.on("message", (data) => {
    let msg;                               // no contract: validate everything
    try { msg = JSON.parse(data); } catch { return ws.close(1007, "bad JSON"); }
    if (!hello) {
      if (msg.type !== "hello") return ws.close(1008, "hello first");
      hello = msg;
      return replayAndTail(ws, msg.docId, msg.resumeAfter);
    }
    if (msg.type === "edit") {
      const seq = log.append(hello.docId, msg.op);
      send(ws, { type: "ack", clientSeq: msg.clientSeq, serverSeq: seq });
    }
  });
});

function send(ws, obj) {
  if (ws.bufferedAmount > HIGH_WATER) {    // the socket will not push back for us
    return ws.close(1013, "slow consumer: reconnect and resume");
  }
  ws.send(JSON.stringify(obj));
}

setInterval(() => {                         // liveness: ping, drop silent peers
  for (const ws of wss.clients) {
    if (!ws.isAlive) { ws.terminate(); continue; }
    ws.isAlive = false;
    ws.ping();
  }
}, 30_000);

Failure modes and resume semantics

FailureWebSocket symptomgRPC symptomFix in both
Network dropclose 1006 (abnormal, no close frame) or nothing until a ping failsUNAVAILABLE on the next Send or RecvReconnect with backoff and jitter, resume from last acked sequence
Server deployclose 1001GOAWAY, then UNAVAILABLE for new streamsDrain, then reconnect elsewhere
Slow readerbufferedAmount grows, memory climbsSend blocksBounded buffers plus a coalesce or disconnect policy
Oversized messageclose 1009 (message too big)RESOURCE_EXHAUSTEDChunk large payloads, set explicit limits
Handler bugclose 1011 (internal error)INTERNAL or UNKNOWNIdempotent ops, sequence numbers, alert on rate

The column that matters is the last one. Both protocols lose in-flight messages on disconnect and neither tells you which ones arrived. Correctness comes from three things you design: sequence numbers in both directions, acknowledgements that say what was durably applied, and idempotent operations so a replayed message does no harm. Many gRPC implementations default to a maximum received message size of around 4 MB; check yours and set it explicitly, and set maxPayload on WebSocket servers for the same reason.

The hybrid most systems converge on

Browser and mobile clients connect to an edge gateway over WebSocket. The gateway authenticates the upgrade, terminates TLS, parses messages against a schema (often protobuf in binary frames), and opens a gRPC bidi stream to the document shard that owns the session. Inside the data centre you keep typed contracts, deadlines, per-stream flow control and HTTP/2 connection reuse; at the edge you keep universal client support.

The gateway is where the two backpressure models meet, so it needs a bounded queue per client that coalesces or disconnects, never one that grows without limit. It also maps errors: a gRPC UNAVAILABLE becomes a WebSocket close your client treats as reconnect-and-resume.

Decision guide

SituationChooseWhy
Browser clients, two-way trafficWebSocketThe universally supported full-duplex browser option
Service to service inside your networkgRPC bidiContracts, deadlines, status, flow control, connection reuse
Mostly server-to-client updatesNeither: SSESimpler, plain HTTP, automatic reconnect
Browser plus internal servicesHybridWebSocket at the edge, gRPC behind the gateway

The trade-off in one sentence: WebSocket minimises what the network and the client must support and maximises what you must design; gRPC does the reverse. Pick by where your clients are, then design the resume protocol either way, because that is the part neither one gives you. If you settle on WebSocket, the WebSocket architecture article covers the server side end to end.

What to do next

  1. List every client type (browser, mobile, service) and every network hop between it and your servers; if any client is a browser, plan for WebSocket at the edge.
  2. Write the message contract first, as a proto file even if you send it over WebSocket, with sequence numbers and a hello or resume message.
  3. Decide the slow-consumer policy (drop, coalesce or disconnect) and implement the bound: a high-water mark on bufferedAmount, or a bounded queue in front of a blocking gRPC Send.
  4. Set keepalive intervals below the shortest idle timeout on the path, and agree client and server keepalive settings for gRPC.
  5. Implement reconnect with exponential backoff and jitter, and test resume by killing connections mid-stream in a load test.
  6. Set explicit maximum message sizes on both protocols and alert on close codes 1006, 1011 and gRPC INTERNAL rates.
Key takeaway: WebSocket is a minimal framing layer that every browser speaks; gRPC bidi is a typed call on an HTTP/2 stream with contracts, status codes, deadlines and per-stream flow control. Browsers cannot open gRPC bidi streams, so browser-facing systems use WebSocket at the edge and often gRPC behind a gateway. Neither protocol replays lost messages: sequence numbers, acknowledgements, idempotent operations and an explicit slow-consumer policy are your job in both.