HTTP security habits are built around requests: authenticate, authorise, respond, forget. A WebSocket or a bidirectional gRPC stream breaks that model. One authentication event at the start is followed by hours of messages in both directions, on a connection that outlives the token that opened it, the permissions the user had when it opened, and sometimes the employment of the user. Most real incidents with bidirectional connections come from carrying request-era assumptions into that world.

This article builds a threat model for long-lived connections and maps a control to each threat: how to authenticate the handshake when the browser will not let you set headers, how to stop cross-site WebSocket hijacking, how to authorise every message and subscription, how to make expiry and revocation reach open connections, and how to limit message size, rate and connection count. The surrounding architecture, upgrade path and draining are covered in WebSocket architecture, in depth; this page is the security layer on top of it.

Advertisement

Threat model

ThreatWhat happensPrimary control
Cross-site WebSocket hijackinga malicious page opens a socket to your server and the browser attaches the victim's cookiesOrigin allowlist, no ambient cookie auth
Credential leakagea token in the URL lands in proxy and access logssingle-use, seconds-long tickets
Stale authorisationpermissions removed after connect still applyper-message checks, revocation push
Token expiry ignoredconnection stays open long after the token expiredserver-side expiry timer, re-auth
Message floodingone client sends thousands of messages a secondper-connection and per-user rate limits
Oversized or compressed bombshuge frames or tiny compressed frames that inflate to gigabytesmax payload enforced after decompression
Connection exhaustionmany idle sockets hold memory and file descriptorsper-IP and per-user caps, idle timeouts
Subscription snoopingclient subscribes to a topic it should not seeauthorise every subscribe, not only connect
Injection through messagesmessage fields reach queries, HTML or promptsschema validation, treat as untrusted

Two properties make this list different from an HTTP one. Time: anything decided at connect time decays. And amplification: one connection can carry unlimited messages, so a per-request limit at the edge does not apply; the edge saw one request, the upgrade.

Security controls on a long-lived bidirectional connection: at the edge, at the upgrade, and on every messageClientbrowser or serviceEdge / LBTLS, conn limits per IPwss://Upgrade handlerOrigin, ticket, quotaupgradeAuth serviceredeem ticketConnection loopsize, rate, schema101 SwitchingPolicy checkper message + topicPub/sub fan-outauthorised topicsRevocation bususer disabled, expiryClose 1008 / 4xxxwith reasonon violationAuthentication happens once at the upgrade; authorisation, limits and revocation keep running for the whole life of the connection
Authentication happens at the upgrade. The connection loop enforces size, rate and schema on each message and asks a policy check about each action and topic. A revocation bus can close any connection at any time.

Authenticating the handshake

The browser WebSocket constructor takes a URL and an optional list of subprotocols, nothing else. You cannot set an Authorization header, so the options are constrained:

OptionHowProblems
Session cookiebrowser sends cookies on the upgrade requestambient authority, so it is exactly what cross-site hijacking abuses; must be paired with an Origin check
Token in query stringwss://host/ws?token=...URLs are logged by proxies, load balancers and servers; long-lived tokens leak
Short-lived ticket in query stringclient fetches a single-use ticket over authenticated HTTPS, then connects with itneeds a ticket store; a leaked ticket is useless seconds later
Token in first messageconnect unauthenticated, send credentials firstserver holds unauthenticated sockets; needs a strict deadline and size limit
Token in Sec-WebSocket-Protocolabuse the subprotocol list to carry a tokenheader is logged and echoed; misuses the field; fragile with proxies

For browser clients, the ticket pattern is the strongest simple choice. The page calls POST /ws-ticket with its normal authentication, receives a random ticket bound to the user, the intended origin and an expiry of around 30 seconds, and connects with it. The server redeems the ticket atomically, so a second use fails. For service-to-service clients that can set headers, use a normal bearer token or mutual TLS at the upgrade.

Advertisement

Origin checks and cross-site hijacking

The WebSocket handshake is an HTTP GET, but browsers do not apply the same-origin policy to it the way they do to fetch responses: a page on any site can open a socket to your server and read what comes back. If your upgrade authenticates by cookie, that page now holds a fully authenticated connection on behalf of the visitor. This is cross-site WebSocket hijacking.

Browsers always send an Origin header on the handshake, and scripts cannot forge it. Check it against an exact allowlist at the upgrade and reject everything else with 403 before switching protocols. Compare full origins (scheme, host and port), not substrings; example.com.attacker.net contains example.com. Non-browser clients can send any Origin they like, so the check defends browsers against other sites, not your server against arbitrary clients. SameSite cookie settings help but differ in detail between browsers; do not rely on them alone.

Code: a defended upgrade in Node

With the ws library, take over the HTTP server's upgrade event so every check runs before the protocol switch:

import http from "node:http";
import { WebSocketServer } from "ws";

const ALLOWED_ORIGINS = new Set(["https://app.example.com"]);
const wss = new WebSocketServer({ noServer: true, maxPayload: 64 * 1024 });
const server = http.createServer();

function reject(socket, status, text) {
  socket.write(`HTTP/1.1 ${status} ${text}\r\nConnection: close\r\n\r\n`);
  socket.destroy();
}

server.on("upgrade", async (req, socket, head) => {
  if (!ALLOWED_ORIGINS.has(req.headers.origin)) return reject(socket, 403, "Forbidden");
  const ticket = new URL(req.url, "https://placeholder").searchParams.get("ticket");
  const user = ticket && await tickets.redeemOnce(ticket, req.headers.origin); // atomic, single use
  if (!user) return reject(socket, 401, "Unauthorized");
  if (!(await connQuota.tryAcquire(user.id))) return reject(socket, 429, "Too Many Requests");
  wss.handleUpgrade(req, socket, head, (ws) => wss.emit("connection", ws, req, user));
});

wss.on("connection", (ws, req, user) => {
  const session = { user, expiresAt: user.tokenExpiry, bucket: new TokenBucket(20, 40) };
  const timer = setTimeout(() => ws.close(4001, "session expired"), session.expiresAt - Date.now());
  ws.on("message", (data, isBinary) => handleMessage(ws, session, data, isBinary));
  ws.on("close", () => { clearTimeout(timer); connQuota.release(user.id); });
});

maxPayload is important: the library default is far larger than most applications need, and the limit should be set from your real message sizes. In ws it is also checked against the decompressed size when permessage-deflate is negotiated, which is what stops a small compressed frame inflating into an enormous one; if you use another library, confirm it does the same. Compression trade-offs are covered in WebSocket compression architecture.

Authorise every message, not just the connection

Authentication tells you who is on the other end. It does not tell you whether this particular message is allowed. Every inbound message should be validated against a schema, then checked by the same policy function your HTTP API uses, with the user, the action and the resource: can(user, 'post', channel). Subscriptions are the most commonly missed case: a client that sends {"type":"subscribe","topic":"tenant-42/orders"} must be authorised for tenant 42 at that moment, not merely be logged in.

async function handleMessage(ws, session, data, isBinary) {
  if (isBinary || !session.bucket.take(1)) return ws.close(1008, "rate limit");
  let msg;
  try { msg = Schema.parse(JSON.parse(data)); } catch { return ws.close(1008, "invalid message"); }
  if (!(await policy.can(session.user, msg.type, msg.resource))) {
    return ws.send(JSON.stringify({ id: msg.id, error: "forbidden" }));  // deny, keep connection
  }
  if (msg.type === "subscribe") {
    if (session.topics.size >= 50) return ws.send(JSON.stringify({ id: msg.id, error: "too many topics" }));
    session.topics.add(msg.resource);
    pubsub.subscribe(msg.resource, ws);
  }
  // ... dispatch other message types
}

Decide which violations close the connection and which only reject the message. Malformed or flooding traffic suggests a broken or hostile client and earns a close with code 1008 (policy violation); 1009 means a message was too big. A forbidden action from a well-behaved client, perhaps after a permission change, is better answered with an error message so the UI can react. Codes 4000 to 4999 are reserved for private use, so applications can define their own, such as 4001 for "session expired, re-authenticate", and clients can tell expiry apart from a server fault.

On the outbound side, filter fan-out too. If a topic's audience can change, check authorisation when delivering, or rebuild subscriptions when permissions change, so a user removed from a channel stops receiving it.

Expiry and revocation on open connections

A connection opened with a token that expires in 15 minutes must not live for 8 hours on that authority. Three mechanisms work together. First, a server-side timer at the token's expiry, as in the code above, closes the connection with a private code unless it has been renewed. Second, an in-band re-authentication message lets the client send a fresh token before expiry and extend the session without reconnecting. Third, a revocation bus: when an account is disabled, a password changed or a role removed, the identity system publishes an event, every connection server subscribes to it, and each closes the matching connections or drops their affected subscriptions.

Worked example. At 14:00:00 an administrator disables a contractor's account. Without revocation, their dashboard socket opened at 09:12 keeps streaming customer orders until it disconnects on its own, potentially for hours. With the bus: at 14:00:00.2 the identity service publishes user.disabled; at 14:00:00.3 each of the 40 connection servers receives it and looks the user up in its local index of connections by user id; at 14:00:00.4 the two matching sockets are closed with code 4003 and a reason; the client's reconnect attempt fails ticket issuance because the account is disabled. The cost was one index per server and one subscription. Test this path regularly; it only matters on the day nobody is watching.

Rate, size and connection limits

Limits must exist per connection, per user and per source address, because an attacker can open many connections and a user can have many devices. A token bucket per connection, with a sustained rate and a burst, is simple and fair:

class TokenBucket {
  constructor(ratePerSec, burst) { this.rate = ratePerSec; this.cap = burst; this.tokens = burst; this.at = Date.now(); }
  take(n) {
    const now = Date.now();
    this.tokens = Math.min(this.cap, this.tokens + (now - this.at) / 1000 * this.rate);
    this.at = now;
    if (this.tokens < n) return false;
    this.tokens -= n;
    return true;
  }
}
  • Messages and bytes. Limit both; a client within the message rate can still send large messages.
  • Connections per user and per IP. Enforced at the upgrade with a shared counter; return 429 before switching protocols. Be generous per IP where users share carrier-grade NAT.
  • Subscriptions per connection. Each subscription costs fan-out work; cap it.
  • Outbound buffer. A client that stops reading makes the server buffer. Check ws.bufferedAmount before sending and drop or close slow consumers past a threshold.
  • Unauthenticated time. If you authenticate in the first message, close any socket that has not done so within a few seconds.
  • Idle timeout. Heartbeats both detect dead peers and reclaim abandoned connections; intervals are discussed in heartbeat and keep-alive strategies.

When the whole service is overloaded rather than one client misbehaving, close with 1013 (try again later) and have clients reconnect with jittered backoff so they do not return all at once.

gRPC bidirectional streams

The same model applies to gRPC streams with different mechanics. Credentials arrive in metadata when the stream opens, and an interceptor authenticates them, so header-based tokens and mutual TLS work naturally. But the stream is still long-lived, so per-message authorisation, expiry timers and revocation are just as necessary. Use the server's maximum receive message size (4 MiB by default in most gRPC implementations) and lower it where messages are small, limit concurrent streams per connection through the HTTP/2 setting, and set deadlines so streams cannot be held open indefinitely. Return PERMISSION_DENIED for a forbidden message and UNAUTHENTICATED when the credential has expired, so clients know whether to refresh or give up.

Failure modes

  • Cookie auth without an Origin check. Exploitable cross-site hijacking. Add the allowlist at the upgrade.
  • Long-lived tokens in URLs. Found later in log archives. Switch to single-use tickets and scrub existing logs.
  • Authorisation only at connect. Removed users keep receiving data. Check every message and subscription; wire revocation.
  • Default payload limits. Memory exhaustion from a handful of large or compressed frames.
  • Limits only at the edge. The load balancer counts one request per connection and cannot see the thousands of messages behind it.
  • Close without a reason code. Clients cannot distinguish expiry from failure and reconnect in a tight loop. Use specific codes and backoff.
Key takeaway: <p>On a long-lived connection, authentication is a moment and authorisation is a process. Authenticate the upgrade with an Origin allowlist and a single-use ticket, validate and authorise every message and subscription, make expiry and revocation reach open sockets, and limit size, rate, connections and buffers per connection and per user.</p><p><strong>What to do next:</strong></p><ol><li>Check that the upgrade handler rejects origins outside an exact allowlist, and test it from a page on another domain.</li><li>Replace tokens in WebSocket URLs with single-use tickets valid for about 30 seconds.</li><li>Route every inbound message and subscribe request through the same policy function as your HTTP API.</li><li>Add an expiry timer per connection and an in-band re-authentication message.</li><li>Wire a revocation event from the identity system to every connection server and rehearse the disabled-user drill.</li><li>Set maxPayload, per-connection rate limits, per-user connection caps and an outbound buffer limit from measured traffic.</li><li>Define private close codes for expiry and revocation and make clients back off with jitter on reconnect.</li></ol>