HTTP security habits are built around requests: authenticate, authorise, respond, forget. A WebSocket or a bidirectional gRPC stream breaks that model. One authentication event at the start is followed by hours of messages in both directions, on a connection that outlives the token that opened it, the permissions the user had when it opened, and sometimes the employment of the user. Most real incidents with bidirectional connections come from carrying request-era assumptions into that world.
This article builds a threat model for long-lived connections and maps a control to each threat: how to authenticate the handshake when the browser will not let you set headers, how to stop cross-site WebSocket hijacking, how to authorise every message and subscription, how to make expiry and revocation reach open connections, and how to limit message size, rate and connection count. The surrounding architecture, upgrade path and draining are covered in WebSocket architecture, in depth; this page is the security layer on top of it.
Threat model
| Threat | What happens | Primary control |
|---|---|---|
| Cross-site WebSocket hijacking | a malicious page opens a socket to your server and the browser attaches the victim's cookies | Origin allowlist, no ambient cookie auth |
| Credential leakage | a token in the URL lands in proxy and access logs | single-use, seconds-long tickets |
| Stale authorisation | permissions removed after connect still apply | per-message checks, revocation push |
| Token expiry ignored | connection stays open long after the token expired | server-side expiry timer, re-auth |
| Message flooding | one client sends thousands of messages a second | per-connection and per-user rate limits |
| Oversized or compressed bombs | huge frames or tiny compressed frames that inflate to gigabytes | max payload enforced after decompression |
| Connection exhaustion | many idle sockets hold memory and file descriptors | per-IP and per-user caps, idle timeouts |
| Subscription snooping | client subscribes to a topic it should not see | authorise every subscribe, not only connect |
| Injection through messages | message fields reach queries, HTML or prompts | schema validation, treat as untrusted |
Two properties make this list different from an HTTP one. Time: anything decided at connect time decays. And amplification: one connection can carry unlimited messages, so a per-request limit at the edge does not apply; the edge saw one request, the upgrade.
Authenticating the handshake
The browser WebSocket constructor takes a URL and an optional list of subprotocols, nothing else. You cannot set an Authorization header, so the options are constrained:
| Option | How | Problems |
|---|---|---|
| Session cookie | browser sends cookies on the upgrade request | ambient authority, so it is exactly what cross-site hijacking abuses; must be paired with an Origin check |
| Token in query string | wss://host/ws?token=... | URLs are logged by proxies, load balancers and servers; long-lived tokens leak |
| Short-lived ticket in query string | client fetches a single-use ticket over authenticated HTTPS, then connects with it | needs a ticket store; a leaked ticket is useless seconds later |
| Token in first message | connect unauthenticated, send credentials first | server holds unauthenticated sockets; needs a strict deadline and size limit |
| Token in Sec-WebSocket-Protocol | abuse the subprotocol list to carry a token | header is logged and echoed; misuses the field; fragile with proxies |
For browser clients, the ticket pattern is the strongest simple choice. The page calls POST /ws-ticket with its normal authentication, receives a random ticket bound to the user, the intended origin and an expiry of around 30 seconds, and connects with it. The server redeems the ticket atomically, so a second use fails. For service-to-service clients that can set headers, use a normal bearer token or mutual TLS at the upgrade.
Origin checks and cross-site hijacking
The WebSocket handshake is an HTTP GET, but browsers do not apply the same-origin policy to it the way they do to fetch responses: a page on any site can open a socket to your server and read what comes back. If your upgrade authenticates by cookie, that page now holds a fully authenticated connection on behalf of the visitor. This is cross-site WebSocket hijacking.
Browsers always send an Origin header on the handshake, and scripts cannot forge it. Check it against an exact allowlist at the upgrade and reject everything else with 403 before switching protocols. Compare full origins (scheme, host and port), not substrings; example.com.attacker.net contains example.com. Non-browser clients can send any Origin they like, so the check defends browsers against other sites, not your server against arbitrary clients. SameSite cookie settings help but differ in detail between browsers; do not rely on them alone.
Code: a defended upgrade in Node
With the ws library, take over the HTTP server's upgrade event so every check runs before the protocol switch:
import http from "node:http";
import { WebSocketServer } from "ws";
const ALLOWED_ORIGINS = new Set(["https://app.example.com"]);
const wss = new WebSocketServer({ noServer: true, maxPayload: 64 * 1024 });
const server = http.createServer();
function reject(socket, status, text) {
socket.write(`HTTP/1.1 ${status} ${text}\r\nConnection: close\r\n\r\n`);
socket.destroy();
}
server.on("upgrade", async (req, socket, head) => {
if (!ALLOWED_ORIGINS.has(req.headers.origin)) return reject(socket, 403, "Forbidden");
const ticket = new URL(req.url, "https://placeholder").searchParams.get("ticket");
const user = ticket && await tickets.redeemOnce(ticket, req.headers.origin); // atomic, single use
if (!user) return reject(socket, 401, "Unauthorized");
if (!(await connQuota.tryAcquire(user.id))) return reject(socket, 429, "Too Many Requests");
wss.handleUpgrade(req, socket, head, (ws) => wss.emit("connection", ws, req, user));
});
wss.on("connection", (ws, req, user) => {
const session = { user, expiresAt: user.tokenExpiry, bucket: new TokenBucket(20, 40) };
const timer = setTimeout(() => ws.close(4001, "session expired"), session.expiresAt - Date.now());
ws.on("message", (data, isBinary) => handleMessage(ws, session, data, isBinary));
ws.on("close", () => { clearTimeout(timer); connQuota.release(user.id); });
});maxPayload is important: the library default is far larger than most applications need, and the limit should be set from your real message sizes. In ws it is also checked against the decompressed size when permessage-deflate is negotiated, which is what stops a small compressed frame inflating into an enormous one; if you use another library, confirm it does the same. Compression trade-offs are covered in WebSocket compression architecture.
Authorise every message, not just the connection
Authentication tells you who is on the other end. It does not tell you whether this particular message is allowed. Every inbound message should be validated against a schema, then checked by the same policy function your HTTP API uses, with the user, the action and the resource: can(user, 'post', channel). Subscriptions are the most commonly missed case: a client that sends {"type":"subscribe","topic":"tenant-42/orders"} must be authorised for tenant 42 at that moment, not merely be logged in.
async function handleMessage(ws, session, data, isBinary) {
if (isBinary || !session.bucket.take(1)) return ws.close(1008, "rate limit");
let msg;
try { msg = Schema.parse(JSON.parse(data)); } catch { return ws.close(1008, "invalid message"); }
if (!(await policy.can(session.user, msg.type, msg.resource))) {
return ws.send(JSON.stringify({ id: msg.id, error: "forbidden" })); // deny, keep connection
}
if (msg.type === "subscribe") {
if (session.topics.size >= 50) return ws.send(JSON.stringify({ id: msg.id, error: "too many topics" }));
session.topics.add(msg.resource);
pubsub.subscribe(msg.resource, ws);
}
// ... dispatch other message types
}Decide which violations close the connection and which only reject the message. Malformed or flooding traffic suggests a broken or hostile client and earns a close with code 1008 (policy violation); 1009 means a message was too big. A forbidden action from a well-behaved client, perhaps after a permission change, is better answered with an error message so the UI can react. Codes 4000 to 4999 are reserved for private use, so applications can define their own, such as 4001 for "session expired, re-authenticate", and clients can tell expiry apart from a server fault.
On the outbound side, filter fan-out too. If a topic's audience can change, check authorisation when delivering, or rebuild subscriptions when permissions change, so a user removed from a channel stops receiving it.
Expiry and revocation on open connections
A connection opened with a token that expires in 15 minutes must not live for 8 hours on that authority. Three mechanisms work together. First, a server-side timer at the token's expiry, as in the code above, closes the connection with a private code unless it has been renewed. Second, an in-band re-authentication message lets the client send a fresh token before expiry and extend the session without reconnecting. Third, a revocation bus: when an account is disabled, a password changed or a role removed, the identity system publishes an event, every connection server subscribes to it, and each closes the matching connections or drops their affected subscriptions.
Worked example. At 14:00:00 an administrator disables a contractor's account. Without revocation, their dashboard socket opened at 09:12 keeps streaming customer orders until it disconnects on its own, potentially for hours. With the bus: at 14:00:00.2 the identity service publishes user.disabled; at 14:00:00.3 each of the 40 connection servers receives it and looks the user up in its local index of connections by user id; at 14:00:00.4 the two matching sockets are closed with code 4003 and a reason; the client's reconnect attempt fails ticket issuance because the account is disabled. The cost was one index per server and one subscription. Test this path regularly; it only matters on the day nobody is watching.
Rate, size and connection limits
Limits must exist per connection, per user and per source address, because an attacker can open many connections and a user can have many devices. A token bucket per connection, with a sustained rate and a burst, is simple and fair:
class TokenBucket {
constructor(ratePerSec, burst) { this.rate = ratePerSec; this.cap = burst; this.tokens = burst; this.at = Date.now(); }
take(n) {
const now = Date.now();
this.tokens = Math.min(this.cap, this.tokens + (now - this.at) / 1000 * this.rate);
this.at = now;
if (this.tokens < n) return false;
this.tokens -= n;
return true;
}
}- Messages and bytes. Limit both; a client within the message rate can still send large messages.
- Connections per user and per IP. Enforced at the upgrade with a shared counter; return 429 before switching protocols. Be generous per IP where users share carrier-grade NAT.
- Subscriptions per connection. Each subscription costs fan-out work; cap it.
- Outbound buffer. A client that stops reading makes the server buffer. Check
ws.bufferedAmountbefore sending and drop or close slow consumers past a threshold. - Unauthenticated time. If you authenticate in the first message, close any socket that has not done so within a few seconds.
- Idle timeout. Heartbeats both detect dead peers and reclaim abandoned connections; intervals are discussed in heartbeat and keep-alive strategies.
When the whole service is overloaded rather than one client misbehaving, close with 1013 (try again later) and have clients reconnect with jittered backoff so they do not return all at once.
gRPC bidirectional streams
The same model applies to gRPC streams with different mechanics. Credentials arrive in metadata when the stream opens, and an interceptor authenticates them, so header-based tokens and mutual TLS work naturally. But the stream is still long-lived, so per-message authorisation, expiry timers and revocation are just as necessary. Use the server's maximum receive message size (4 MiB by default in most gRPC implementations) and lower it where messages are small, limit concurrent streams per connection through the HTTP/2 setting, and set deadlines so streams cannot be held open indefinitely. Return PERMISSION_DENIED for a forbidden message and UNAUTHENTICATED when the credential has expired, so clients know whether to refresh or give up.
Failure modes
- Cookie auth without an Origin check. Exploitable cross-site hijacking. Add the allowlist at the upgrade.
- Long-lived tokens in URLs. Found later in log archives. Switch to single-use tickets and scrub existing logs.
- Authorisation only at connect. Removed users keep receiving data. Check every message and subscription; wire revocation.
- Default payload limits. Memory exhaustion from a handful of large or compressed frames.
- Limits only at the edge. The load balancer counts one request per connection and cannot see the thousands of messages behind it.
- Close without a reason code. Clients cannot distinguish expiry from failure and reconnect in a tight loop. Use specific codes and backoff.