Authenticating a WebSocket is awkward for one structural reason: the browser's WebSocket constructor accepts a URL and a list of subprotocols and nothing else, so the familiar Authorization header is not available to browser code. Every pattern in use is a way of working around that constraint, and each moves the credential to a different layer with different logging, replay and expiry behaviour.
The bidi security article covers the threat model: cross-site hijacking, origin checks, per-message authorisation, revocation and limits. This article is the implementation companion. It answers which layer should authenticate for each kind of client, then builds the main patterns properly: a one-time ticket service, first-frame authentication, and an in-band refresh protocol so long-lived connections do not outlive their credentials. It closes with how four common stacks implement the same ideas and a worked migration away from long-lived tokens in URLs.
Where authentication can live
A WebSocket connection passes through up to four points where a credential can be checked. The edge or gateway sees the TLS session, client certificates, IP address and request headers. The upgrade handler in your application sees the HTTP GET that asks to switch protocols, including cookies, the query string and the Origin header. After the switch, the first application frame can carry credentials in your own protocol. Finally, every later message must be authorised against the identity established earlier.
The design rule is to authenticate at the earliest layer the client can support, because each later layer means holding an unauthenticated socket open for longer. Rejecting at the upgrade costs one HTTP response. Rejecting after the first frame means you have already allocated a connection, a buffer and probably a file descriptor for an anonymous peer, which is why first-frame authentication needs a strict deadline.
Choosing a pattern by client type
| Client | Recommended pattern | Why |
|---|---|---|
| Browser, same site as the API | Session cookie at upgrade + exact Origin allowlist | Cookie is sent automatically; Origin check stops cross-site hijacking |
| Browser SPA with bearer tokens | One-time ticket in the query string | Nothing long-lived appears in URLs or logs |
| Browser via a framework (Socket.IO) | Token in the framework's connect payload | Credential travels inside the protocol, not the URL |
| Mobile or desktop app | Authorization header at upgrade | Native clients can set headers; reuse API validation |
| Service to service | mTLS at the edge, or a bearer header | Workload identity without user tokens |
| Devices behind a gateway | Gateway authorizer on connect | Reject before your servers spend anything |
Two patterns are missing on purpose. Long-lived access tokens in the query string end up in load balancer logs, proxy logs, browser history and error trackers, so the credential outlives its purpose in a dozen places. Smuggling a token through Sec-WebSocket-Protocol works but misuses a negotiation header that servers echo and proxies log; WebSocket subprotocols explains the trade-offs if you inherit it.
Building a one-time ticket service
A ticket is a short random string that stands in for the user's real credential for a few seconds. The page calls an ordinary HTTPS endpoint, authenticated the way the rest of the API is, and receives a ticket. It then opens wss://rt.example.com/ws?ticket=.... The upgrade handler redeems the ticket and attaches the stored identity to the connection. Because the ticket is single-use and expires in about thirty seconds, a copy in a log is useless by the time anyone reads it.
// Ticket issue and redemption (Node, node-redis v4). GETDEL needs Redis 6.2+.
import crypto from "node:crypto";
const TICKET_TTL_S = 30;
// POST /ws-ticket -- behind your normal session or bearer authentication
export async function issueTicket(req, res) {
const ticket = crypto.randomBytes(32).toString("base64url");
const claims = { sub: req.user.id, tenant: req.user.tenant, scopes: req.user.scopes,
origin: req.get("Origin"), exp: req.user.tokenExp };
await redis.set(`wst:${ticket}`, JSON.stringify(claims), { EX: TICKET_TTL_S });
res.set("Cache-Control", "no-store").json({ ticket });
}
// Called from the HTTP server's "upgrade" event, before handleUpgrade
export async function redeemTicket(req) {
const ticket = new URL(req.url, "http://x").searchParams.get("ticket");
if (!ticket || ticket.length > 64) return null;
const raw = await redis.getDel(`wst:${ticket}`); // atomic: a second use finds nothing
if (!raw) return null;
const claims = JSON.parse(raw);
if (claims.origin !== req.headers.origin) return null; // bound to the page that asked for it
return claims;
}Four details make the difference between a ticket and a weaker token. Redemption is atomic: GETDEL reads and deletes in one command, so two concurrent upgrades cannot both succeed. The ticket is bound to the origin of the page that requested it, so a ticket stolen by a script on another site fails the comparison. The claims carry the expiry of the underlying access token, so the connection inherits a deadline instead of living forever. And the issuing response is no-store so no cache keeps it.
The stateless alternative is a signed ticket: an HMAC or a short JWT with a few seconds of lifetime and a unique id. It removes the Redis round trip but cannot be single-use without storing the ids it has seen, at which point you have rebuilt the stateful version. Use signed tickets when the issuing and redeeming services cannot share a store, keep their lifetime under ten seconds, and accept that a replay inside that window is possible.
First-frame authentication
Some environments cannot do anything at the upgrade: a gateway that does not forward query strings, or a protocol such as STOMP or a GraphQL-over-WebSocket protocol that defines its own connect message. In those cases the connection opens anonymously and the first frame carries the token. The connection is a state machine with two states, unauthenticated and authenticated, and the unauthenticated state must be tiny.
- Arm a deadline at accept time, typically five seconds, and close with code 1008 (policy violation) if no valid auth frame arrives.
- Accept exactly one frame type while unauthenticated; anything else closes the connection.
- Apply a small maximum frame size before authentication, a few kilobytes, separate from the larger limit after.
- Do not subscribe the socket to any topic, room or broadcast until the auth frame has been validated.
- Count unauthenticated sockets per IP and globally, and shed load when the count spikes.
The cost is that the origin check and the credential check now happen at different times. Keep the Origin allowlist at the upgrade regardless, so browsers on other sites are rejected before they can even attempt the first frame.
Keeping long-lived connections honest: in-band refresh
Access tokens usually live for minutes; dashboard and chat connections live for hours. Without a refresh path you either let the connection outlive its credential or force a reconnect every few minutes, which drops in-flight state and multiplies handshake load. An in-band refresh lets the client present a new token over the existing socket. The protocol has three messages and one timer.
// In-band refresh on the server: the connection keeps its identity, only exp moves.
function armExpiry(ws) {
clearTimeout(ws.expiryTimer);
const ms = ws.identity.exp * 1000 - Date.now();
ws.expiryTimer = setTimeout(() => ws.close(4001, "credential expired"), Math.max(ms, 0));
}
ws.on("message", async (data) => {
const msg = JSON.parse(data);
if (msg.type === "reauth") {
const claims = await verifyAccessToken(msg.token); // same validation as your HTTP API
if (!claims || claims.sub !== ws.identity.sub) return ws.close(4003, "identity mismatch");
ws.identity.exp = claims.exp;
ws.identity.scopes = claims.scopes; // scopes may shrink; honour it
armExpiry(ws);
return ws.send(JSON.stringify({ type: "reauth_ok", exp: claims.exp }));
}
if (!authorise(ws.identity, msg)) return ws.send(JSON.stringify({ type: "error", code: "forbidden" }));
handle(ws, msg);
});The client schedules a refresh at about 80 percent of the token lifetime, obtains a new access token through its normal refresh flow, and sends {type: "reauth", token}. The server validates it exactly as the HTTP API would, JWT validation included, insists that the subject has not changed, replaces the scopes and moves the timer. If the refresh never arrives the timer closes the socket with a private code in the 4000-4999 range, and the client knows to obtain a fresh ticket rather than retry blindly. Note the scope replacement: if an administrator removed a permission, the next refresh is where the connection loses it.
How common stacks implement it
Frameworks wrap these ideas in their own vocabulary. The mapping matters because each one places the credential at a different layer, which changes what appears in logs and when an invalid token is detected.
// --- Socket.IO: the auth object travels in the CONNECT packet, not the URL ---
const socket = io("https://rt.example.com", { auth: (cb) => cb({ token: getAccessToken() }) });
io.use(async (socket, next) => {
const claims = await verifyAccessToken(socket.handshake.auth.token);
if (!claims) return next(new Error("unauthorized")); // client receives connect_error
socket.data.identity = claims;
next();
});
// --- ASP.NET Core SignalR: browsers send access_token in the query string ---
// client: new HubConnectionBuilder().withUrl("/hubs/chat", { accessTokenFactory: getAccessToken })
options.Events = new JwtBearerEvents {
OnMessageReceived = ctx => {
var token = ctx.Request.Query["access_token"];
if (!string.IsNullOrEmpty(token) && ctx.HttpContext.Request.Path.StartsWithSegments("/hubs"))
ctx.Token = token;
return Task.CompletedTask;
}
};
// --- Spring STOMP over WebSocket: authenticate the CONNECT frame ---
registration.interceptors(new ChannelInterceptor() {
@Override public Message<?> preSend(Message<?> message, MessageChannel channel) {
StompHeaderAccessor acc = MessageHeaderAccessor.getAccessor(message, StompHeaderAccessor.class);
if (StompCommand.CONNECT.equals(acc.getCommand()))
acc.setUser(authenticate(acc.getFirstNativeHeader("Authorization")));
return message;
}
});- Socket.IO sends the
authobject inside its CONNECT packet once the transport is open, so the token stays out of URLs. Passing a function, as above, means every reconnect fetches a current token. - SignalR documents sending the token as
access_tokenin the query string for WebSockets and Server-Sent Events because browsers cannot set headers. Restrict the query lookup to hub paths, keep token lifetimes short, and make sure request logging does not record full URLs. - Spring STOMP authenticates the STOMP CONNECT frame with a
ChannelInterceptoron the client inbound channel and sets the user on the session, which later@MessageMappingmethods and destination security rules use. - AWS API Gateway WebSocket APIs run authorizers only on the
$connectroute, using IAM or a Lambda authorizer that can read headers and query string parameters. Persist the resulting identity against the connection id during connect so later routes can look it up.
Many servers, one identity
Once you run more than one connection server, identity becomes distributed state. Keep the authoritative record on the connection object in memory and index it by user id on each node, so a revocation event can be fanned out and each node closes its own matching sockets. Tickets must be redeemable on any node, which is why the store is shared; sticky routing at the load balancer is not an authentication mechanism. Load balancer pitfalls covers idle timeouts and header handling that silently break the upgrade path.
Reconnection interacts with authentication in a way that bites at scale. After a deploy or network blip, every client reconnects, and with tickets every reconnect is first an HTTPS call to the issuing endpoint. Rate limit issuance per user, add jitter to reconnects as described in reconnection strategies, and make sure the issuing service scales with the connection fleet rather than with ordinary API traffic.
Worked example: removing tokens from URLs
A team runs a trading dashboard that connects with wss://rt.example.com/ws?token=<15-minute JWT>. A log review finds those JWTs in the CDN access logs, retained for 90 days. The migration runs in four steps. First, add the ticket endpoint and the redemption path while still accepting token, and emit a metric for each method. Second, ship the client change: fetch a ticket, connect with it, and send reauth at 12 minutes. Third, once the metric shows under one percent of connections using token, reject it at the upgrade with HTTP 401; browser clients only observe a failed connection, commonly close code 1006, so the client falls back to fetching a ticket. Fourth, purge or redact the historical logs and add a log-pipeline rule that strips query strings on the WebSocket path. The result: no credential in any URL lasts longer than 30 seconds, and every connection has a deadline tied to a real token.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Clients loop reconnecting with no error detail | Browser hides the HTTP status of a failed upgrade | Expose a pre-flight auth check endpoint; log rejection reasons server-side |
| Ticket accepted twice | GET then DEL as two commands | Atomic GETDEL or a Lua script |
| Connections survive account disable | No expiry timer or revocation fan-out | Arm a timer at exp; subscribe nodes to revocation events |
| Spike of anonymous sockets | First-frame auth without a deadline | Five-second deadline, tiny pre-auth frame limit, per-IP caps |
| Tokens found in logs | Query-string tokens and default URL logging | Tickets, URL redaction, shorter token lifetimes |
| Permissions removed but still effective | Scopes captured at connect, never refreshed | Replace scopes on reauth; push revocation events |
What to do next
- List every WebSocket entry point and record which layer authenticates it and what credential it accepts.
- Remove long-lived tokens from URLs: introduce a ticket endpoint with atomic, origin-bound, 30-second tickets.
- Store the credential expiry on each connection, arm a close timer, and implement an in-band reauth message.
- For first-frame protocols, enforce a connect deadline, a pre-auth frame size limit and per-IP unauthenticated caps.
- Confirm request logs on the WebSocket path strip query strings, at the CDN, load balancer and application.
- Fan out revocation events to every connection node and test disabling a user with an open connection.
- Re-read bidi security to confirm per-message authorisation and Origin checks are in place.