WebSocket bugs are hard to see. Once the handshake succeeds, the connection stops being HTTP: there is no request log per message, many proxies log a single line when the connection finally ends, and the browser's onclose handler often reports only code 1006 with an empty reason. Teams then guess, add reconnect loops and ship. The fix is a method, not a single tool: decide which layer could explain the symptom, and use the tool that sees that layer.

This page walks through that method from the bottom up. You will test a handshake with curl, read frames in browser DevTools, talk to an endpoint by hand with wscat and websocat, decrypt TLS traffic in Wireshark, read close codes as evidence, automate captures with Playwright, and add the one server log line that answers most questions. A worked example traces a disconnect that happens every sixty seconds.

Advertisement

The method: isolate the layer

A WebSocket connection is a stack of five things that can each fail: reachability and TLS, the HTTP Upgrade handshake, frames, your application protocol, and the lifecycle over minutes or hours. "Never connects" lives in the bottom two layers. "Connects, then nothing arrives" is usually the application protocol, such as a rejected subscribe message. "Drops after a fixed time" is almost always an idle timer in a proxy or load balancer.

Reproduce outside your application early. If a command-line client against the same URL fails the same way, your frontend is cleared. If it works and the browser does not, look at what the browser adds: cookies, Origin, extensions, a service worker or a corporate proxy.

Debug a WebSocket one layer at a time, with the tool that sees that layer1. DNS, TCP, TLScan we reach and trust the host?2. HTTP Upgrade handshake101 Switching Protocols or not?3. Framestext, binary, ping, pong, close4. Application messagesyour JSON or protobuf protocol5. Lifecycleidle timeouts, reconnects, close codesopenssl s_client, curl -vcertificate, SNI, ALPN, reachabilitycurl with Upgrade headers, DevTools Headersstatus, Sec-WebSocket-Accept, cookiesWireshark + SSLKEYLOGFILE, tcpdumpopcodes, masking, who sent the closeDevTools Messages, wscat, websocatsend and read real payloads by handServer logs and metrics, Playwright captureduration, close code, reason, countsStart at the lowest layer that could explain the symptom. A handshake failure is never fixed in message code,and a disconnect every 60 seconds is almost always a proxy timer, not your application.
Five layers, and the tool that can observe each one.

Test the handshake with curl

A WebSocket starts as an ordinary HTTP/1.1 GET with Upgrade: websocket, Connection: Upgrade, Sec-WebSocket-Version: 13 and a random base64 Sec-WebSocket-Key. The server must answer 101 Switching Protocols with Sec-WebSocket-Accept, a hash of the key and a fixed GUID. You can perform exactly that request with curl and see the raw response, which takes the browser and your client library out of the picture.

# Does the endpoint upgrade at all? Force HTTP/1.1 so the Upgrade header is legal.
curl -i -N --http1.1 \
  -H "Connection: Upgrade" \
  -H "Upgrade: websocket" \
  -H "Sec-WebSocket-Version: 13" \
  -H "Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==" \
  -H "Origin: https://app.example.com" \
  https://api.example.com/ws

# Healthy answer:
# HTTP/1.1 101 Switching Protocols
# Upgrade: websocket
# Connection: Upgrade
# Sec-WebSocket-Accept: s3pPLMBiTxaQ9kYGzzhZRbK+xOo=

The key is the sample from RFC 6455, so the expected accept value is known. Use --http1.1 against https endpoints, because HTTP/2 forbids the Upgrade header (WebSockets over HTTP/2 use RFC 8441's extended CONNECT instead). After the 101, press Ctrl-C.

ResponseWhat it usually means
101 with Sec-WebSocket-AcceptThe handshake path is healthy. Move up a layer.
200 with an HTML or JSON bodySomething in the path ignored the Upgrade: a proxy not configured for WebSockets, or a route served by the wrong backend.
400Missing or malformed upgrade headers, often stripped by a proxy that speaks HTTP/1.0 or drops hop-by-hop headers.
401 or 403Authentication or Origin check. Repeat with and without the Origin and cookie headers the browser sends.
404Wrong path, or the WebSocket route exists only on some backends behind the load balancer.
426 Upgrade RequiredThe server wants a WebSocket upgrade or a different version on this path.
502 or 504The proxy could not reach the backend or timed out waiting for the 101.

In nginx, the classic cause of a 200 or 400 is a location block missing proxy_http_version 1.1 and the two lines that forward Upgrade and set Connection, because those are hop-by-hop headers that a proxy drops unless told otherwise. The frame format deep dive covers what happens on the wire after the 101.

Advertisement

Browser DevTools

In Chrome, open the Network panel, filter it to WebSocket connections and reload; each connection is one row. Selecting it shows a Headers tab with the upgrade request and the 101 response (check the cookies and Origin actually sent), a Messages tab listing every frame with direction, payload, length and time, and a Timing tab. The Messages view can filter by text and by direction and expands JSON payloads. Since Chrome 99, the Network panel's throttling presets also apply to WebSocket traffic, so you can reproduce slow-network behaviour such as a client that sends faster than a 3G link drains. Firefox's Network panel has an equivalent WebSocket message inspector.

Preserve the log across navigations so a reconnect does not wipe the evidence, and check the Initiator to see which script opened each socket. For problems you cannot reproduce, ask the user to record a network log at chrome://net-export, which saves connection-level events, including socket errors, to a file.

Talk to the endpoint by hand: wscat and websocat

Command-line clients let you speak the application protocol yourself. wscat (installed from npm) is interactive: it prints what arrives and sends each line you type. Its flags cover most debugging needs: -H adds headers such as authorization, -s requests a subprotocol, -o sets the Origin, -P prints a notice when a ping or pong arrives, --slash enables /ping, /pong and /close commands for control frames, -x sends a message after connecting and -w waits a number of seconds before exiting. websocat is a netcat-style tool that connects standard input and output to a socket, which makes it easy to pipe a fixture file in and grep the replies.

# Interactive session with auth, a subprotocol and visible pings
wscat -c wss://api.example.com/ws \
  -H "Authorization: Bearer $TOKEN" \
  -s chat.v2 -P --slash
> {"type":"subscribe","channel":"orders"}
< {"type":"subscribed","channel":"orders"}
> /ping hello          # send a ping control frame
> /close 1000 done     # close with a code and reason

# Scripted: send one message, keep the socket open 10 seconds, print replies
wscat -c wss://api.example.com/ws -x '{"type":"snapshot"}' -w 10

# websocat is a netcat-style client; pipe lines in, read lines out
echo '{"type":"snapshot"}' | websocat -v wss://api.example.com/ws

Use them for precise questions: does the server close with 1008 when the token is expired, answer a ping, or reject a 2 MB message with 1009? Each takes a minute by hand.

Close codes are evidence

Every intentional close carries a status code. The ones you will see most, with what they usually mean in practice:

CodeMeaningDebugging hint
1000Normal closureSomeone chose to close; check which side sent the close frame.
1001Going awayServer shutdown or deploy, or a browser tab navigating away.
1002 / 1007Protocol error / invalid payloadMalformed frames, or invalid UTF-8 in a text frame. Suspect a custom framing layer.
1006Abnormal closure (never sent on the wire)The TCP connection ended without a close frame: a proxy timeout, a crash, a network drop.
1008Policy violationCommonly used for authentication or authorization failures.
1009Message too bigA frame or message exceeded the peer's size limit.
1011Internal errorThe server hit an unexpected condition; look in server logs at that time.

1006 is the important one. It is reserved: no endpoint sends it. The browser reports it when the socket closed without a close handshake, and CloseEvent.wasClean will be false. So 1006 tells you to stop looking at application logic and start looking at whatever sat in between, and at the exact duration of the connection.

Go to the wire: Wireshark and TLS keys

When both ends insist they did nothing wrong, a packet capture settles it. For TLS, start Chrome, Firefox or curl with SSLKEYLOGFILE=/tmp/keys.log and point Wireshark's TLS pre-master-secret log preference at that file. The display filter websocket then shows decoded frames with opcodes, the FIN bit and unmasked payloads. On a server, capture with tcpdump -i any -w ws.pcap port 443 and analyse elsewhere.

A capture answers what nothing else can: which side sent the FIN or RST, whether a close frame was sent, whether pongs came back, and the exact gap between the last data and the disconnect. Compressed payloads look like noise until you account for the extension; the permessage-deflate article explains the negotiation.

Treat the key log file as a secret: anyone holding it can decrypt the captured session. Never enable it on production servers or share it with a capture.

Automate capture with Playwright

Intermittent problems need recordings, not someone staring at DevTools. Playwright exposes WebSocket traffic for pages it drives: the page emits a websocket event per connection, and each socket emits framesent, framereceived, socketerror and close. A small script turns that into a timestamped JSON log you can attach to a ticket or run in CI on a schedule.

import { chromium } from 'playwright';

const browser = await chromium.launch();
const page = await browser.newPage();
const t0 = Date.now();

page.on('websocket', ws => {
  const log = (kind, data) =>
    console.log(JSON.stringify({ t: Date.now() - t0, url: ws.url(), kind, data }));
  log('open');
  ws.on('framesent', f => log('sent', String(f.payload).slice(0, 200)));
  ws.on('framereceived', f => log('recv', String(f.payload).slice(0, 200)));
  ws.on('socketerror', e => log('error', e));
  ws.on('close', () => log('close'));
});

await page.goto('https://app.example.com/orders');
await page.waitForTimeout(180_000);   // long enough to see an idle timeout
await browser.close();

The same data is available lower down through the Chrome DevTools Protocol's Network domain events for WebSocket creation, handshake, frames and closure, if you need it from another automation stack. For load and reconnect behaviour at scale, load testing WebSockets with k6 is the next step.

Instrument the server

Most WebSocket incidents are diagnosed from server-side data, so build it first. The minimum is one structured log line per connection at close: connection id, user, subprotocol, duration, message counts, close code and reason, and which side closed. Add metrics for open connections, closes per second by code, and a duration histogram. A spike at one exact duration is an idle timer; a burst of 1006 across all users at once is a deploy or network event.

import asyncio, logging, time
import websockets

logging.basicConfig(level=logging.INFO)
# DEBUG on this logger prints every frame in and out; use it in staging only.
logging.getLogger("websockets").setLevel(logging.DEBUG)
log = logging.getLogger("ws.lifecycle")

async def handler(ws):
    opened, msgs_in = time.monotonic(), 0
    try:
        async for message in ws:
            msgs_in += 1
            await ws.send(message)
    finally:
        # One structured line per connection is the most useful debugging artifact you can own.
        log.info("closed conn=%s code=%s reason=%r duration_s=%.1f msgs_in=%d",
                 ws.id, ws.close_code, ws.close_reason,
                 time.monotonic() - opened, msgs_in)

async def main():
    async with websockets.serve(handler, "0.0.0.0", 8080,
                                ping_interval=20, ping_timeout=20):
        await asyncio.Future()

The Python websockets library logs each frame at DEBUG on its own logger, which is invaluable in staging and far too verbose for production. Its server also sends pings on an interval and closes connections whose pongs do not return in time, which is the keepalive pattern described in WebSocket ping, pong and keepalive.

Worked example: a disconnect every sixty seconds

Users of an order-tracking page report that live updates stop and the page reconnects about once a minute during quiet periods. In the browser, onclose reports 1006 with wasClean false. DevTools shows the connection's last message, a status update, and the close roughly 60 seconds later. Busy periods, when updates arrive every few seconds, are fine.

Following the layers: the handshake is healthy (101 every time), and wscat against the public URL reproduces the drop at 60 seconds, so the frontend is cleared. Running wscat against the backend's internal address from inside the network keeps the connection open for ten minutes, so the server is cleared too. That leaves the path: the load balancer and the nginx tier. A 60-second idle timeout is the default for both an AWS Application Load Balancer and nginx's proxy_read_timeout. A tcpdump on the nginx host shows nginx closing the upstream connection exactly 60 seconds after the last frame.

Apply both fixes: raise the proxy read timeout for the WebSocket location, and have the server ping every 20 to 30 seconds so no hop sees an idle connection. The heartbeat is the durable fix, because the next proxy will have its own timer. Confirm that the 60-second spike leaves the duration histogram, and keep client reconnects with backoff and resubscribe (WebSocket reconnection strategies), because some disconnects are genuine.

Failure modes and trade-offs

  • Debugging the wrong layer. Rewriting reconnect logic for a proxy timeout. Always reproduce with a command-line client first.
  • Frame logging in production. It is slow, enormous and leaks message content, including tokens. Log lifecycle events in production; frames only in staging or under a time-limited debug flag for one connection.
  • Missing close reasons. Servers that close without a code and reason throw away the cheapest diagnostic there is. Always send a specific code and a short reason.
  • Captures with secrets. Key logs and pcaps contain session data; store them like credentials and delete them after the investigation.
  • Tool-specific behaviour. Command-line clients may not send the same Origin, cookies or compression offer as the browser, so a passing wscat test is evidence about the path, not proof that the browser will succeed.

What to do next

  1. Write the curl handshake test for each WebSocket endpoint into your runbook, with the expected 101 response.
  2. Install wscat or websocat on a jump host inside the network so you can test past the load balancer.
  3. Add one structured log line per connection at close time with duration, counts, close code, reason and initiator.
  4. Graph connection duration as a histogram and alert on new spikes at a fixed duration.
  5. Set a server ping interval shorter than the smallest idle timeout in your path, and document every proxy's timeout.
  6. Make every intentional server close send a specific code and reason.
  7. Keep a Playwright capture script in the repository for intermittent problems reported from the field.
Key takeaway: Debug WebSockets by layer: reachability and TLS, the Upgrade handshake, frames, your application protocol and the connection lifecycle. Test the handshake with curl, read frames in DevTools, speak the protocol by hand with wscat or websocat, and settle disputes with a Wireshark capture decrypted through SSLKEYLOGFILE. Treat close codes as evidence, especially 1006, which means no close frame arrived, and log one structured line per connection so fixed-duration disconnects reveal the proxy timer behind them.