A load test for a REST API sends requests and measures responses. A bidirectional server does not work that way. Its cost is held connections, each with memory, a file descriptor and a heartbeat, plus a flow of messages in both directions and fanout between them. The failures are different too: memory creeping up over hours, a reconnect storm after a deploy, one slow consumer backing up a room. A request-per-second test finds none of these.

k6 handles this well once you model the workload correctly. This article explains what to measure on a bidi server, how to size a test from real traffic, and how to write k6 scenarios for WebSocket and gRPC bidirectional streams using the current k6/websockets and k6/net/grpc stream APIs. It then covers the test types that find real problems, the limits of the load generator itself, and how to read results. General harness design is in load testing architecture; this page is specific to long-lived connections.

Advertisement

What to measure on a bidi server

Think in three dimensions rather than one. Concurrency is how many connections are open at once; it drives memory, file descriptors and heartbeat traffic even when nobody is typing. Message rate is how many messages arrive per second and how many leave after fanout; it drives CPU, broker load and serialization cost. Churn is how fast connections open and close; it drives TLS handshakes, authentication and the per-connection setup work that dominates during a reconnect storm.

The client-side numbers that matter are connection time, the share of connections that fail or drop, and delivery latency: the time from one client sending a message to another client receiving it. HTTP response time is almost irrelevant. On the server, watch open connections, memory per connection, event-loop or scheduler lag, outbound queue depth per connection and the broker's own latency. Outbound queues deserve particular attention, because a slow consumer is where bidi servers run out of memory; backpressure in bidirectional streams covers the mechanisms under test.

The test architecture

A bidi load test: generators, the system under test, and two sets of metricsk6 generator AVUs, many conns eachk6 generator Bsecond source IPLoad balanceridle timeout, stickinessBidi server 1WebSocket or gRPCBidi server 2Bidi server 3Brokerpub/sub for fanoutMetrics storek6 output + server statspublishclient metricsconns, heap, loop lag, queue depthClient metrics say what users would feel. Server metrics say why.A test without both only tells you that something broke, not what.
Two or more k6 generators open connections through the same load balancer real clients use. Client metrics from k6 and server metrics from the fleet land in one store, on one time axis.

Test through the real path, including the load balancer, because idle timeouts, connection stickiness and TLS termination there are common failure points. Many cloud load balancers close idle connections after a default timeout, 60 seconds for AWS Application Load Balancers, so a test with heartbeats slower than that will measure disconnects that are an artifact of configuration. Heartbeat design is in heartbeats and keepalive.

Advertisement

Sizing the workload: a worked example

Start from production numbers, not round figures. Suppose a chat service expects 50,000 concurrent connections at peak. Each active client sends one message every 10 seconds, and rooms average 20 members, so each message fans out to 20 recipients.

QuantityCalculationResult
Inbound messages50,000 / 10 s5,000 msg/s
Outbound deliveries5,000 x 20100,000 msg/s
Outbound bandwidth at 300 bytes per frame100,000 x 300 B30 MB/s, about 240 Mbit/s
Connection churn at 30-minute sessions50,000 / 1,800 sabout 28 new connections/s

Two things stand out. The server's real work is the outbound side, twenty times the inbound, so a test that only counts messages sent is watching the wrong number. And churn is modest in steady state but becomes the dominant cost after an outage, when all 50,000 clients reconnect within seconds.

The generator has limits of its own. A client connection to one destination address and port needs a local ephemeral port, and the default Linux range, 32768 to 60999, gives 28,232 per source IP. 50,000 connections therefore need at least two source IPs, in practice two or more generator machines, each with raised file-descriptor limits. k6's k6/websockets module runs a global event loop per VU, so one VU can hold several connections, which is far cheaper than one VU per connection.

A WebSocket scenario

The script below ramps to 25,000 connections per generator, 2,500 VUs holding 10 each, holds for 20 minutes and ramps down. Each connection joins a room, sends a timestamped message every 10 seconds with a random initial offset, and measures delivery latency when the room broadcast comes back. It uses k6/websockets, the stable module; k6/experimental/websockets is deprecated.

import { WebSocket } from 'k6/websockets';
import { setTimeout, setInterval, clearInterval } from 'k6/timers';
import { Trend, Counter } from 'k6/metrics';

const deliveryMs = new Trend('chat_delivery_ms', true);
const lost = new Counter('chat_msgs_lost');
const CONNS_PER_VU = 10;
const SESSION_MS = 5 * 60 * 1000;
const SEND_EVERY_MS = 10 * 1000;

export const options = {
  scenarios: {
    chat: {
      executor: 'ramping-vus',
      startVUs: 0,
      stages: [
        { duration: '10m', target: 2500 },   // 25,000 connections
        { duration: '20m', target: 2500 },
        { duration: '5m', target: 0 },
      ],
      gracefulRampDown: '6m',               // let sessions finish instead of cutting them
    },
  },
  thresholds: {
    ws_connecting: ['p(95)<500'],
    chat_delivery_ms: ['p(95)<250', 'p(99)<1000'],
    chat_msgs_lost: ['count<100'],
  },
};

function session(room) {
  const ws = new WebSocket(`wss://chat.test.internal/ws?room=${room}`);
  const pending = new Map();
  let seq = 0, timer;

  ws.onopen = () => {
    // jitter the first send so 25,000 clients do not tick in lockstep
    setTimeout(() => {
      timer = setInterval(() => {
        const id = `${__VU}-${seq++}`;
        pending.set(id, Date.now());
        ws.send(JSON.stringify({ type: 'msg', id, room, body: 'x'.repeat(200) }));
      }, SEND_EVERY_MS);
    }, Math.random() * SEND_EVERY_MS);
    setTimeout(() => { clearInterval(timer); ws.close(); }, SESSION_MS);
  };

  ws.onmessage = (e) => {
    const m = JSON.parse(e.data);
    const sentAt = pending.get(m.id);          // server echoes the room broadcast to the sender too
    if (sentAt !== undefined) {
      deliveryMs.add(Date.now() - sentAt);
      pending.delete(m.id);
    }
  };

  ws.onclose = () => { lost.add(pending.size); };
}

export default function () {
  const room = `room-${__VU % 1250}`;       // about 20 connections per room
  for (let i = 0; i < CONNS_PER_VU; i++) session(room);
}

Several details are deliberate. The first send is jittered, because thousands of clients started in the same second would otherwise send in synchronized waves no real population produces. Messages carry their own ID, and the pending map turns unanswered IDs into a loss count when the connection closes. gracefulRampDown is longer than a session so ramp-down does not cut sessions mid-flight and report them as failures. The thresholds make the run fail automatically when latency or loss exceeds the service objective, so the test can gate a release.

Measuring delivery latency by echo keeps both timestamps on one machine's clock. Measuring from a sender on generator A to a receiver on generator B needs synchronized clocks, and ordinary NTP error of a millisecond or more can swamp a fast server's real latency, so prefer echo or same-generator pairs.

k6 also records built-in metrics for this module: ws_connecting, ws_sessions, ws_msgs_sent, ws_msgs_received, ws_session_duration and ws_ping. The message counters are useful cross-checks: received divided by sent should approach the fanout factor.

A gRPC bidirectional streaming scenario

For gRPC, k6/net/grpc provides a Stream created from a connected client and a method name, with data, error and end events, write to send and end to half-close. The example drives a conversational assistant service where each client message produces a stream of chunks, and records time to first chunk per turn.

import { Client, Stream } from 'k6/net/grpc';
import { setTimeout } from 'k6/timers';
import { Trend } from 'k6/metrics';

const turnMs = new Trend('assistant_first_chunk_ms', true);
const client = new Client();
client.load(['protos'], 'assistant.proto');         // init context only

export const options = {
  scenarios: {
    streams: { executor: 'constant-vus', vus: 500, duration: '15m' },
  },
  thresholds: { assistant_first_chunk_ms: ['p(95)<400'] },
};

export default function () {
  client.connect('assistant.test.internal:443', {});
  const stream = new Stream(client, 'assistant.v1.Assistant/Converse');
  let sentAt = 0, first = true, done = false;

  stream.on('data', (chunk) => {
    if (first) { turnMs.add(Date.now() - sentAt); first = false; }
    if (chunk.final && !done) {
      sentAt = Date.now(); first = true;
      stream.write({ text: 'and the next step?' });
    }
  });
  stream.on('error', (e) => console.error(`stream error: ${e.message}`));
  stream.on('end', () => client.close());

  sentAt = Date.now();
  stream.write({ text: 'summarise the incident' });
  setTimeout(() => { done = true; stream.end(); }, 60 * 1000);   // half-close after one minute
}

The protobuf definitions load in the init context, while connecting happens inside the VU function. The client writes the next turn only after the server marks a response final, which models a conversational user, so throughput here is limited by the server's response time; if you need a fixed message rate regardless of response time, send on a timer as in the WebSocket script. k6 reports grpc_streams, grpc_streams_msgs_sent and grpc_streams_msgs_received for streams, from version 0.49.0. HTTP/2 flow control shapes what you will see; the protocol side is in gRPC bidi streaming architecture.

Four tests that find real problems

  • Capacity ramp. Increase connections in steps, holding each step long enough for memory to settle, until latency or errors breach the objective. The knee tells you connections per server, and memory at each step gives memory per connection.
  • Soak. Hold realistic load for hours. Leaks in per-connection state, timers never cleared, room maps never pruned, only show over time. Plot server memory against open connections: a flat ratio is healthy, a rising one is a leak.
  • Reconnect storm. With the fleet at steady load, restart a server or drop connections at the load balancer, and measure how long clients take to reconnect and what happens to authentication, TLS CPU and the broker meanwhile. Model real client backoff, including jitter; a test that reconnects instantly measures a worse storm than your clients will cause, and one without jitter hides herd effects.
  • Slow consumer. Make a fraction of VUs read slowly or stop reading, then check that the server's outbound queues stay bounded and that other clients' latency is unaffected. This is where load-shedding policy is proved; see load shedding on bidirectional streams.

Reading the results

Put client and server metrics on one time axis before drawing conclusions. A latency jump that coincides with a garbage collection pause, broker failover or connection ramp step has a different fix from one that tracks CPU.

Check the generator before blaming the server. A k6 process near full CPU delays its own timers and callbacks, which inflates measured latency and slows the send rate, so the server looks slow and lightly loaded at once. Keep generators well below saturation, and add machines rather than VUs when they approach it. Watch for errors that are really generator limits: port exhaustion, file-descriptor limits and DNS resolution under churn.

Finally, compare message counts end to end. If k6 sent 5,000 messages per second and the fleet reports 4,000 inbound, messages are being dropped or queued somewhere between, and the latency percentiles are computed over survivors, which flatters the server.

Failure modes of the test itself

SymptomCauseFix
Connections fail around 28,000 per generatorEphemeral port exhaustion to one destinationMore source IPs or generators; spread across addresses
All clients disconnect at the same intervalLoad balancer idle timeout shorter than heartbeatHeartbeat below the timeout, or raise it to match production
Latency rises while server CPU is lowGenerator saturated, delaying its own timersMore generators, fewer connections per process
Session failures at the end of every runRamp-down cuts live sessionsgracefulRampDown longer than session length
Periodic spikes every 10 sSynchronized clientsJitter first send and reconnect backoff

What to do next

  1. Pull peak concurrency, per-client message rate, fanout and session length from production, and compute inbound, outbound and churn rates.
  2. Plan generators from the port and file-descriptor limits, with at least two source IPs past about 28,000 connections to one address.
  3. Write a k6/websockets or Stream scenario from k6/net/grpc, with jitter, several connections per VU, a delivery-latency Trend, a loss Counter and thresholds.
  4. Collect server connections, memory, loop lag, outbound queue depth and broker latency on the same time axis as k6 output.
  5. Run a capacity ramp, a multi-hour soak, a reconnect storm and a slow-consumer test, and record connections per server and memory per connection.
  6. Wire the thresholds into CI or the release process so a regression fails the build.
Key takeaway: Bidi servers are loaded by concurrency, message flow and churn, not requests per second, so a useful k6 test models all three. Size from production numbers, remembering that outbound fanout is where the work is. Use k6/websockets with several connections per VU and jittered sends, or the Stream class from k6/net/grpc, measure delivery latency on one clock, count losses, and set thresholds so the run passes or fails on its own. Run capacity, soak, reconnect-storm and slow-consumer tests, collect server metrics beside client metrics, and check the generator's own limits before blaming the server.