Raw WebSockets give you a pipe that carries bytes in both directions. Applications want more than that: call a named method on the server and get a result back, have the server call a method on one client, a group of clients or every client, fall back gracefully when a proxy blocks WebSockets, and survive a dropped connection. ASP.NET Core SignalR packages exactly that set of features, and its design has become a reference shape for real-time frameworks generally.
This article dissects that shape. It uses SignalR's published protocol specifications and documentation for concrete wire formats and defaults, then shows how to build the same pattern elsewhere. The pattern, rather than the .NET product, is the subject; Socket.IO solves the same problems with different choices, compared in Socket.IO at scale.
The pattern in one picture
Four ideas define a SignalR-style system. Hubs expose named methods that clients invoke, and clients register named handlers that the server invokes; both directions are remote procedure calls over one connection. Transport negotiation picks WebSockets when possible and falls back to Server-Sent Events or long polling when something in the path prevents it. Addressing lets server code target a connection, an authenticated user across all their devices, a named group, or everyone. A backplane forwards those sends between servers so addressing works in a farm.
Application code is written in terms of methods and audiences, not sockets; the cost is a runtime with its own timeouts, limits and scale-out rules.
Connection lifecycle: negotiate, transport, handshake
A connection starts with an HTTP POST to the hub URL plus /negotiate. In negotiate version 1 the response carries a public connectionId, which other code may use to address this client, a secret connectionToken, the negotiated version, and a list of availableTransports with the transfer formats each supports. The distinction matters: HTTP-based transports identify the connection by the token in an id query parameter, so the token must not leak, while the connection id is safe to share. Negotiate can instead return a url and accessToken to redirect the client elsewhere, which is how a managed service takes over client connections.
The client then opens the first transport it and the server both support, in the order WebSockets, ServerSentEvents, LongPolling. SSE and long polling are half transports: they carry server-to-client traffic, and the client sends with separate HTTP POSTs. With long polling, the client issues a GET that the server holds open until it has messages; a 200 delivers them, a 204 signals that the connection is closing, and a 404 means the connection no longer exists. The background on both fallbacks is in long-polling fallback and Server-Sent Events.
Over whichever transport is chosen, the client first sends a handshake naming the hub protocol, json or messagepack, with version 1, and the server answers with an empty object or an error. In the JSON protocol every message is a JSON object terminated by the ASCII record separator, byte 0x1E, shown here as <RS>:
-> {"protocol":"json","version":1}<RS>
<- {}<RS>
-> {"type":1,"invocationId":"7","target":"PlaceOrder","arguments":[{"sku":"A-12","qty":2}]}<RS>
<- {"type":1,"target":"OrderUpdated","arguments":[{"id":"42","status":"accepted"}]}<RS>
<- {"type":3,"invocationId":"7","result":{"id":"42"}}<RS>
<- {"type":6}<RS>The client invokes PlaceOrder with id 7. Before the completion arrives, the server pushes OrderUpdated with no invocationId, because it wants no reply. Then the completion for id 7 carries the result. Calls interleave freely; the invocation id pairs each completion with its call.
The hub protocol's message types
| Type | Name | Purpose |
|---|---|---|
| 1 | Invocation | Call a method; with an invocationId a Completion is expected, without one it is fire-and-forget |
| 2 | StreamItem | One item of a streaming result or argument |
| 3 | Completion | Result or error for an invocation, or end of a stream |
| 4 | StreamInvocation | Call a method that returns a stream |
| 5 | CancelInvocation | Client cancels a stream it started |
| 6 | Ping | Keepalive, no payload |
| 7 | Close | Orderly shutdown, optional error and reconnect hint |
| 8 | Ack | Stateful reconnect: messages received up to a sequenceId |
| 9 | Sequence | Stateful reconnect: first message after reconnecting, giving the sequenceId sending resumes from |
Two consequences are worth internalising. First, fire-and-forget invocations give no delivery guarantee: the JavaScript client's send resolves when the message has been written, not when the server has processed it. Use invoke when you need to know. Second, streaming is first class, so long-running server output such as progress or a token stream from a model can flow as StreamItems instead of being chopped into separate invocations.
Hubs, groups and users in code
On the server, a strongly typed hub declares what clients can call and an interface declares what the server can call on clients. Groups are named sets of connections, and users are all connections whose authenticated identity maps to the same user id.
public interface IOrderClient
{
Task OrderUpdated(OrderDto order);
}
[Authorize]
public class OrdersHub : Hub<IOrderClient>
{
private readonly IOrderAccess _access;
public OrdersHub(IOrderAccess access) => _access = access;
public async Task WatchOrder(string orderId)
{
// Authorise per resource: joining a group is a subscription, not a formality.
if (!await _access.CanView(Context.UserIdentifier, orderId))
throw new HubException("forbidden");
await Groups.AddToGroupAsync(Context.ConnectionId, "order:" + orderId);
}
}
// Elsewhere in the app, outside any hub method:
public class ShippingEvents(IHubContext<OrdersHub, IOrderClient> hub)
{
public Task Shipped(OrderDto o) => hub.Clients.Group("order:" + o.Id).OrderUpdated(o);
}const connection = new signalR.HubConnectionBuilder()
.withUrl("/hubs/orders", { accessTokenFactory: () => auth.getAccessToken() })
.withAutomaticReconnect() // waits 0, 2, 10, 30 s, then gives up
.build();
connection.on("OrderUpdated", order => render(order)); // register before start
connection.onreconnected(async () => {
for (const id of watchedOrders) await connection.invoke("WatchOrder", id);
});
await connection.start();
await connection.invoke("WatchOrder", "42");Two details in that client are deliberate. Handlers are registered before start so no early message is dropped. And group subscriptions are replayed after reconnecting, because automatic reconnect produces a connection the server treats as entirely new, with a new connection id, and group membership belongs to connections. If you forget this, the UI reconnects and silently stops updating.
Authentication has a transport wrinkle. Browsers cannot set an Authorization header on a WebSocket or EventSource request, so the JavaScript client sends the access token from accessTokenFactory as a query string parameter for those transports, and the server's bearer handler must be configured to read it there. Keep tokens short-lived and make sure proxies and access logs do not record query strings on the hub path.
Keepalives and timeouts
A persistent connection that carries no traffic is indistinguishable from a dead one. SignalR's defaults are a server KeepAliveInterval of 15 seconds and a ClientTimeoutInterval of 30 seconds; the JavaScript client also pings every 15 seconds and times out the server after 30. The rule that keeps these consistent is that each side's timeout should be about double the other side's keepalive interval, so one lost ping does not kill the connection. Change one side and you must change the other.
Intermediaries add their own idle limits. A reverse proxy or cloud load balancer that closes idle connections after 60 seconds is harmless with 15-second pings, but if you lengthen the keepalive to save traffic, the proxy will start cutting connections first. Long polling also needs the proxy's read timeout to exceed the server's poll duration. The general discipline is covered in heartbeat design. Note too the default MaximumReceiveMessageSize of 32 KB: large client uploads belong on ordinary HTTP endpoints, not hub methods.
Reconnect: automatic and stateful
Automatic reconnect, as configured above, retries after 0, 2, 10 and 30 seconds and then fires onclose. It does not retry a failed initial start, so the start call needs its own retry loop with backoff and jitter. Because each reconnect is a new connection, messages sent while the client was away are lost unless the application recovers them, typically by fetching current state over HTTP after onreconnected or by replaying from a per-user sequence number.
Stateful reconnect, added in .NET 8, narrows that gap for short drops. When enabled on the server endpoint with AllowStatefulReconnects and on the client with withStatefulReconnect, each side buffers unacknowledged messages, acknowledges with Ack messages, and after reconnecting sends a Sequence message stating where it resumes; duplicates are possible and must be ignored. The buffer defaults to 100,000 bytes. It helps a phone that changes networks for a second; it does not survive a server restart or a reconnect that lands on a different server, so it complements application-level recovery rather than replacing it.
Scale-out: sticky sessions and backplanes
SignalR requires every HTTP request for one connection to reach the same server process, because negotiate, the SSE or long-polling GETs and the client's POSTs all refer to in-memory connection state. In a farm that means sticky sessions. Microsoft's documentation lists only three exceptions: a single server, a managed service such as Azure SignalR Service that holds the client connections itself, or clients restricted to WebSockets with negotiation skipped, since one WebSocket is one long-lived request.
Stickiness solves routing, not addressing. A send to a group must reach members held by every server, which is what the backplane does: with the Redis backplane, each server publishes sends through Redis pub/sub and every server delivers to its local members. That makes every broadcast a cross-server operation, so backplane traffic grows with servers times message rate, and you must still scale servers by connection count. The alternative is a managed service that terminates client connections and leaves your servers with a few connections to it, letting you scale on message volume. The underlying capacity questions, per-node limits, rebalancing and reconnect storms, are worked through in WebSocket scaling patterns.
Building the pattern without SignalR
The pattern is small enough to implement on any WebSocket library. The core is framing, a method table, invocation ids, and an addressing layer. The sketch below omits authentication, negotiate, the handshake and transport fallback, and assumes an object ws with async recv and send methods.
import asyncio, json
from collections import defaultdict
RS = chr(0x1E)
METHODS = {} # name -> async fn(conn, *args)
GROUPS = defaultdict(set) # group -> set of Conn
CONNS = {} # connection id -> Conn
TASKS = set() # strong refs so running dispatch tasks are not collected
def hub_method(fn):
METHODS[fn.__name__] = fn
return fn
class Conn:
def __init__(self, cid, ws):
self.cid, self.ws, self.groups = cid, ws, set()
async def call(self, target, *args): # server -> client, fire-and-forget
await self.ws.send(json.dumps({"type": 1, "target": target, "arguments": list(args)}) + RS)
async def serve(cid, ws):
conn = CONNS[cid] = Conn(cid, ws)
buf = ""
try:
while True:
buf += await ws.recv()
*frames, buf = buf.split(RS) # keep any partial frame
for f in frames:
t = asyncio.create_task(dispatch(conn, json.loads(f)))
TASKS.add(t)
t.add_done_callback(TASKS.discard)
finally: # disconnect: drop memberships
for g in conn.groups:
GROUPS[g].discard(conn)
CONNS.pop(cid, None)
async def dispatch(conn, msg):
if msg.get("type") == 6:
return # ping
fn, inv = METHODS.get(msg.get("target")), msg.get("invocationId")
try:
if fn is None:
raise LookupError("unknown method")
result = await fn(conn, *msg.get("arguments", []))
reply = {"type": 3, "invocationId": inv, "result": result}
except Exception as e:
reply = {"type": 3, "invocationId": inv, "error": str(e)}
if inv is not None: # no id means no completion
await conn.ws.send(json.dumps(reply) + RS)
async def send_group(group, target, *args): # a backplane would publish here
await asyncio.gather(*(c.call(target, *args) for c in list(GROUPS[group])),
return_exceptions=True)
@hub_method
async def WatchOrder(conn, order_id):
conn.groups.add("order:" + order_id)
GROUPS["order:" + order_id].add(conn)Even this sketch surfaces the real design work: dispatching each invocation as a task means completions can arrive out of order, so clients must match on id; one slow client must not stall a group send, hence gather with exceptions captured, and a production version needs per-connection send queues with limits; and disconnect cleanup must be in a finally block or groups leak dead connections.
Failure modes
- Silent loss of subscriptions after reconnect. Group membership is per connection. Replay subscriptions in onreconnected and refetch state.
- Negotiate succeeds, transport fails. Without stickiness, the WebSocket or SSE request lands on a server that never saw the negotiate, and the connection fails with a not-found error. Check affinity first when failures are intermittent and scale with server count.
- Mismatched timeouts. Changing keepalive on one side only, or a proxy idle timeout below the ping interval, produces regular disconnects at a fixed period. Look for that period in your logs.
- Reconnect storms. A server restart drops thousands of clients, which all retry on the same schedule. Use a retry policy with jitter and drain servers gradually.
- Tokens in URLs. Query-string access tokens end up in logs. Scrub them and keep lifetimes short.
Trade-offs
| Choice | SignalR-style framework | Socket.IO | Raw WebSocket plus your protocol |
|---|---|---|---|
| Programming model | Typed RPC both ways, streams | Named events with acknowledgements | Whatever you build |
| Fallback transports | SSE, long polling | HTTP long polling | None unless you add it |
| Interoperability | Published protocol, official clients | Own protocol, many clients | Fully yours |
| Scale-out | Sticky sessions plus backplane or service | Sticky sessions plus adapter | Your design |
| Overhead | Runtime and protocol layer | Runtime and protocol layer | Lowest, most work |
What to do next
- Draw your connection path, client to proxy to load balancer to server, and write down each hop's idle timeout next to your keepalive and client timeout values.
- Confirm session affinity is configured, or that every client is WebSockets-only with negotiation skipped.
- Replay group subscriptions and refetch state in onreconnected, then test by killing a server under load.
- Replace the default retry delays with a jittered policy and add a retry loop around the initial start.
- Decide where access tokens are read from on the hub path and scrub them from logs.
- Load-test a broadcast to your largest group across all servers and measure backplane throughput.