The first remote transport in the Model Context Protocol was HTTP with Server-Sent Events, defined in protocol version 2024-11-05. It let an MCP server run as an independent web service rather than a subprocess, which is what made hosted MCP servers possible. Protocol version 2025-03-26 replaced it with Streamable HTTP and marked it deprecated. It has not disappeared: older clients and servers still speak it, many tutorials still teach it, and anyone who runs a public MCP server has to decide how long to keep serving it.

This article treats the old transport as something you operate rather than something you choose. It explains exactly what the specification requires, how Server-Sent Events are framed, what a session looks like on the wire, and why the design is awkward behind load balancers. It then covers security, the compatibility probe the newer specification defines, and a migration plan. The replacement transport itself is covered in MCP streaming transport architecture, and the wider transport landscape in MCP transport architecture.

Advertisement

What the 2024-11-05 specification requires

The specification text for this transport is short, and it is worth knowing exactly what is in it, because most of what people believe about the transport comes from SDK behaviour rather than the spec. The server runs as an independent process that can handle multiple client connections, and it must provide two endpoints: an SSE endpoint, where a client opens a connection and receives messages from the server, and a regular HTTP POST endpoint, where the client sends messages to the server.

When a client connects to the SSE endpoint, the server must first send an event named endpoint whose data is a URI. Every subsequent message from the client must be sent as an HTTP POST to that URI. Every message from the server, whether a response, a request or a notification, is sent on the SSE stream as an event named message with the JSON-RPC message encoded as JSON in the event data.

That is the whole wire contract. Notice what it leaves out: no session header, no event ids, no rule for reconnection and no requirement about what the POST returns. Those gaps are why the transport was replaced.

HTTP+SSE (protocol version 2024-11-05): two endpoints, one long-lived stream per clientMCP clientSSE readerparses eventsRequest senderone POST per messagePending mapJSON-RPC id to futureProxy / load balancermust not buffermust route POSTsto the stream's instanceMCP server instanceSSE endpointGET, held openPOST endpointURI from endpoint eventSession tablesession to stream1. GET opens stream2. event: endpoint, then event: message3. POST JSON-RPC messageResponses to POSTed requests come back on the SSE stream, not in the POST response.Lose the stream and you lose the session: the old transport defines no resumption.
The client holds a GET open to the SSE endpoint and sends each message as a separate POST to the URI it was given. The server routes every reply onto that client's stream, so the POST and the stream must reach the same server state.

Server-Sent Events from first principles

Server-Sent Events is a plain-text streaming format defined by the HTML standard. The server answers a GET with Content-Type: text/event-stream and keeps the response open, writing events as they occur. Each event is a few lines of field: value pairs ended by a blank line. The fields that matter here are event, which names the event type, data, which carries the payload, and id, which sets an event id. A line starting with a colon is a comment and is ignored, which makes it a convenient keepalive.

HTTP/1.1 200 OK
Content-Type: text/event-stream
Cache-Control: no-cache

event: endpoint
data: /messages?sessionId=4f9c2b7e

: keepalive

event: message
data: {"jsonrpc":"2.0","id":1,"result":{"protocolVersion":"2024-11-05","capabilities":{},"serverInfo":{"name":"inventory","version":"1.4.0"}}}

SSE is one-directional: the server can push, but the client cannot reply on the same connection. That is why the transport needs a second endpoint. It also explains a practical limit. The browser EventSource API cannot set custom request headers, so browser-based clients that need an Authorization header on the stream had to use a fetch-based SSE parser instead, or carry credentials some other way.

Advertisement

A session on the wire

Here is a complete initialization over the old transport. The paths /sse and /messages?sessionId= are conventions used by the official SDKs and most servers, not names the specification requires; the client must use whatever URI the endpoint event gives it. Likewise, answering the POST with 202 Accepted and an empty body is common SDK behaviour, not a spec rule.

# 1. Client opens the stream
GET /sse HTTP/1.1
Accept: text/event-stream

# 2. Server's first event tells the client where to POST
event: endpoint
data: /messages?sessionId=4f9c2b7e

# 3. Client sends initialize as a POST to that URI
POST /messages?sessionId=4f9c2b7e HTTP/1.1
Content-Type: application/json

{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2024-11-05","capabilities":{},"clientInfo":{"name":"desk-agent","version":"0.9"}}}

HTTP/1.1 202 Accepted

# 4. The response arrives on the SSE stream, not in the POST reply
event: message
data: {"jsonrpc":"2.0","id":1,"result":{...}}

# 5. Client confirms with a notification, then calls tools the same way
POST /messages?sessionId=4f9c2b7e HTTP/1.1
{"jsonrpc":"2.0","method":"notifications/initialized"}

The client therefore needs a correlation table: when it POSTs a request with id 7, it stores a pending future under id 7 and resolves it when an SSE message event with id 7 arrives. Server-initiated requests, such as sampling, arrive on the same stream and are answered by POSTing a JSON-RPC response.

A minimal server

The server's core data structure is a table from session id to an outbound queue feeding that client's stream. The sketch below uses aiohttp and omits authentication and the MCP method handlers so that the transport mechanics stand out.

import asyncio, json, secrets
from aiohttp import web

sessions: dict[str, asyncio.Queue] = {}

async def sse(request: web.Request) -> web.StreamResponse:
    sid = secrets.token_urlsafe(16)
    queue: asyncio.Queue = asyncio.Queue(maxsize=1000)
    sessions[sid] = queue
    resp = web.StreamResponse(headers={
        "Content-Type": "text/event-stream",
        "Cache-Control": "no-cache",
        "X-Accel-Buffering": "no",          # ask nginx not to buffer
    })
    await resp.prepare(request)
    await resp.write(f"event: endpoint\ndata: /messages?sessionId={sid}\n\n".encode())
    try:
        while True:
            try:
                msg = await asyncio.wait_for(queue.get(), timeout=15)
                await resp.write(f"event: message\ndata: {json.dumps(msg)}\n\n".encode())
            except asyncio.TimeoutError:
                await resp.write(b": keepalive\n\n")
    except ConnectionResetError:
        pass
    finally:
        sessions.pop(sid, None)              # the stream is the session
    return resp

async def messages(request: web.Request) -> web.Response:
    queue = sessions.get(request.query.get("sessionId", ""))
    if queue is None:
        return web.Response(status=404)
    msg = await request.json()
    asyncio.create_task(handle(msg, queue))  # reply later, on the stream
    return web.Response(status=202)

async def handle(msg: dict, queue: asyncio.Queue) -> None:
    if "id" in msg and "method" in msg:
        result = await dispatch(msg["method"], msg.get("params", {}))
        await queue.put({"jsonrpc": "2.0", "id": msg["id"], "result": result})

app = web.Application()
app.add_routes([web.get("/sse", sse), web.post("/messages", messages)])

Two choices here matter in production. The bounded queue means a slow client eventually blocks the handlers producing its messages, which is the backpressure you want rather than unbounded memory growth. And removing the session when the stream ends means a POST after a disconnect gets a 404, which tells the client it must start again. The dispatch function stands for your method handlers.

Why the POST must reach the stream's instance

In the sketch above, the session table lives in one process. Put two replicas behind a round-robin load balancer and a client's GET may land on replica A while its POST lands on replica B, which has never heard of the session and answers 404. This is the defining operational problem of the old transport. There are three ways to deal with it.

ApproachHow it worksCost
Sticky routingLoad balancer hashes on the sessionId query parameter, or uses a cookie, so both requests reach one replicaUneven load, and a replica restart drops every session on it
Shared busAny replica accepts the POST and publishes the message on a channel keyed by session id; the replica holding the stream subscribesExtra hop and a dependency such as Redis pub/sub
Single instanceRun one replica per serverNo horizontal scaling

Proxies cause the second class of trouble. Many reverse proxies buffer responses, so events sit in a buffer and the client sees nothing until it fills. Turn buffering off for the SSE route, either in proxy configuration or with a header the proxy honours, such as X-Accel-Buffering: no for nginx. Idle timeouts on load balancers will also close a quiet stream; send a comment line every 15 to 30 seconds, comfortably inside the shortest timeout on the path. A gateway in front of many servers, as described in MCP gateway architecture, has to apply all of this per route.

What the old transport cannot do

Because the specification defines no event ids or resumption, a dropped stream loses any messages in flight and, in practice, the session itself. The client's only safe move is to open a new stream, receive a new endpoint, and run initialize again. Requests that were outstanding when the stream died must be treated as failed, and the client cannot tell whether a tool call with side effects completed.

ConcernHTTP+SSE (2024-11-05)Streamable HTTP (2025-03-26 and later)
EndpointsTwo: SSE stream plus POST URIOne MCP endpoint accepting POST and GET
Where a response arrivesAlways on the long-lived streamIn the POST reply, as JSON or a short SSE stream
Session identityWhatever the endpoint URI encodesMcp-Session-Id header
ResumptionNot definedOptional, using event ids and Last-Event-ID
Always-open connectionRequiredOptional

The lesson for anyone still running the old transport is to keep tool calls idempotent where you can, so that a client that retries after a dropped stream does no harm. MCP sessions covers what state a server should and should not tie to a connection.

Security on both endpoints

The 2024-11-05 specification carries three explicit warnings: servers must validate the Origin header on all incoming connections to prevent DNS rebinding attacks, should bind only to 127.0.0.1 when running locally, and should authenticate all connections. Without these, a web page in the user's browser could reach a local MCP server.

Apply authentication to both endpoints, not just the stream. A common mistake is to check a token on the GET and then accept any POST that carries a valid session id. Session ids in a query string end up in access logs and proxy logs, so treat them as identifiers, not credentials, and bind each session to the authenticated principal that opened it. Reject a POST whose credentials belong to someone else.

Migrating with the compatibility probe

The 2025-03-26 specification defines how the two transports coexist. A server that wants to keep supporting older clients continues to host the old SSE and POST endpoints alongside the new MCP endpoint. A client that wants to support older servers takes a single URL from the user and probes it: it POSTs an InitializeRequest with an Accept header listing both application/json and text/event-stream. If that succeeds, the server speaks Streamable HTTP. If it fails with a 4xx status, such as 405 or 404, the client issues a GET to the same URL and expects an SSE stream whose first event is endpoint; if it gets one, it uses the old transport from then on.

from urllib.parse import urljoin

async def connect(url: str, http) -> "Transport":
    init = {"jsonrpc": "2.0", "id": 0, "method": "initialize", "params": INIT_PARAMS}
    r = await http.post(url, json=init,
                        headers={"Accept": "application/json, text/event-stream"})
    if r.status_code < 400:
        return StreamableHttpTransport(url, first_response=r)
    if 400 <= r.status_code < 500:
        stream = await open_sse(http, url)           # GET with Accept: text/event-stream
        first = await stream.next_event(timeout=10)
        if first.event == "endpoint":
            post_uri = urljoin(url, first.data)
            return LegacySseTransport(stream, post_uri)  # then send initialize here
    raise ConnectionError(f"{url}: no supported MCP transport ({r.status_code})")

On the server side, plan the deprecation with data. Count sessions per transport and per client name from the clientInfo in initialize, announce a date, and keep the old endpoints until legacy traffic is small enough to accept breaking it.

Failure modes

  • Silent stream. A proxy buffers events and the client times out waiting for initialize. Disable buffering on the route and test through the real proxy chain, not just locally.
  • 404 on POST after scaling out. The POST reached a replica without the session. Add sticky routing on the session id or a shared message bus.
  • Idle disconnects. A load balancer closes quiet streams after its idle timeout. Send keepalive comments.
  • Lost in-flight work. The stream drops during a long tool call and the result is gone. Make tools idempotent and let clients re-initialize and retry.
  • Unbounded queues. A stalled client makes the server buffer without limit. Bound per-session queues and close sessions that stay full.
  • Session hijack. A leaked session id is accepted without a credential check. Authenticate every POST and bind sessions to principals.

What to do next

  1. Inventory your MCP servers and clients and record which transport each speaks; anything on protocol version 2024-11-05 over HTTP is using this transport.
  2. For each server still offering HTTP+SSE, confirm Origin validation, authentication on both endpoints, and sessions bound to the authenticated principal.
  3. Test through the real proxy and load balancer path: disable response buffering on the SSE route, set keepalives inside the shortest idle timeout, and route POSTs to the replica holding the stream.
  4. Add Streamable HTTP alongside the old endpoints, and update clients to probe with a POST initialize and fall back on a 4xx.
  5. Measure sessions per transport and client, announce a retirement date, and remove the old endpoints when legacy traffic is negligible.
  6. Review MCP security for the controls that apply whatever the transport.
Key takeaway: The HTTP+SSE transport from protocol version 2024-11-05 has two endpoints. The client holds a GET open for server messages and POSTs every message it sends to a URI announced in the first endpoint event. Because replies come back on that stream, the POST must reach the same server state, proxies must not buffer, and a dropped stream ends the session, with no resumption defined. Secure both endpoints, run it behind sticky routing or a shared bus while you must, and migrate using the 2025-03-26 probe: POST initialize first, and fall back to GET on a 4xx.