The stdio transport is how most local MCP servers run. The host application starts the server as a child process, writes JSON-RPC messages to its standard input, and reads replies from its standard output. There is no port, no TLS and no login: the operating system's process model provides the connection, the identity and the lifetime. That simplicity is why the specification says clients should support stdio whenever possible, and it is also why stdio bugs look strange, because they live in pipes, buffers, encodings and process trees rather than in the protocol.
This article treats stdio as process engineering. It lists what the specification actually requires, then builds a server loop and a client in plain Python, explains why an unread stderr pipe can freeze a session, walks through the shutdown sequence and process-tree cleanup, and shows a small proxy that records every message for debugging. The protocol-level comparison of transports is in MCP transport architecture, and the remote alternative in Streamable HTTP. Requirements quoted here come from the MCP specification revision 2025-11-25.
What the specification requires
The stdio section of the transports page is short, and every line of it matters in practice.
- The client launches the server as a subprocess. The server reads JSON-RPC messages from stdin and writes messages to stdout.
- Messages are individual JSON-RPC requests, notifications or responses, delimited by newlines, and they must not contain embedded newlines.
- The server may write UTF-8 strings to stderr for logging. The client may capture, forward or ignore them, and should not assume stderr output indicates an error.
- The server must not write anything to stdout that is not a valid MCP message, and the client must not write anything to stdin that is not a valid MCP message.
- All JSON-RPC messages must be UTF-8 encoded.
Two more rules live on other pages. The lifecycle page says the client should shut down by closing the server's stdin, waiting for it to exit, then sending SIGTERM, then SIGKILL, and that the server may end the session by closing stdout and exiting. The authorization page says implementations using stdio should not follow the OAuth-based authorization flow and should instead retrieve credentials from the environment. Everything else in this article is engineering practice for meeting those rules reliably.
Launching the server
A host needs three things to start a server: a command, its arguments and an environment. Many hosts read them from a configuration file shaped roughly like the one below; the exact file name and keys depend on the host, so check its documentation.
{
"servers": {
"tickets": {
"command": "/opt/mcp/tickets-server",
"args": ["--read-only", "--project", "OPS"],
"env": { "TICKETS_TOKEN": "${secret:tickets_token}" }
}
}
}Three details cause most launch failures. First, the child inherits the host's PATH, not your shell's; a desktop application started from a dock often has a minimal PATH, so a bare node that works in a terminal fails with file-not-found. Use absolute paths. Second, on Windows, tools installed by npm are .cmd batch shims that process APIs bypassing the shell cannot execute directly, so launching npx needs the shim extension or a cmd /c wrapper. Third, passing the host's whole environment leaks unrelated secrets to every server; pass a minimal base plus the server's own variables.
Credentials follow from the authorization rule: the server reads its token from the environment at startup, scoped to what its tools need.
A server loop that cannot corrupt the stream
The most common stdio bug is a stray write to stdout: a debug print, a library banner, a progress bar. The receiver tries to parse it as JSON and the session breaks. The defence is structural: keep a private handle to the real stdout for protocol messages, and point the language's default stdout at stderr so stray prints land in the log instead. Work in bytes, so the platform's text-mode newline translation and default encoding never touch the protocol.
import json
import sys
PROTO_OUT = sys.stdout.buffer # private handle: only protocol messages go here
sys.stdout = sys.stderr # any stray print() now lands in the log channel
def send(msg: dict) -> None:
# Compact separators and json.dumps escaping guarantee no raw newline inside the message.
data = json.dumps(msg, separators=(",", ":"), ensure_ascii=False).encode("utf-8")
PROTO_OUT.write(data + b"\n")
PROTO_OUT.flush() # pipes are block-buffered: flush every message
def main() -> None:
for raw in sys.stdin.buffer: # loop ends at EOF, i.e. when the client closes our stdin
line = raw.rstrip(b"\r\n")
if not line:
continue
try:
msg = json.loads(line)
except ValueError:
send({"jsonrpc": "2.0", "id": None,
"error": {"code": -32700, "message": "Parse error"}})
continue
reply = handle(msg) # returns None for notifications
if reply is not None:
send(reply)
sys.exit(0) # EOF on stdin is the polite shutdown signalThe flush matters as much as the redirect. When stdout is a pipe rather than a terminal, most runtimes buffer output in blocks, so an unflushed response can sit in the server's memory while the client waits for it until it times out. Stripping a trailing carriage return on input costs nothing and makes the reader tolerant of peers on Windows whose text-mode stdout turned a newline into CRLF. This loop is sequential; a real server dispatches long tool calls to worker tasks so that a ping or a cancellation can be read while a tool runs.
The client: correlation, timeouts and a size limit
On the client side, the transport's job is to write messages without interleaving them, read lines and route each one: responses go to the waiting caller by id, while server requests and notifications go to handlers. The sketch below uses only asyncio.
import asyncio, itertools, json, logging, os, signal, subprocess
log = logging.getLogger("mcp.stdio")
class StdioClient:
# on_server_message(msg) is supplied by the host: it answers server requests
# such as sampling and handles notifications such as progress and logging.
def __init__(self, name, argv, env):
self.name, self.argv, self.env = name, argv, env
self.ids = itertools.count(1)
self.pending = {}
self.write_lock = asyncio.Lock()
async def start(self):
self.proc = await asyncio.create_subprocess_exec(
*self.argv, env=self.env,
stdin=asyncio.subprocess.PIPE, stdout=asyncio.subprocess.PIPE,
stderr=asyncio.subprocess.PIPE,
limit=16 * 1024 * 1024, # default line limit is 64 KiB
start_new_session=(os.name != "nt")) # own process group, for clean kills
self.tasks = [asyncio.create_task(self._read_stdout()),
asyncio.create_task(self._drain_stderr())]
async def _send(self, msg):
data = json.dumps(msg, separators=(",", ":")).encode("utf-8") + b"\n"
async with self.write_lock: # one whole line at a time
self.proc.stdin.write(data)
await self.proc.stdin.drain() # backpressure if the server is slow
async def request(self, method, params=None, timeout=30.0):
rid = next(self.ids)
fut = asyncio.get_running_loop().create_future()
self.pending[rid] = fut
await self._send({"jsonrpc": "2.0", "id": rid, "method": method, "params": params or {}})
try:
return await asyncio.wait_for(fut, timeout)
except asyncio.TimeoutError:
self.pending.pop(rid, None)
await self._send({"jsonrpc": "2.0", "method": "notifications/cancelled",
"params": {"requestId": rid, "reason": "timeout"}})
raise
async def _read_stdout(self):
while line := await self.proc.stdout.readline():
line = line.rstrip(b"\r\n")
if not line:
continue
msg = json.loads(line)
if "method" not in msg: # a response: route by id
fut = self.pending.pop(msg.get("id"), None)
if fut and not fut.done():
fut.set_result(msg)
else:
asyncio.create_task(self.on_server_message(msg))
for fut in self.pending.values(): # EOF: the server exited or closed stdout
if not fut.done():
fut.set_exception(ConnectionError("server closed stdout"))
self.pending.clear()The limit argument is easy to miss. asyncio's stream reader refuses lines longer than its limit, 64 KiB by default, and a tool that returns a large file or a base64 image exceeds that on the first call. Set the limit explicitly and also cap message size deliberately, so a misbehaving server cannot make the client buffer gigabytes. The timeout path sends a cancellation notification, as the lifecycle page recommends; the race between a late response and the cancellation is covered in MCP cancellation. Before the initialize response arrives the client should send only pings; before the notifications/initialized notification arrives the server should send only pings and logging.
stderr and the pipe-buffer deadlock
A pipe is a fixed-size kernel buffer; on Linux the default capacity is 64 KiB. When it fills, the writer blocks until the reader takes data out. If the client never reads the server's stderr, a chatty server eventually fills that buffer, and its next log write blocks forever. The server is now stuck inside a log call, it stops reading stdin and writing stdout, and from the client's side every request simply times out. Nothing in the protocol traffic explains it.
The fix is to always drain stderr, even if you throw the bytes away. Read it in chunks rather than lines, so a log line longer than the reader's limit cannot crash the drain task, and forward it to the host's log with the server's name attached.
async def _drain_stderr(self):
while chunk := await self.proc.stderr.read(65536):
log.info("[%s] %s", self.name, chunk.decode("utf-8", "replace").rstrip())If you do not want the output at all, start the child with stderr redirected to the null device instead of a pipe. Never leave an unread pipe attached. Treat stderr as diagnostics only: the specification says it does not signal errors, so do not fail a session because a server wrote a warning. Structured, level-filtered logs that the host can show the user belong in MCP logging notifications instead.
Shutdown and process trees
The recommended shutdown is an escalation. Close the server's stdin; a well-behaved server sees end-of-file, finishes or abandons in-flight work and exits. If it has not exited within a grace period, send SIGTERM; if it still has not exited, send SIGKILL.
async def close(self, grace=5.0):
self.proc.stdin.close() # step 1: EOF on the server's stdin
for step in ("SIGTERM", "SIGKILL", None):
try:
await asyncio.wait_for(self.proc.wait(), grace)
return self.proc.returncode
except asyncio.TimeoutError:
if step is None:
raise RuntimeError("server did not exit after SIGKILL")
self._signal(step)
def _signal(self, name):
if os.name == "nt":
# No POSIX signals here: terminate the whole tree so launcher children die too.
subprocess.run(["taskkill", "/T", "/F", "/PID", str(self.proc.pid)],
capture_output=True)
else:
os.killpg(self.proc.pid, getattr(signal, name)) # the group made by start_new_sessionSignalling the process group rather than the single process is the important part. Many servers are started through a launcher: a package runner that starts a runtime, which starts the real server. Killing only the launcher leaves an orphaned grandchild still holding files, ports or database connections, and after a few restarts the machine has a dozen of them. Starting the child in its own session on POSIX and killing the tree on Windows avoids that. Also handle the other direction: if the server exits on its own, the reader sees EOF, fails every pending request, and the host should report the exit code and the last lines of stderr instead of a generic timeout.
Worked example: a server that fails at startup
A user adds a server and the host reports a parse error immediately, before any tool appears. To see the bytes, put a small tee proxy in front of the server. It is the command the host launches; it starts the real server and copies both directions through unchanged while appending each line to a log file.
#!/usr/bin/env python3
# Usage as the configured command: mcp-tee /tmp/mcp.log -- /opt/mcp/tickets-server --read-only
import subprocess, sys, threading, time
log_path = sys.argv[1]
argv = sys.argv[sys.argv.index("--") + 1:]
log, lock = open(log_path, "ab"), threading.Lock()
child = subprocess.Popen(argv, stdin=subprocess.PIPE, stdout=subprocess.PIPE) # stderr passes through
def pump(src, dst, tag):
for line in iter(src.readline, b""):
with lock:
log.write(b"%.3f %b " % (time.time(), tag) + line)
log.flush()
dst.write(line)
dst.flush()
dst.close() # propagate EOF in this direction
threading.Thread(target=pump, args=(sys.stdin.buffer, child.stdin, b">>"), daemon=True).start()
pump(child.stdout, sys.stdout.buffer, b"<<")
sys.exit(child.wait())The log shows the client's initialize request going in, then a line << Connected to tickets API v3 coming out before the JSON response. A dependency prints a banner to stdout when it connects. The fix is the redirect from the server-loop section, or configuring that library's logger to use stderr. With the banner gone, the log shows the expected sequence: initialize, its result carrying the negotiated protocolVersion, notifications/initialized, then tools/list.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| Parse error at startup | Banner or print on stdout | Redirect default stdout to stderr; keep a private protocol handle |
| Requests time out after a while, server alive | Unread stderr pipe filled and blocked the server | Always drain stderr, or redirect it to the null device |
| Response never arrives, server log shows it was sent | Output not flushed from a block buffer | Flush after every message |
| Orphaned processes after restarts | Only the launcher was killed | Own process group; kill the tree |
| Garbled non-ASCII text | Platform default encoding used on a text stream | Read and write bytes, encode as UTF-8 explicitly |
Security and trade-offs
A stdio server runs with the full privileges of the user who started the host, so adding one to a configuration file is installing software. Review what each server executes, pin versions rather than letting a package runner fetch the latest release on every start, and scope its credentials. The wider threat model is in MCP security.
Against Streamable HTTP, stdio gives one client per process, no network exposure and trivial deployment, at the cost of a process per user per server and a lifetime tied to the host. Choose HTTP when many users need the same server or it must outlive any one client.
What to do next
- Audit your server for any path that writes to stdout outside the send function, including third-party libraries, and add the stdout-to-stderr redirect at startup.
- Flush after every message and confirm that EOF on stdin makes the server exit cleanly.
- In your client, drain stderr continuously, set an explicit line limit and enforce a per-request timeout that sends a cancellation.
- Start servers in their own process group and implement the close, SIGTERM, SIGKILL escalation with a tree kill on Windows.
- Use absolute paths and a minimal environment in server configuration, with credentials resolved from a secret store and scoped per server.
- Keep a tee proxy in your toolbox and use it before guessing at any startup or hang problem.