The Model Context Protocol (MCP) is usually shown connecting a desktop assistant to cloud services. It is just as useful at the edge: a small model on a phone, kiosk or robot calls tools exposed by a gateway that sits next to sensors, motors and local files, with no cloud round trip. Edge devices bring constraints that desktop tutorials ignore: they reboot, lose the network, run models with short context windows, and control physical things where a wrong call has real consequences.

This article explains where MCP fits in an edge system, how revision 2026-07-28 of the specification changes the picture, and how to build and run a gateway server that a small model can use safely. The protocol details were checked against the published specification on 2026-10-02. SDK support for the new revision varies, so the example server uses only the Python standard library.

Advertisement

Three tiers, and only two of them speak MCP

Three tiers: who speaks MCP at the edgeOn-device hostsmall model + MCP clientphone, kiosk, robot, laptopGateway MCP serverLinux-class box on the LANtools, interlocks, cache, auditMicrocontrollerssensors and actuatorsno MCP: MQTT, Modbus, serialstdio (same box) orStreamable HTTP (LAN)device protocolPer-request (2026-07-28)_meta: protocolVersion, clientInfo, clientCapabilitiesno initialize, no session id: restart = retryEnforced in the gateway, not the modelrange limits, interlocks, rate limitsOrigin check, auth, audit logThe model proposes tool calls; the gateway decides what the hardware actually does.
Figure 1. Microcontrollers keep their native protocols. A Linux-class gateway translates them into MCP tools and enforces safety. The model runs on the host, or on the gateway itself, as an MCP client.

MCP messages are JSON-RPC 2.0 encoded as JSON text, and every tool call needs a JSON Schema check. That is a poor fit for a microcontroller with kilobytes of RAM and a better fit for a Linux-class gateway or single-board computer. So the usual design has three tiers. Microcontrollers speak what they already speak: MQTT, Modbus, BLE or a serial line. A gateway runs an MCP server that turns those devices into a small set of tools. The host runs the model and an MCP client; on a compact device the host and gateway can be the same box.

The small models in IoT guide covers choosing and running the model itself. This article is about the protocol layer between that model and the hardware.

Why the 2026-07-28 revision suits devices

Earlier revisions, up to 2025-11-25, opened every connection with an initialize handshake and, on HTTP, could bind state to an Mcp-Session-Id header. Revision 2026-07-28 removes both. Every request now carries its own protocol version and client capabilities in _meta under the keys io.modelcontextprotocol/protocolVersion, io.modelcontextprotocol/clientCapabilities and, recommended, io.modelcontextprotocol/clientInfo. Servers must implement server/discover, which returns supported versions, capabilities and identity in one call.

For edge hardware this matters a lot. A gateway that reboots after a power cut has no session table to rebuild, and the specification says plainly that after an unexpected exit the client restarts the server and retries lost requests. A server that only speaks the new revision rejects an unknown version with UnsupportedProtocolVersionError (code -32022) and lists what it supports in data.supported. If your fleet still has clients or servers on 2025-11-25, the stdio rule is to probe with server/discover first and fall back to initialize only when the reply is not a recognised modern error. The specification also says to cache that result per server process rather than probing on every call.

Advertisement

Choosing a transport

stdioStreamable HTTP
TopologyClient launches the server as a child process on the same deviceServer listens on a single MCP endpoint; clients POST to it
FramingOne JSON-RPC message per line, no embedded newlines; stdout carries only MCP messagesOne request per POST; reply is JSON or a per-request SSE stream
MetadataInline in _meta only_meta plus MCP-Protocol-Version, Mcp-Method and Mcp-Name headers that must match the body
Cancellationnotifications/cancelledClose the response stream
Use at the edgeModel and tools on one box; no open portSeveral hosts on a LAN sharing one gateway

Prefer stdio whenever the model and the gateway share hardware, because it opens no network port. Use Streamable HTTP when phones or several robots share one gateway. Then the specification's security rules apply directly: validate the Origin header and return 403 if it is invalid, bind to 127.0.0.1 when only local clients need access, and authenticate every connection. The remote MCP servers guide covers the HTTP side in detail.

A gateway server in standard-library Python

The server below exposes two tools for a greenhouse controller over stdio. A bridge thread, not shown, subscribes to MQTT and fills LAST with the latest readings. Look at three things: version checking on every request, the required resultType on results, and the caching fields ttlMs and cacheScope that the revision requires on list results.

import json, sys, time

VERSION = "2026-07-28"
INFO = {"name": "greenhouse-gw", "version": "0.4.0"}
LAST = {}                    # sensor_id -> (value, unit, epoch_s), filled by the MQTT bridge
VENT = {"percent": 0, "wind_lock": False}

TOOLS = [                    # fixed order helps client caches and prompt caches
  {"name": "read_sensor",
   "description": "Latest reading of one sensor: value, unit, age_s.",
   "inputSchema": {"type": "object", "required": ["sensor_id"],
     "properties": {"sensor_id": {"type": "string", "enum": ["t1", "t2", "rh1", "soil1"]}}},
   "annotations": {"readOnlyHint": True}},
  {"name": "set_vent",
   "description": "Open the roof vent to 0-100 percent. Refused while the wind lock is on.",
   "inputSchema": {"type": "object", "required": ["percent"],
     "properties": {"percent": {"type": "integer", "minimum": 0, "maximum": 100}}},
   "annotations": {"readOnlyHint": False, "idempotentHint": True}},
]

def reply(id_, result):
    result.setdefault("resultType", "complete")
    result.setdefault("_meta", {})["io.modelcontextprotocol/serverInfo"] = INFO
    return {"jsonrpc": "2.0", "id": id_, "result": result}

def error(id_, code, msg, data=None):
    e = {"code": code, "message": msg}
    if data is not None:
        e["data"] = data
    return {"jsonrpc": "2.0", "id": id_, "error": e}

def tool_result(text, structured, is_error=False):
    return {"content": [{"type": "text", "text": text}],
            "structuredContent": structured, "isError": is_error}

def call_tool(name, args):
    if name == "read_sensor":
        v = LAST.get(args.get("sensor_id"))
        if v is None:
            return tool_result("no reading yet", {"ok": False}, True)
        age = round(time.time() - v[2])
        return tool_result(f"{v[0]} {v[1]}, {age}s old", {"value": v[0], "unit": v[1], "age_s": age})
    if name == "set_vent":
        pct = args.get("percent")
        if not isinstance(pct, int) or not 0 <= pct <= 100:
            return tool_result("percent must be an integer 0-100", {"ok": False}, True)
        if VENT["wind_lock"] and pct > 0:
            return tool_result("refused: wind lock active", {"ok": False}, True)
        VENT["percent"] = pct     # real code publishes to the actuator and waits for an ack
        return tool_result(f"vent set to {pct}%", {"ok": True, "percent": pct})
    return None

def handle(msg):
    id_, method = msg.get("id"), msg.get("method")
    params = msg.get("params") or {}
    if id_ is None:
        return None                               # notification, e.g. notifications/cancelled
    ver = params.get("_meta", {}).get("io.modelcontextprotocol/protocolVersion")
    if ver != VERSION:                            # also catches a legacy initialize
        return error(id_, -32022, "Unsupported protocol version",
                     {"supported": [VERSION], "requested": ver})
    if method == "server/discover":
        return reply(id_, {"supportedVersions": [VERSION], "capabilities": {"tools": {}},
                           "ttlMs": 3600000, "cacheScope": "public"})
    if method == "tools/list":
        return reply(id_, {"tools": TOOLS, "ttlMs": 3600000, "cacheScope": "public"})
    if method == "tools/call":
        res = call_tool(params.get("name"), params.get("arguments") or {})
        if res is None:
            return error(id_, -32602, "Unknown tool")
        return reply(id_, res)
    return error(id_, -32601, "Method not found")

for line in sys.stdin:                            # EOF on stdin is the shutdown signal
    try:
        out = handle(json.loads(line))
    except Exception as exc:                      # never let a bad line kill the loop
        print(f"bad message: {exc}", file=sys.stderr)
        continue
    if out is not None:
        sys.stdout.write(json.dumps(out, separators=(",", ":")) + "\n")
        sys.stdout.flush()

Logs go to stderr, because stdout may carry only MCP messages. The loop ends when stdin closes, which the specification calls the primary graceful shutdown signal. A production version validates arguments against the full schema, handles the actuator acknowledgement with a timeout, and, if you choose dual-era support, answers initialize for legacy clients instead of rejecting it. The MCP security guide covers authorisation in depth.

Designing tools for small models

A small model has a short context window and weaker tool-selection skills than a frontier model. Every tool definition is sent to the model as prompt text, so the tool list is a budget. A practical rule is to expose the fewest tools that cover the job, give each a one-line description that says when to use it, and make arguments hard to get wrong.

  • Use enums and bounds. "enum": ["t1", "t2", "rh1", "soil1"] turns a free-text guess into a choice, and guided decoding can enforce it; see small models for tool calling.
  • Return compact structured results. A value, a unit and an age are enough. Do not return a 2 KB status dump that pushes the conversation out of context.
  • Merge chatty tools. One read_sensor with an enum is better than four near-identical tools.
  • Keep order deterministic. The revision asks servers to return tools in a deterministic order, so clients and prompt caches can reuse the prefix.
  • Cache the list. Honour ttlMs instead of calling tools/list before every turn on a slow link.

Worked example: an overheating greenhouse

A user asks the kiosk: "It feels hot in bay 2, do something." The host model sees two tools. It calls read_sensor with t2 and gets 34.1 C, 12 seconds old. It calls set_vent with 60. The gateway checks the range and the wind lock. Wind is above the lock threshold, so the call returns isError: true with "refused: wind lock active". The model tells the user the vent stays closed because of wind and suggests the shade screen, which a person must operate. No model output reached the motor without passing the gateway's checks.

Now the gateway reboots mid-conversation. Under the old handshake, the client would need to notice the broken session and re-initialise. Under 2026-07-28, the client restarts the stdio server and re-sends the lost request with a new id. Because set_vent is idempotent, a retry after an uncertain outcome is safe. Mark tools that are not idempotent, and have the gateway reject duplicates using a request key you pass as an argument.

Safety, security and operations

Put every physical limit in the gateway. The model may be wrong, prompt-injected through a sensor label or a document, or simply confused, so ranges, interlocks, rate limits and allowed operating hours belong in code that the model cannot change. Write each actuation to an append-only audit log with the client identity and arguments. For a destructive action, a person should confirm. Under 2026-07-28 a server asks for input by returning an interim result with resultType: "input_required" that carries the request, and the client retries the original call with the answer; check that your client supports that pattern before relying on it.

Operate the gateway like any other service. Put a timeout on every actuator acknowledgement and report a timeout as an error, never as success. Serve cached readings with their age, so the model can tell a stale value from a fresh one when the field bus drops. Propagate trace context: the revision documents traceparent in _meta, which lets you follow one user request from the host to the motor. Note also that Roots, Sampling and Logging are deprecated in this revision. New edge servers should take paths through tool arguments or configuration, log to stderr or OpenTelemetry, and not depend on the client's model.

Failure modes

SymptomCauseFix
Client hangs at start-upModern client, legacy server that ignores unknown methodsProbe with server/discover and a timeout, then fall back to initialize
Every call fails with -32022Client sends a version the gateway does not supportRead data.supported and retry; pin versions in fleet config
Garbled stream, parse errorsA library printed to stdoutRedirect all logging to stderr; test with a strict line parser
Model picks the wrong tool or invents argumentsToo many tools, vague descriptions, free-text argumentsFewer tools, enums, bounds, guided decoding
Actuator moves twiceRetry after an unknown outcome on a non-idempotent toolIdempotency keys and duplicate rejection in the gateway
Remote web page drives the gatewayHTTP server on 0.0.0.0 without Origin checks or authBind localhost, validate Origin, require auth
Confident answers from stale dataBus outage; cache returns old values silentlyReturn the age with every reading; refuse above a maximum age

What to do next

  1. Draw your three tiers and decide which box runs the MCP server and which runs the model.
  2. List every physical action and write its limits and interlocks as gateway code before you write any tool description.
  3. Start from the stdlib server above, replace the bridge with your bus, and test it with a line-by-line JSON client.
  4. Keep the tool list under the budget your model's context allows, and use enums and bounds everywhere.
  5. Decide whether you need dual-era support for 2025-11-25 clients, and implement the discover-then-fallback probe if you do.
  6. Kill the gateway mid-request in testing and confirm that clients restart, retry and never repeat a non-idempotent action. The edge deployment guide covers packaging and updates for the host side.
Key takeaway: At the edge, MCP belongs on a Linux-class gateway that turns microcontroller protocols into a small set of well-bounded tools, with the model on the host as the client. Revision 2026-07-28 fits devices well: there is no handshake or session, so a reboot means restart and retry. Use stdio when model and tools share a box and Streamable HTTP with Origin checks, localhost binding and authentication when they do not. Keep the tool list small for small models, and enforce every physical limit in gateway code rather than in the prompt.