The Model Context Protocol connects a language model to tools and data, and that makes it a security boundary in the most literal sense: on one side a model that will follow instructions found in almost any text it reads, on the other side file systems, databases, payment APIs and internal networks. A deployment is secure only if every path from untrusted content to a privileged action has a deterministic control on it.

This article is a threat model. It walks through the attacks that the MCP project's own Security Best Practices page documents, which as of the 2026-07-28 specification revision cover OAuth proxies, token handling, SSRF, state handles, local servers, authorization URLs, mix-up attacks and scopes, and adds the model-layer threats that sit outside the protocol but inside every real deployment. For each one you get the mechanism, the defence and, where it helps, code. The OAuth machinery itself, dynamic registration, PKCE and audience-bound tokens, is explained in MCP authorization architecture, and a shorter overview of the model is in MCP security model.

Advertisement

Three parties and four boundaries

MCP trust boundaries: every arrow that crosses a dashed line carries untrusted dataUser's hostModeltreats all text as possible instructionsMCP client (host)consent, pinning, URL checksLocal MCP serverstdio, sandboxed processFiles, keys, shellwhat a local server can reachAuthorization serverissues audience-bound tokensRemote MCP servervalidates every requestDownstream APIown token, never passthroughInternal network169.254.x, 10.x, localhostWeb contenttool results from the worldtool callsstdioOAuthtools, resultsSSRF riskfetchedDescriptions, results and metadata URLs all come from outside;validate them before a model or a socket sees them.
The host contains the model and the MCP client. Remote servers, authorization servers and everything a tool fetches sit across a trust boundary, and so, less obviously, does a local server process.

An MCP deployment has three kinds of party: the host application, containing the model and the MCP client; MCP servers, local or remote; and authorization servers that issue tokens for remote servers. Four boundaries matter: model to client, because the model's output is shaped by everything in its context; client to each server, because a server may be malicious or compromised; local server to machine, because it is an arbitrary program with the user's privileges; and remote server to its downstream APIs, because it acts on the user's behalf.

Keep one principle in mind throughout: anything that arrives from a server is data, never instructions, and never a trusted URL. Tool descriptions, tool results, error messages and OAuth metadata documents are all authored by the other side of a boundary.

Tokens: the rules the specification already settles

Two server-side rules come straight from the specification. First, an MCP server must not accept any token that was not issued for it: it validates the audience of every token, so a token minted for a different service cannot be replayed against it. Second, token passthrough, forwarding the client's token to a downstream API, is forbidden. A server that needs to call a downstream API holds its own credentials for that API, obtained in its own OAuth flow, which keeps audit trails and revocation meaningful at every hop.

The subtler case is the MCP proxy server, an MCP server that fronts a third-party API using a single static OAuth client ID. If it lets clients register dynamically and the third-party authorization server remembers consent in a cookie, an attacker can register a client with their own redirect URI, send the user a link, and have the consent screen skipped because the cookie says the user already approved the proxy. The best-practices page requires such proxies to keep a per-user registry of approved client IDs and to show their own consent page, naming the client, the scopes and the redirect URI, before forwarding to the third party. Redirect URIs must match exactly, and the OAuth state value must be single-use, short-lived, and stored only after the user has approved consent.

Advertisement

SSRF during OAuth discovery

Connecting to a protected remote server, a client follows URLs the server supplies: the resource metadata URL in the WWW-Authenticate header, the authorization server URLs in that metadata, and the endpoints in the authorization server's metadata. A malicious server can point any of them at http://169.254.169.254/ or an internal admin panel; for a client running in a cloud VPC, that reads instance credentials.

The specification's guidance is to require HTTPS except for loopback during development, block private, loopback and link-local ranges, apply the same checks to every redirect hop, prefer an egress proxy that enforces the policy centrally, and beware of DNS changing between the check and the connection. It also warns explicitly against hand-parsing IP addresses, because octal, hexadecimal and IPv4-mapped IPv6 encodings slip past custom parsers. The sketch below follows that advice by letting the resolver produce addresses and the standard library classify them, then pinning the connection to the checked address:

import ipaddress, socket
from urllib.parse import urlsplit

class UnsafeURL(Exception):
    pass

def resolve_public(url: str, allow_loopback_http: bool = False):
    """Validate an OAuth metadata URL and return (host, pinned_ip, port).

    Connect to pinned_ip (sending the original Host header / SNI) so DNS
    cannot change between this check and the request.
    """
    parts = urlsplit(url)
    if parts.scheme != "https":
        if not (allow_loopback_http and parts.scheme == "http"
                and parts.hostname in ("localhost", "127.0.0.1", "::1")):
            raise UnsafeURL(f"scheme not allowed: {parts.scheme}")
    if not parts.hostname:
        raise UnsafeURL("no host")
    port = parts.port or (443 if parts.scheme == "https" else 80)
    infos = socket.getaddrinfo(parts.hostname, port, type=socket.SOCK_STREAM)
    addrs = {ipaddress.ip_address(info[4][0]) for info in infos}
    for addr in addrs:
        if isinstance(addr, ipaddress.IPv6Address) and addr.ipv4_mapped:
            addr = addr.ipv4_mapped
        if not addr.is_global and not (allow_loopback_http and addr.is_loopback):
            raise UnsafeURL(f"{parts.hostname} resolves to non-public {addr}")
    return parts.hostname, str(sorted(addrs, key=str)[0]), port

# Fetch with redirects disabled; run every Location header through
# resolve_public() again before following it, and cap the number of hops.

Authorization servers that fetch client metadata documents from client-supplied URLs need identical controls. For server-side clients an egress proxy is stronger still, because it covers every HTTP call the process makes.

State handles, not sessions

The 2026-07-28 revision describes MCP as stateless, with no protocol-level sessions. A server that needs state across calls, a shopping cart or a long-running job, mints a handle, returns it in a tool result, and receives it back as an ordinary tool argument. The attack is obvious once stated: a handle that is guessable, or that the server treats as proof of ownership, lets one user operate on another user's state. The specification says servers must verify every inbound request and must not treat possession of a handle as authentication, and recommends random handles bound server-side to the user identity taken from the verified token.

import secrets

def create_cart(token_claims: dict, store) -> dict:
    user = token_claims["sub"]                    # from the verified token, never an argument
    handle = secrets.token_urlsafe(24)            # unguessable, not sequential
    store.put(f"{user}:{handle}", {"items": []}, ttl_seconds=3600)
    return {"cart_id": handle}

def add_item(token_claims: dict, store, cart_id: str, sku: str) -> dict:
    key = f"{token_claims['sub']}:{cart_id}"
    cart = store.get(key)
    if cart is None:                              # wrong owner and unknown handle look identical
        raise PermissionError("unknown cart")
    cart["items"].append(sku)
    store.put(key, cart, ttl_seconds=3600)
    return cart

Keying storage by the token's subject plus the handle makes a stolen or guessed handle useless to anyone else, and one error for unknown and foreign handles avoids leaking which exist. Earlier revisions' server-assigned session IDs need the same discipline.

Local servers: consent, sandboxes and stdio proxies

A local MCP server is a program the client launches, usually from a command line in a configuration file. That makes one-click installation a code-execution feature. The best-practices page requires a client that supports one-click local server configuration to show the exact, untruncated command before running it, flag it as executing code on the user's machine, and require explicit approval. It recommends highlighting dangerous patterns such as sudo, rm -rf or access to SSH keys, and running servers in a sandbox with minimal file system and network access, granting more only on explicit request.

Servers meant to run locally should prefer the stdio transport, which only the launching client can talk to. A local server that listens on HTTP is reachable by any process on the machine and, through DNS rebinding, by web pages in the user's browser; if it must use HTTP, it should require a token or use a restricted IPC channel such as a Unix domain socket.

A local proxy between a web-based client and stdio servers adds an escalation path: cross-site scripting in the client can steal the proxy's token and spawn arbitrary commands. Prevent the XSS, and sandbox and log whatever the proxy spawns.

Authorization URLs, mix-up and scopes

A malicious server can also supply an authorization URL with a javascript: scheme, hoping the client passes it to a browser API, or a string that a shell will interpret as extra commands. Clients must allow only https, or http for loopback in development, must reject other schemes, and must never open URLs by invoking a shell; use the platform's non-shell URL opener. Web-based clients should add a Content Security Policy that forbids inline script.

In a mix-up attack, one hostile authorization server among the many a client uses tries to obtain a code issued by an honest one. PKCE does not help here, because the client would send its code verifier to the attacker's token endpoint. The defence is to record which authorization server a flow was started with and check the iss parameter in the authorization response against it before redeeming the code.

Finally, scopes. A token that carries every scope a server offers turns any leak into full compromise. The recommended model is progressive: start with a minimal read-only scope set, and when a privileged tool is first called, have the server answer with a WWW-Authenticate challenge naming the extra scope, so the client asks the user for it at the moment it is needed. Servers should not advertise omnibus scopes, and should still enforce authorization logic server-side rather than trusting the scope string alone.

The model layer: poisoned tools and poisoned results

The harder threats live at the model layer; they are not protocol requirements but the ecosystem's recurring attack classes. Tool poisoning hides instructions in a tool's description, which the model reads but the user rarely does: a weather tool whose description says to read the user's SSH config and pass it as a parameter. Definition drift, sometimes called a rug pull, is a server that presents a harmless tool at approval time and changes its description or schema later. Indirect prompt injection arrives through tool results: a fetched web page or email that tells the model to call another tool.

No filter reliably detects injected instructions, so design for the case where the model has been subverted. Pin tool definitions: fingerprint each tool's name, description and input schema at approval time, and require a fresh approval when any of them changes.

import hashlib, json

def fingerprint(tool: dict) -> str:
    canonical = json.dumps({k: tool.get(k) for k in ("name", "description", "inputSchema")},
                           sort_keys=True, separators=(",", ":"))
    return hashlib.sha256(canonical.encode()).hexdigest()

def review_tools(server_id: str, listed: list, approved: dict) -> list:
    """Expose only tools whose definitions match what the user approved."""
    usable, changed = [], []
    for tool in listed:
        fp = fingerprint(tool)
        if approved.get((server_id, tool["name"])) == fp:
            usable.append(tool)
        else:
            changed.append(tool["name"])          # new or modified: ask again
    if changed:
        request_user_review(server_id, changed)
    return usable

Then limit what a subverted model can do. Namespace tools by server so a malicious server cannot shadow a trusted tool's name. Require human confirmation for irreversible or externally visible actions, and treat annotations such as read-only or destructive hints as untrusted unless the server itself is trusted; MCP tool annotations architecture covers where those hints can and cannot be enforced. Break the lethal combination of private data access, untrusted content and an outbound channel within one session wherever possible. The general defence architecture, trust separation and least-privilege tools, is in prompt-injection defense architecture.

Worked example: an enterprise MCP gateway

Consider a company that lets employees use a desktop agent with a dozen MCP servers: internal wiki, ticketing, a CRM, a code host and a web fetcher. Rather than enforcing each control in every client, it routes remote servers through a gateway, as described in MCP gateway architecture, and applies the threat model at that choke point.

The gateway obtains a separate, audience-bound token per user per upstream server, so no token crosses servers. It holds an allowlist of servers and pinned tool fingerprints, hiding any changed tool until a reviewer approves it. Outbound OAuth discovery and web-fetch traffic goes through an egress proxy that blocks private ranges. Every tool result is tagged with its origin server, so policy can say that once web-fetcher content enters a conversation, CRM write tools need explicit confirmation. Local servers are limited to an approved set, launched through the consent flow in a container with only the project directory mounted. Every call is logged with user, server, tool, argument digest and token scopes, which is what incident response needs.

Detection and response

  • Alert on unusual sequences, such as a read of sensitive data followed in the same session by a tool that can send data out.
  • Alert on definition changes and elevation events: scope requested, subset granted.
  • Keep revocation fast: revoke one user's token for one server without disrupting the rest.
  • Rehearse a trusted server being compromised and returning injected instructions; confirm which controls stop it.

What to do next

  1. Draw your deployment's trust boundaries and list every path from untrusted content to a privileged action.
  2. On each remote server, validate token audience on every request and remove any token passthrough.
  3. If you run an OAuth proxy server, add per-client consent, exact redirect URI matching and single-use state.
  4. In clients, enforce HTTPS and block non-public addresses for all discovery URLs and redirects, or route through an egress proxy.
  5. Bind every state handle to the authenticated user and make handles unguessable.
  6. Show exact commands and require approval before launching local servers; sandbox them.
  7. Validate authorization URL schemes, never open URLs through a shell, and check iss on authorization responses.
  8. Move to minimal initial scopes with step-up challenges.
  9. Pin tool definitions, namespace tools, and require confirmation for irreversible actions.
Key takeaway: MCP security is a set of deterministic controls placed on every path from untrusted data to a privileged action. Servers validate token audience, never pass tokens through, bind state handles to users and request scopes progressively. Clients treat metadata URLs, authorization URLs, tool definitions and tool results as hostile input: block SSRF, validate schemes, check the issuer, pin tool definitions and gate local server launches behind explicit consent and a sandbox. Because the model itself can be subverted, confirmation for irreversible actions and logs good enough to investigate are part of the design, not optional extras.