The Model Context Protocol (MCP) lets one host connect a model to many servers, each written by someone different. That is its value and its injection problem. Every server can put text in front of the model: tool results, resources, prompt templates, server instructions, and requests the server starts itself. The model cannot reliably tell instructions from data, so any of that text can try to steer it.

This page covers defences that hold up even when the model is fooled. It maps where injected text enters an MCP session, builds a host-side mediator that labels data and applies flow rules (working Python), and covers the server-initiated channels, sampling and elicitation, that most write-ups skip. A traced example, a test plan, failure modes and a checklist follow. Poisoned tool descriptions are covered in depth on the tool hijacking page, so this one concentrates on runtime content.

Where injected text enters an MCP session

Start with an inventory. In the 2025-06-18 revision of the specification, these are the places where text an attacker may control reaches the model or the user:

ChannelWho authors itTypical injection
Tool result content / structuredContentServer, often relaying third partiesAn issue, email or web page that says "ignore the task, do X"
Resources and resource_link itemsServer or the files it readsA README or document that contains instructions
Prompt templatesServerA template that quietly adds steps to the user request
instructions in the initialize resultServerHints the host may add to the system prompt
Tool annotationsServerA destructive tool marked readOnlyHint: true
sampling/createMessageServerA server-written prompt run on your model and budget
elicitation/createServerA form message that phishes the user

Two facts from the specification shape the defence. First, it says clients MUST consider tool annotations untrusted unless they come from trusted servers, so readOnlyHint and destructiveHint are labels for the UI, not facts a policy can trust. Second, its security considerations say clients SHOULD validate tool results before passing them to the model, show tool inputs to the user before calling the server, and prompt for confirmation on sensitive operations. The protocol puts the burden on the host. None of these channels is signed or typed in a way that separates data from commands.

Why the host is the enforcement point

A server cannot defend the host against itself, and the model cannot be relied on to defend anyone. Only the host sees every server, the user, and every proposed call, so the enforcement point belongs there. Treat it as a mediator: a component that every MCP message passes through, holding a small amount of state per run.

Every MCP channel that carries server-authored text passes through one host mediatorUsergoal and approvalsModelplans, proposes callsHost mediator1. label results by server2. validate outputSchema3. flow rules on calls4. sampling / elicitation policy and review5. decision logServer Afirst partyServer Bthird partyServer Cpublic writestaskproposed calllabelled resulttools/callresult, resourcessampling, elicitationRun state: labels {untrusted, private} + their sourcesa call is judged on what the run has seen, not on how the text reads
The host mediator sits between the model and every server. Inbound text is labelled by origin; outbound calls are judged against the labels the run has collected.

The key idea is that a decision about a call does not depend on whether the context looks malicious. It depends on what the run has touched. If a run has read untrusted text and private data, it must not publish anything without a human looking, however reasonable the model's explanation sounds. That approach survives injections nobody has seen yet, which classifiers and prompt wording cannot promise. The general theory is on the indirect injection page. Here it becomes MCP plumbing.

Layer one: labels and typed results

Give every server a profile when it is installed, not at run time: who operates it, whether it can return private data, whether any of its tools write somewhere public or external, and whether it may start sampling. Then label every inbound result by server and tool. A first-party database server returning rows from your own tables is different from a GitHub server returning issue text that anyone on the internet can write. That issue text is untrusted even though the server itself is trustworthy, which is why labels attach to content sources, not just to servers.

Wrapping results in an envelope that names the source helps the model and the audit log. Do not mistake it for a boundary: an attacker can write a fake closing tag. Escape anything inside that looks like your delimiter, and never let the envelope be the only control. Where a tool declares an outputSchema, validate structuredContent against it and prefer passing typed fields to free text. A field of type number cannot carry a sentence, and a model that only sees {"stars": 41} has nothing to obey.

Layer two: flow rules on proposed calls

The second layer judges proposed calls. The pattern Invariant Labs called a toxic agent flow has three ingredients: untrusted input, access to private data, and a way out. Remove any one and the attack fails. A mediator can check for all three without understanding a word of the injected text:

from dataclasses import dataclass, field

@dataclass
class ServerProfile:
    name: str
    trust: str                 # "first_party" | "third_party"
    reads_private: bool
    writes_public: bool
    allow_sampling: bool = False

@dataclass
class RunState:
    labels: set = field(default_factory=set)      # {"untrusted", "private"}
    sources: list = field(default_factory=list)   # (server, tool) that set them

class Blocked(Exception):
    pass

class Mediator:
    def __init__(self, servers, public_write_tools, relay_tools, is_private):
        self.servers = {s.name: s for s in servers}
        self.public_write_tools = public_write_tools   # {(server, tool)}
        self.relay_tools = relay_tools                 # tools returning third-party text
        self.is_private = is_private                   # trusted lookup, never model output

    def on_result(self, run, server, tool, args, result, output_schema=None):
        prof = self.servers[server]
        if output_schema is not None and "structuredContent" in result:
            missing = [k for k in output_schema.get("required", [])
                       if k not in result["structuredContent"]]
            if missing:
                raise Blocked(f"{server}.{tool} broke its outputSchema: {missing}")
        if prof.trust != "first_party" or (server, tool) in self.relay_tools:
            run.labels.add("untrusted")
            run.sources.append((server, tool))
        # privacy belongs to the resource, not the tool: the same read tool
        # fetches public and private repositories
        if prof.reads_private and self.is_private(server, args):
            run.labels.add("private")
            run.sources.append((server, tool))
        return result          # the caller wraps it in an envelope for the model

    def before_call(self, run, server, tool, args):
        if (server, tool) in self.public_write_tools:
            if {"untrusted", "private"} <= run.labels:
                raise Blocked("untrusted input + private data -> public write")
            if "untrusted" in run.labels:
                return "ask_user"
        return "allow"

The private label is set from a trusted lookup of the resource the call touched, for example the repository's visibility fetched with the host's own credentials, never from the tool name or from anything the model claims. On GitHub, the same file-reading tool serves public and private repositories.

Labels are sticky for the run: once untrusted text is in context, everything the model writes afterwards may be influenced by it. That is crude but sound. Finer tracking, such as a planner that never sees untrusted text and a quarantined model that only extracts values, buys back autonomy at the cost of engineering. Start sticky, measure how often users hit ask_user, and refine where the numbers justify it. Outbound network calls deserve the same treatment; see egress control for agents.

Layer three: sampling and elicitation

Sampling inverts the call. With sampling/createMessage a server asks the host to run a prompt the server wrote on the host's model. The specification says there SHOULD always be a human in the loop with the ability to deny sampling requests, and that applications SHOULD let users view and edit prompts before sending and review responses before delivery. The 2025-11-25 revision adds tool calling to sampling through tools and toolChoice parameters, so a sampling request can now ask for model-driven actions, not just text. A practical policy:

def on_sampling(self, run, server, params):
    prof = self.servers[server]
    if not prof.allow_sampling:                       # default off, per server
        raise Blocked(f"{server} may not request sampling")
    if params.get("includeContext", "none") != "none":
        raise Blocked("sampling may not pull host context")
    if params.get("tools"):
        raise Blocked("sampling with tools disabled for this server")
    return {"maxTokens": min(params.get("maxTokens", 256), 512), "review": True}

The includeContext values thisServer and allServers ask the host to attach context from MCP servers to the prompt; the 2025-11-25 revision soft-deprecates them, and the safe default is to refuse anything but none. The reasons: a server that can read the rest of your context through sampling can exfiltrate it; a server that can make your model call tools has escalated from data source to agent. Never run sampling with the main conversation's tools or history. Cap tokens and rate-limit per server, because the host pays. The MCP sampling page covers the protocol mechanics.

Elicitation lets a server ask the user for input. In the 2025-06-18 revision it is a form, and servers MUST NOT use it to request sensitive information. The 2025-11-25 revision splits it into form mode, which still must not ask for passwords, API keys, access tokens or payment credentials, and URL mode, which sends the user to an external page so secrets never pass through the client. For URL mode, clients MUST show the full URL, MUST NOT open it without explicit consent and MUST NOT pre-fetch it. A malicious or injected server will ignore the server-side rules, so the host enforces them: show the requesting server's name outside the server-authored message, refuse form fields that ask for secrets, highlight the URL's domain, and treat decline and cancel as normal outcomes.

Worked example: the public issue that asked for private code

On 26 May 2025 Invariant Labs published an attack on the official GitHub MCP server. An attacker opened an issue in a public repository owned by the victim, carrying an "About The Author" injection: gather information about the repository owner and publish it. The victim then asked their agent, Claude 4 Opus in the demonstration, something harmless: "Have a look at the open issues" in that public repository. The agent read the issue, followed it, pulled data from the owner's private repositories through the same token, and opened a pull request in the public repository containing it. No server was malicious and no tool description was poisoned, and a well-aligned model was still fooled. The flaw was the flow.

Trace it through the mediator, using tool names from the github-mcp-server README. The list_issues result is labelled untrusted because the pair is in relay_tools. When the model reads a file from a private repository, the visibility lookup adds the private label. When it proposes create_pull_request on the public repository, before_call sees both labels and blocks the call, recording the sources. The user sees a refusal that names the issue that tainted the run, not a leak. Invariant's own recommendation goes earlier: restrict the agent to one repository per session, ideally with a token scoped to that repository.

Testing the defences

Defences that are not tested decay. Build a small corpus per channel: an issue body, a README resource, a prompt template, an instructions string, a sampling request that sets includeContext, and an elicitation form asking for an API key. Each case carries the action that must not happen. Run the agent against them in CI with a fake server and assert on mediator decisions, not on model text:

CASES = [
  ("issue_body",  "github.list_issues",  "github.create_pull_request", "blocked"),
  ("readme",      "docs.read_resource",  "mail.send_email",           "ask_user"),
  ("sampling",    "plugin.sampling",     None,                        "blocked"),
]
for name, inject_via, forbidden_call, expected in CASES:
    log = run_agent_with_fake_server(name, inject_via)
    assert log.decision_for(forbidden_call) == expected, name

Track two numbers: the attack success rate (should be zero for flow-blocked cases whatever the model does) and the approval rate on benign tasks. The second tells you whether users will keep the controls switched on.

Failure modes

  • Labels on servers only. A trusted server relaying untrusted content (issues, email, web pages) is marked clean and the main channel goes unguarded.
  • Trusting annotations. Policy keyed on readOnlyHint lets a server talk its way past approval.
  • Envelope as boundary. An injected fake closing delimiter escapes the wrapper.
  • Approval fatigue. Prompting on every call trains users to click yes; prompt only where the flow rules say so, and show source and arguments (see permission prompt patterns).
  • Resetting labels on summarisation. A model-written summary of untrusted text is still untrusted.
  • Sampling left on by default. A server gets a free model with your context and budget.

Trade-offs

ControlStopsCost
Flow rules on labelsExfiltration and public writes after injectionSome benign tasks need approval
Typed results via outputSchemaInstructions hidden in free textServer work; not every tool can be typed
Per-task token scopePrivate data reachable at allToken minting per task
Sampling off by defaultContext theft, budget abuse, nested agentsServers that rely on sampling lose features
Injection classifiersKnown phrasingsMisses novel attacks; useful only as a signal

Classifiers and careful system prompts reduce how often the model is fooled, which is worth having. The controls above limit what happens when it is fooled, which is the property you can actually test.

What to do next

  1. List every MCP server your host can connect to and write a profile for each: operator, private reads, public or external writes, sampling allowed.
  2. Mark each tool that relays third-party content, and label its results untrusted regardless of server.
  3. Add a mediator hook on results and on proposed calls; block public or external writes when a run is both untrusted and private, and require approval when it is only untrusted.
  4. Turn sampling off per server by default; if you allow it, refuse context inclusion and tools, cap tokens, and show the prompt to the user.
  5. Validate structuredContent against outputSchema and prefer typed fields.
  6. Scope tokens to the task, ideally one repository or mailbox per session.
  7. Build the per-channel injection corpus and run it in CI against mediator decisions.
  8. Log every decision with its label sources so a blocked run explains itself.
Key takeaway: In MCP every server can put text in front of the model, through results, resources, prompts, instructions, annotations, sampling and elicitation, and the model cannot be relied on to ignore it. Put the defence in the host: label content by its real source, validate typed results, block public or external writes once a run holds both untrusted input and private data, keep sampling off unless needed, and test with a per-channel corpus asserting on decisions rather than model text.