The Model Context Protocol (MCP) lets one host connect a model to many servers, each written by someone different. That is its value and its injection problem. Every server can put text in front of the model: tool results, resources, prompt templates, server instructions, and requests the server starts itself. The model cannot reliably tell instructions from data, so any of that text can try to steer it.
This page covers defences that hold up even when the model is fooled. It maps where injected text enters an MCP session, builds a host-side mediator that labels data and applies flow rules (working Python), and covers the server-initiated channels, sampling and elicitation, that most write-ups skip. A traced example, a test plan, failure modes and a checklist follow. Poisoned tool descriptions are covered in depth on the tool hijacking page, so this one concentrates on runtime content.
Where injected text enters an MCP session
Start with an inventory. In the 2025-06-18 revision of the specification, these are the places where text an attacker may control reaches the model or the user:
| Channel | Who authors it | Typical injection |
|---|---|---|
Tool result content / structuredContent | Server, often relaying third parties | An issue, email or web page that says "ignore the task, do X" |
Resources and resource_link items | Server or the files it reads | A README or document that contains instructions |
| Prompt templates | Server | A template that quietly adds steps to the user request |
instructions in the initialize result | Server | Hints the host may add to the system prompt |
| Tool annotations | Server | A destructive tool marked readOnlyHint: true |
sampling/createMessage | Server | A server-written prompt run on your model and budget |
elicitation/create | Server | A form message that phishes the user |
Two facts from the specification shape the defence. First, it says clients MUST consider tool annotations untrusted unless they come from trusted servers, so readOnlyHint and destructiveHint are labels for the UI, not facts a policy can trust. Second, its security considerations say clients SHOULD validate tool results before passing them to the model, show tool inputs to the user before calling the server, and prompt for confirmation on sensitive operations. The protocol puts the burden on the host. None of these channels is signed or typed in a way that separates data from commands.
Why the host is the enforcement point
A server cannot defend the host against itself, and the model cannot be relied on to defend anyone. Only the host sees every server, the user, and every proposed call, so the enforcement point belongs there. Treat it as a mediator: a component that every MCP message passes through, holding a small amount of state per run.
The key idea is that a decision about a call does not depend on whether the context looks malicious. It depends on what the run has touched. If a run has read untrusted text and private data, it must not publish anything without a human looking, however reasonable the model's explanation sounds. That approach survives injections nobody has seen yet, which classifiers and prompt wording cannot promise. The general theory is on the indirect injection page. Here it becomes MCP plumbing.
Layer one: labels and typed results
Give every server a profile when it is installed, not at run time: who operates it, whether it can return private data, whether any of its tools write somewhere public or external, and whether it may start sampling. Then label every inbound result by server and tool. A first-party database server returning rows from your own tables is different from a GitHub server returning issue text that anyone on the internet can write. That issue text is untrusted even though the server itself is trustworthy, which is why labels attach to content sources, not just to servers.
Wrapping results in an envelope that names the source helps the model and the audit log. Do not mistake it for a boundary: an attacker can write a fake closing tag. Escape anything inside that looks like your delimiter, and never let the envelope be the only control. Where a tool declares an outputSchema, validate structuredContent against it and prefer passing typed fields to free text. A field of type number cannot carry a sentence, and a model that only sees {"stars": 41} has nothing to obey.
Layer two: flow rules on proposed calls
The second layer judges proposed calls. The pattern Invariant Labs called a toxic agent flow has three ingredients: untrusted input, access to private data, and a way out. Remove any one and the attack fails. A mediator can check for all three without understanding a word of the injected text:
from dataclasses import dataclass, field
@dataclass
class ServerProfile:
name: str
trust: str # "first_party" | "third_party"
reads_private: bool
writes_public: bool
allow_sampling: bool = False
@dataclass
class RunState:
labels: set = field(default_factory=set) # {"untrusted", "private"}
sources: list = field(default_factory=list) # (server, tool) that set them
class Blocked(Exception):
pass
class Mediator:
def __init__(self, servers, public_write_tools, relay_tools, is_private):
self.servers = {s.name: s for s in servers}
self.public_write_tools = public_write_tools # {(server, tool)}
self.relay_tools = relay_tools # tools returning third-party text
self.is_private = is_private # trusted lookup, never model output
def on_result(self, run, server, tool, args, result, output_schema=None):
prof = self.servers[server]
if output_schema is not None and "structuredContent" in result:
missing = [k for k in output_schema.get("required", [])
if k not in result["structuredContent"]]
if missing:
raise Blocked(f"{server}.{tool} broke its outputSchema: {missing}")
if prof.trust != "first_party" or (server, tool) in self.relay_tools:
run.labels.add("untrusted")
run.sources.append((server, tool))
# privacy belongs to the resource, not the tool: the same read tool
# fetches public and private repositories
if prof.reads_private and self.is_private(server, args):
run.labels.add("private")
run.sources.append((server, tool))
return result # the caller wraps it in an envelope for the model
def before_call(self, run, server, tool, args):
if (server, tool) in self.public_write_tools:
if {"untrusted", "private"} <= run.labels:
raise Blocked("untrusted input + private data -> public write")
if "untrusted" in run.labels:
return "ask_user"
return "allow"The private label is set from a trusted lookup of the resource the call touched, for example the repository's visibility fetched with the host's own credentials, never from the tool name or from anything the model claims. On GitHub, the same file-reading tool serves public and private repositories.
Labels are sticky for the run: once untrusted text is in context, everything the model writes afterwards may be influenced by it. That is crude but sound. Finer tracking, such as a planner that never sees untrusted text and a quarantined model that only extracts values, buys back autonomy at the cost of engineering. Start sticky, measure how often users hit ask_user, and refine where the numbers justify it. Outbound network calls deserve the same treatment; see egress control for agents.
Layer three: sampling and elicitation
Sampling inverts the call. With sampling/createMessage a server asks the host to run a prompt the server wrote on the host's model. The specification says there SHOULD always be a human in the loop with the ability to deny sampling requests, and that applications SHOULD let users view and edit prompts before sending and review responses before delivery. The 2025-11-25 revision adds tool calling to sampling through tools and toolChoice parameters, so a sampling request can now ask for model-driven actions, not just text. A practical policy:
def on_sampling(self, run, server, params):
prof = self.servers[server]
if not prof.allow_sampling: # default off, per server
raise Blocked(f"{server} may not request sampling")
if params.get("includeContext", "none") != "none":
raise Blocked("sampling may not pull host context")
if params.get("tools"):
raise Blocked("sampling with tools disabled for this server")
return {"maxTokens": min(params.get("maxTokens", 256), 512), "review": True}The includeContext values thisServer and allServers ask the host to attach context from MCP servers to the prompt; the 2025-11-25 revision soft-deprecates them, and the safe default is to refuse anything but none. The reasons: a server that can read the rest of your context through sampling can exfiltrate it; a server that can make your model call tools has escalated from data source to agent. Never run sampling with the main conversation's tools or history. Cap tokens and rate-limit per server, because the host pays. The MCP sampling page covers the protocol mechanics.
Elicitation lets a server ask the user for input. In the 2025-06-18 revision it is a form, and servers MUST NOT use it to request sensitive information. The 2025-11-25 revision splits it into form mode, which still must not ask for passwords, API keys, access tokens or payment credentials, and URL mode, which sends the user to an external page so secrets never pass through the client. For URL mode, clients MUST show the full URL, MUST NOT open it without explicit consent and MUST NOT pre-fetch it. A malicious or injected server will ignore the server-side rules, so the host enforces them: show the requesting server's name outside the server-authored message, refuse form fields that ask for secrets, highlight the URL's domain, and treat decline and cancel as normal outcomes.
Worked example: the public issue that asked for private code
On 26 May 2025 Invariant Labs published an attack on the official GitHub MCP server. An attacker opened an issue in a public repository owned by the victim, carrying an "About The Author" injection: gather information about the repository owner and publish it. The victim then asked their agent, Claude 4 Opus in the demonstration, something harmless: "Have a look at the open issues" in that public repository. The agent read the issue, followed it, pulled data from the owner's private repositories through the same token, and opened a pull request in the public repository containing it. No server was malicious and no tool description was poisoned, and a well-aligned model was still fooled. The flaw was the flow.
Trace it through the mediator, using tool names from the github-mcp-server README. The list_issues result is labelled untrusted because the pair is in relay_tools. When the model reads a file from a private repository, the visibility lookup adds the private label. When it proposes create_pull_request on the public repository, before_call sees both labels and blocks the call, recording the sources. The user sees a refusal that names the issue that tainted the run, not a leak. Invariant's own recommendation goes earlier: restrict the agent to one repository per session, ideally with a token scoped to that repository.
Testing the defences
Defences that are not tested decay. Build a small corpus per channel: an issue body, a README resource, a prompt template, an instructions string, a sampling request that sets includeContext, and an elicitation form asking for an API key. Each case carries the action that must not happen. Run the agent against them in CI with a fake server and assert on mediator decisions, not on model text:
CASES = [
("issue_body", "github.list_issues", "github.create_pull_request", "blocked"),
("readme", "docs.read_resource", "mail.send_email", "ask_user"),
("sampling", "plugin.sampling", None, "blocked"),
]
for name, inject_via, forbidden_call, expected in CASES:
log = run_agent_with_fake_server(name, inject_via)
assert log.decision_for(forbidden_call) == expected, nameTrack two numbers: the attack success rate (should be zero for flow-blocked cases whatever the model does) and the approval rate on benign tasks. The second tells you whether users will keep the controls switched on.
Failure modes
- Labels on servers only. A trusted server relaying untrusted content (issues, email, web pages) is marked clean and the main channel goes unguarded.
- Trusting annotations. Policy keyed on
readOnlyHintlets a server talk its way past approval. - Envelope as boundary. An injected fake closing delimiter escapes the wrapper.
- Approval fatigue. Prompting on every call trains users to click yes; prompt only where the flow rules say so, and show source and arguments (see permission prompt patterns).
- Resetting labels on summarisation. A model-written summary of untrusted text is still untrusted.
- Sampling left on by default. A server gets a free model with your context and budget.
Trade-offs
| Control | Stops | Cost |
|---|---|---|
| Flow rules on labels | Exfiltration and public writes after injection | Some benign tasks need approval |
| Typed results via outputSchema | Instructions hidden in free text | Server work; not every tool can be typed |
| Per-task token scope | Private data reachable at all | Token minting per task |
| Sampling off by default | Context theft, budget abuse, nested agents | Servers that rely on sampling lose features |
| Injection classifiers | Known phrasings | Misses novel attacks; useful only as a signal |
Classifiers and careful system prompts reduce how often the model is fooled, which is worth having. The controls above limit what happens when it is fooled, which is the property you can actually test.
What to do next
- List every MCP server your host can connect to and write a profile for each: operator, private reads, public or external writes, sampling allowed.
- Mark each tool that relays third-party content, and label its results untrusted regardless of server.
- Add a mediator hook on results and on proposed calls; block public or external writes when a run is both untrusted and private, and require approval when it is only untrusted.
- Turn sampling off per server by default; if you allow it, refuse context inclusion and tools, cap tokens, and show the prompt to the user.
- Validate
structuredContentagainstoutputSchemaand prefer typed fields. - Scope tokens to the task, ideally one repository or mailbox per session.
- Build the per-channel injection corpus and run it in CI against mediator decisions.
- Log every decision with its label sources so a blocked run explains itself.