An MCP server has two kinds of users, and a test suite has to serve both. The first is the client software, which needs valid JSON-RPC, correct schemas and predictable errors. The second is the model, which reads your tool names, descriptions and results and decides what to call. A server can pass every protocol check and still fail the model, for example because an error message tells it nothing, or because a renamed argument silently breaks an agent that worked yesterday.
This article lays out a layered test strategy for MCP servers, with runnable examples using the MCP Python SDK v2, where the high-level server class is MCPServer and a single Client class can connect to a server object in memory, to a subprocess or to a URL. If you are new to how tools are declared, start with MCP tools.
A test pyramid for MCP servers
Each layer catches a different class of bug. Unit tests catch wrong business logic. In-memory protocol tests catch wrong results, error mapping and serialization. Contract snapshots catch accidental breaking changes to what clients and models see. Transport tests catch deployment problems: routing, authentication, session handling and proxies. Model evals catch the problem no other layer can see, which is whether a real model understands your tools well enough to use them.
The common mistake is to test only at the top, by chatting with the server in a desktop client. That is slow, cannot run in CI, and fails for reasons unrelated to your code. The second most common mistake is to test only at the bottom and never check the protocol surface, which is where most MCP-specific bugs live.
Layer 1: keep tool logic testable without MCP
Write each tool as a thin adapter around a plain function or service. The adapter handles MCP concerns: argument types, turning domain exceptions into tool errors, and shaping results. The plain function holds the logic and has ordinary unit tests with fakes for databases and APIs. This split keeps the fast tests fast, and it means a change of protocol version or SDK major version touches the adapter, not the logic.
The example server below follows that shape. lookup_order lives in a separate module. The tool function maps a missing order to a ToolError with a message written for the model, and returns a Pydantic model so the SDK can publish an output schema and return structured content alongside the text.
# orders_server.py
from mcp.server import MCPServer
from mcp.server.mcpserver.exceptions import ToolError
from pydantic import BaseModel
from orders_core import OrderNotFound, lookup_order # plain business logic, tested on its own
mcp = MCPServer("orders")
class OrderStatus(BaseModel):
order_id: int
status: str
total: float
@mcp.tool()
async def get_order_status(order_id: int) -> OrderStatus:
"""Return the status and total of one order by its numeric id."""
try:
o = await lookup_order(order_id)
except OrderNotFound:
raise ToolError(f"No order {order_id}. Order ids are numeric, e.g. 1182.")
return OrderStatus(order_id=o.id, status=o.status, total=o.total)
Layer 2: in-memory protocol tests
The Python SDK's Client accepts the server object itself and talks to it in process. The request still goes through the SDK's protocol handling, including initialization, argument validation against the input schema, the tool wrapper and result serialization, but there is no subprocess or socket. Tests run in milliseconds and are deterministic. The SDK documentation notes that the in-memory client is era-neutral: it probes the server and picks the appropriate protocol version, currently 2026-07-28 or the older 2025-11-25.
Pass raise_exceptions=True in most tests. Without it, failures outside the tool body are sanitized the way they are in production, which hides the real traceback while you debug. That brings us to the test most suites forget. Because the flag changes error behaviour, at least one test must run without it and prove that production error handling leaks nothing.
# test_orders_server.py
import pytest
from mcp import Client
import orders_server
from orders_server import mcp
@pytest.fixture
def anyio_backend():
return "asyncio"
@pytest.fixture
def fake_db(monkeypatch):
orders = {1182: type("O", (), {"id": 1182, "status": "REFUNDED", "total": 49.0})()}
async def lookup(order_id):
if order_id not in orders:
raise orders_server.OrderNotFound(order_id)
return orders[order_id]
monkeypatch.setattr(orders_server, "lookup_order", lookup)
@pytest.fixture
async def client(fake_db):
async with Client(mcp, raise_exceptions=True) as c:
yield c
@pytest.mark.anyio
async def test_happy_path_returns_structured_content(client):
result = await client.call_tool("get_order_status", {"order_id": 1182})
assert not result.is_error
assert result.structured_content == {"order_id": 1182, "status": "REFUNDED", "total": 49.0}
@pytest.mark.anyio
async def test_unknown_order_is_a_tool_error_the_model_can_read(client):
result = await client.call_tool("get_order_status", {"order_id": 9})
assert result.is_error
assert "No order 9" in result.content[0].text
@pytest.mark.anyio
async def test_crash_is_sanitized(monkeypatch, fake_db):
async def boom(order_id):
raise RuntimeError("password=hunter2 host=db-7.internal")
monkeypatch.setattr(orders_server, "lookup_order", boom)
async with Client(mcp) as c: # production behaviour: no raise_exceptions
result = await c.call_tool("get_order_status", {"order_id": 1182})
assert result.is_error
text = " ".join(getattr(b, "text", "") for b in result.content)
assert "hunter2" not in text and "db-7" not in textThe three tests cover the three outcomes a tool call can have. A success returns is_error false and structured_content that matches the output model. Python attributes are snake_case in SDK v2, while the wire format stays camelCase. A ToolError returns a normal result with is_error true and the message in content, which the model reads and can recover from. Any other exception is treated as a crash, and the model sees only a generic "Error executing tool" message while the traceback goes to the server log. The sanitization test plants a fake secret in the exception and asserts that it never reaches the result.
Testing protocol errors and bad input
There is a fourth outcome. When the request itself should be rejected, for example because the client lacks a capability the tool needs, the SDK's MCPError fails the whole tools/call with a JSON-RPC error. There is no result for the model, and the host application handles it. The SDK's guidance is a useful rule for choosing between them: if a smarter model could have avoided the failure by passing different arguments, raise ToolError; otherwise raise MCPError. Test each deliberate MCPError by asserting that the client call raises and that the error code is the one you chose, such as INVALID_PARAMS. Check the SDK API reference for the exact exception the client raises in your version.
Then test hostile and sloppy input, because models produce both. Send a string where an integer is expected, omit required arguments, add unknown ones, pass empty strings, huge strings and negative ids, and call a tool that does not exist. For each case, assert the outcome class and that the message would help a model fix its call. pytest.mark.parametrize over a table of bad inputs keeps this compact. For the error semantics in the protocol itself, see MCP error handling.
Layer 3: snapshot the contract
Your tools/list response is an API. Clients cache it, agents are prompted around it, and models learn its wording. Renaming an argument, tightening a type or rewording a description can break a working agent without breaking any unit test. A contract snapshot turns those changes into an explicit review step.
import json, pathlib
SNAPSHOT = pathlib.Path(__file__).with_name("tools_contract.json")
@pytest.mark.anyio
async def test_tool_contract_unchanged(client):
tools = (await client.list_tools()).tools
current = {t.name: {"description": t.description, "input_schema": t.input_schema,
"output_schema": t.output_schema} for t in tools}
if not SNAPSHOT.exists(): # first run writes the baseline
SNAPSHOT.write_text(json.dumps(current, indent=2, sort_keys=True))
assert current == json.loads(SNAPSHOT.read_text()), "tool contract changed: review and update"The first run writes a baseline file that you commit. Every later run compares against it, and a diff in a pull request shows exactly what changed. Treat changes the way you would treat a REST API. Adding an optional argument or a new tool is safe. Removing or renaming a tool or argument, making an optional argument required, or changing the output schema of structured results is breaking, and deserves a new tool name or a deprecation period. The structured content article explains why output schemas matter to clients that validate results. Include description text in the snapshot too: description edits change model behaviour, so they should be reviewed as deliberately as schema edits.
Layer 4: transport, auth and deployment
In-memory tests bypass everything between the client and your handler, which is also where deployment bugs live: the HTTP route, session handling, authentication middleware, reverse proxies that buffer streaming responses, and CORS or origin checks. Run a small suite against the real server process on its real transport, for every build.
The same Client class connects to a URL, so the layer 2 tests can be reused with a different fixture. The MCP Inspector also has a CLI mode that makes a good smoke test in any CI system, because it needs no test code:
# Build step: start the real server on its HTTP transport, then probe it like a client would
python -m orders_server --http --port 8765 & # your own entry point
sleep 2
npx @modelcontextprotocol/inspector --cli http://localhost:8765/mcp \
--transport http --method tools/list
npx @modelcontextprotocol/inspector --cli http://localhost:8765/mcp \
--transport http --method tools/call \
--tool-name get_order_status --tool-arg order_id=1182Pass --transport http explicitly for Streamable HTTP servers rather than relying on a default. The Inspector CLI also accepts --config and --server to reuse a server definition from a config file. For authenticated servers, add negative tests: no token, an expired token, and a valid token for the wrong tenant should each be rejected before any tool runs. The MCP security model describes what should be enforced at this layer.
Layer 5: does a model use the tools correctly?
The final question cannot be answered by asserting on JSON: given a realistic request, does a model pick the right tool, pass sensible arguments, and recover from errors? Answer it with a small eval set. Write twenty to fifty prompts that represent real requests, including ambiguous ones and ones that should call no tool at all. Run each through a model connected to your server, record the trajectory of tool calls, and grade it: the right tool, valid arguments, no forbidden calls, and a correct final answer where you can check it.
Treat the results as statistics, not pass or fail. Run each prompt several times, report the success rate per prompt, and compare against the last release. Keep the backend faked, so that evals measure tool usability and not the state of a staging database. This is where description changes prove their worth: a rewording that raises the right-tool rate from 80 to 95 percent is a real improvement, and one that lowers it is a regression the contract snapshot could only flag, not judge.
Worked example: a regression caught at each layer
Suppose a developer changes get_order_status to accept order_ref as a string instead of a numeric order_id, and wraps the database call in a broad except Exception that returns the error text as the result. Here is what each layer catches. The unit tests still pass, because the logic is fine. The layer 2 happy-path test fails because the argument name changed. The not-found test fails because the error now comes back as a normal result with is_error false, which a model would read as an answer. The sanitization test fails because the raw exception text, with a hostname in it, now reaches the result. The contract snapshot shows a breaking rename. Nothing reaches layer 4. If the rename was intended, the snapshot diff forces a decision to ship it as a new tool, and the layer 5 evals confirm that models still find it.
Failure modes of MCP test suites
| Anti-pattern | Consequence | Fix |
|---|---|---|
| Only manual testing in a chat client | Slow, flaky, nothing runs in CI | Layers 1-3 on every commit |
| All tests use raise_exceptions=True | Leaky production errors go unnoticed | At least one sanitization test without it |
| Asserting on full text output | Tests break on harmless wording changes | Assert on is_error and structured_content |
| No contract snapshot | Silent breaking renames | Snapshot tools/list, including descriptions |
| Evals against live backends | Scores reflect data, not tool design | Fake backends; fixed fixtures |
| Single run per eval prompt | Noise mistaken for regressions | Repeat runs; compare rates |
What to do next
- Split each tool into a thin MCP adapter and a plain function, and unit-test the function.
- Add an in-memory
Client(mcp)fixture and one test per outcome: success, tool error, crash. - Write one test without
raise_exceptionsthat plants a secret and proves it never reaches the result. - Parametrize a table of bad inputs and assert each gives a helpful, correctly flagged error.
- Commit a
tools/listsnapshot and require review of every diff to it. - Add an Inspector CLI smoke test with
--transport httpagainst the built server, plus auth negative tests. - Build a small model-in-the-loop eval set, run it before releases, and track the right-tool rate.