An MCP server has two kinds of users, and a test suite has to serve both. The first is the client software, which needs valid JSON-RPC, correct schemas and predictable errors. The second is the model, which reads your tool names, descriptions and results and decides what to call. A server can pass every protocol check and still fail the model, for example because an error message tells it nothing, or because a renamed argument silently breaks an agent that worked yesterday.

This article lays out a layered test strategy for MCP servers, with runnable examples using the MCP Python SDK v2, where the high-level server class is MCPServer and a single Client class can connect to a server object in memory, to a subprocess or to a URL. If you are new to how tools are declared, start with MCP tools.

Advertisement

A test pyramid for MCP servers

Five layers of MCP server tests: fast and many at the bottom, slow and few at the top5. Model evalsdoes a model pick it right?4. Transport + deploymentHTTP, auth, Inspector CLI in CI3. Contract snapshotstools/list: names, schemas, descriptions2. In-memory protocol testsClient(server): results, errors, structured content1. Unit tests of tool logicplain functions, fakes for backends, no MCP at allLayers 1-3 run on every commit in seconds. Layer 4 runs per build against a real process.Layer 5 runs nightly or before releases, because it is slow, costs tokens and is statistical.
Five test layers. Most tests should sit at the bottom, where they are fast and deterministic.

Each layer catches a different class of bug. Unit tests catch wrong business logic. In-memory protocol tests catch wrong results, error mapping and serialization. Contract snapshots catch accidental breaking changes to what clients and models see. Transport tests catch deployment problems: routing, authentication, session handling and proxies. Model evals catch the problem no other layer can see, which is whether a real model understands your tools well enough to use them.

The common mistake is to test only at the top, by chatting with the server in a desktop client. That is slow, cannot run in CI, and fails for reasons unrelated to your code. The second most common mistake is to test only at the bottom and never check the protocol surface, which is where most MCP-specific bugs live.

Layer 1: keep tool logic testable without MCP

Write each tool as a thin adapter around a plain function or service. The adapter handles MCP concerns: argument types, turning domain exceptions into tool errors, and shaping results. The plain function holds the logic and has ordinary unit tests with fakes for databases and APIs. This split keeps the fast tests fast, and it means a change of protocol version or SDK major version touches the adapter, not the logic.

The example server below follows that shape. lookup_order lives in a separate module. The tool function maps a missing order to a ToolError with a message written for the model, and returns a Pydantic model so the SDK can publish an output schema and return structured content alongside the text.

# orders_server.py
from mcp.server import MCPServer
from mcp.server.mcpserver.exceptions import ToolError
from pydantic import BaseModel

from orders_core import OrderNotFound, lookup_order   # plain business logic, tested on its own

mcp = MCPServer("orders")

class OrderStatus(BaseModel):
    order_id: int
    status: str
    total: float

@mcp.tool()
async def get_order_status(order_id: int) -> OrderStatus:
    """Return the status and total of one order by its numeric id."""
    try:
        o = await lookup_order(order_id)
    except OrderNotFound:
        raise ToolError(f"No order {order_id}. Order ids are numeric, e.g. 1182.")
    return OrderStatus(order_id=o.id, status=o.status, total=o.total)
Advertisement

Layer 2: in-memory protocol tests

The Python SDK's Client accepts the server object itself and talks to it in process. The request still goes through the SDK's protocol handling, including initialization, argument validation against the input schema, the tool wrapper and result serialization, but there is no subprocess or socket. Tests run in milliseconds and are deterministic. The SDK documentation notes that the in-memory client is era-neutral: it probes the server and picks the appropriate protocol version, currently 2026-07-28 or the older 2025-11-25.

Pass raise_exceptions=True in most tests. Without it, failures outside the tool body are sanitized the way they are in production, which hides the real traceback while you debug. That brings us to the test most suites forget. Because the flag changes error behaviour, at least one test must run without it and prove that production error handling leaks nothing.

# test_orders_server.py
import pytest
from mcp import Client
import orders_server
from orders_server import mcp

@pytest.fixture
def anyio_backend():
    return "asyncio"

@pytest.fixture
def fake_db(monkeypatch):
    orders = {1182: type("O", (), {"id": 1182, "status": "REFUNDED", "total": 49.0})()}
    async def lookup(order_id):
        if order_id not in orders:
            raise orders_server.OrderNotFound(order_id)
        return orders[order_id]
    monkeypatch.setattr(orders_server, "lookup_order", lookup)

@pytest.fixture
async def client(fake_db):
    async with Client(mcp, raise_exceptions=True) as c:
        yield c

@pytest.mark.anyio
async def test_happy_path_returns_structured_content(client):
    result = await client.call_tool("get_order_status", {"order_id": 1182})
    assert not result.is_error
    assert result.structured_content == {"order_id": 1182, "status": "REFUNDED", "total": 49.0}

@pytest.mark.anyio
async def test_unknown_order_is_a_tool_error_the_model_can_read(client):
    result = await client.call_tool("get_order_status", {"order_id": 9})
    assert result.is_error
    assert "No order 9" in result.content[0].text

@pytest.mark.anyio
async def test_crash_is_sanitized(monkeypatch, fake_db):
    async def boom(order_id):
        raise RuntimeError("password=hunter2 host=db-7.internal")
    monkeypatch.setattr(orders_server, "lookup_order", boom)
    async with Client(mcp) as c:                         # production behaviour: no raise_exceptions
        result = await c.call_tool("get_order_status", {"order_id": 1182})
    assert result.is_error
    text = " ".join(getattr(b, "text", "") for b in result.content)
    assert "hunter2" not in text and "db-7" not in text

The three tests cover the three outcomes a tool call can have. A success returns is_error false and structured_content that matches the output model. Python attributes are snake_case in SDK v2, while the wire format stays camelCase. A ToolError returns a normal result with is_error true and the message in content, which the model reads and can recover from. Any other exception is treated as a crash, and the model sees only a generic "Error executing tool" message while the traceback goes to the server log. The sanitization test plants a fake secret in the exception and asserts that it never reaches the result.

Testing protocol errors and bad input

There is a fourth outcome. When the request itself should be rejected, for example because the client lacks a capability the tool needs, the SDK's MCPError fails the whole tools/call with a JSON-RPC error. There is no result for the model, and the host application handles it. The SDK's guidance is a useful rule for choosing between them: if a smarter model could have avoided the failure by passing different arguments, raise ToolError; otherwise raise MCPError. Test each deliberate MCPError by asserting that the client call raises and that the error code is the one you chose, such as INVALID_PARAMS. Check the SDK API reference for the exact exception the client raises in your version.

Then test hostile and sloppy input, because models produce both. Send a string where an integer is expected, omit required arguments, add unknown ones, pass empty strings, huge strings and negative ids, and call a tool that does not exist. For each case, assert the outcome class and that the message would help a model fix its call. pytest.mark.parametrize over a table of bad inputs keeps this compact. For the error semantics in the protocol itself, see MCP error handling.

Layer 3: snapshot the contract

Your tools/list response is an API. Clients cache it, agents are prompted around it, and models learn its wording. Renaming an argument, tightening a type or rewording a description can break a working agent without breaking any unit test. A contract snapshot turns those changes into an explicit review step.

import json, pathlib

SNAPSHOT = pathlib.Path(__file__).with_name("tools_contract.json")

@pytest.mark.anyio
async def test_tool_contract_unchanged(client):
    tools = (await client.list_tools()).tools
    current = {t.name: {"description": t.description, "input_schema": t.input_schema,
                        "output_schema": t.output_schema} for t in tools}
    if not SNAPSHOT.exists():                            # first run writes the baseline
        SNAPSHOT.write_text(json.dumps(current, indent=2, sort_keys=True))
    assert current == json.loads(SNAPSHOT.read_text()), "tool contract changed: review and update"

The first run writes a baseline file that you commit. Every later run compares against it, and a diff in a pull request shows exactly what changed. Treat changes the way you would treat a REST API. Adding an optional argument or a new tool is safe. Removing or renaming a tool or argument, making an optional argument required, or changing the output schema of structured results is breaking, and deserves a new tool name or a deprecation period. The structured content article explains why output schemas matter to clients that validate results. Include description text in the snapshot too: description edits change model behaviour, so they should be reviewed as deliberately as schema edits.

Layer 4: transport, auth and deployment

In-memory tests bypass everything between the client and your handler, which is also where deployment bugs live: the HTTP route, session handling, authentication middleware, reverse proxies that buffer streaming responses, and CORS or origin checks. Run a small suite against the real server process on its real transport, for every build.

The same Client class connects to a URL, so the layer 2 tests can be reused with a different fixture. The MCP Inspector also has a CLI mode that makes a good smoke test in any CI system, because it needs no test code:

# Build step: start the real server on its HTTP transport, then probe it like a client would
python -m orders_server --http --port 8765 &          # your own entry point
sleep 2
npx @modelcontextprotocol/inspector --cli http://localhost:8765/mcp \
    --transport http --method tools/list
npx @modelcontextprotocol/inspector --cli http://localhost:8765/mcp \
    --transport http --method tools/call \
    --tool-name get_order_status --tool-arg order_id=1182

Pass --transport http explicitly for Streamable HTTP servers rather than relying on a default. The Inspector CLI also accepts --config and --server to reuse a server definition from a config file. For authenticated servers, add negative tests: no token, an expired token, and a valid token for the wrong tenant should each be rejected before any tool runs. The MCP security model describes what should be enforced at this layer.

Layer 5: does a model use the tools correctly?

The final question cannot be answered by asserting on JSON: given a realistic request, does a model pick the right tool, pass sensible arguments, and recover from errors? Answer it with a small eval set. Write twenty to fifty prompts that represent real requests, including ambiguous ones and ones that should call no tool at all. Run each through a model connected to your server, record the trajectory of tool calls, and grade it: the right tool, valid arguments, no forbidden calls, and a correct final answer where you can check it.

Treat the results as statistics, not pass or fail. Run each prompt several times, report the success rate per prompt, and compare against the last release. Keep the backend faked, so that evals measure tool usability and not the state of a staging database. This is where description changes prove their worth: a rewording that raises the right-tool rate from 80 to 95 percent is a real improvement, and one that lowers it is a regression the contract snapshot could only flag, not judge.

Worked example: a regression caught at each layer

Suppose a developer changes get_order_status to accept order_ref as a string instead of a numeric order_id, and wraps the database call in a broad except Exception that returns the error text as the result. Here is what each layer catches. The unit tests still pass, because the logic is fine. The layer 2 happy-path test fails because the argument name changed. The not-found test fails because the error now comes back as a normal result with is_error false, which a model would read as an answer. The sanitization test fails because the raw exception text, with a hostname in it, now reaches the result. The contract snapshot shows a breaking rename. Nothing reaches layer 4. If the rename was intended, the snapshot diff forces a decision to ship it as a new tool, and the layer 5 evals confirm that models still find it.

Failure modes of MCP test suites

Anti-patternConsequenceFix
Only manual testing in a chat clientSlow, flaky, nothing runs in CILayers 1-3 on every commit
All tests use raise_exceptions=TrueLeaky production errors go unnoticedAt least one sanitization test without it
Asserting on full text outputTests break on harmless wording changesAssert on is_error and structured_content
No contract snapshotSilent breaking renamesSnapshot tools/list, including descriptions
Evals against live backendsScores reflect data, not tool designFake backends; fixed fixtures
Single run per eval promptNoise mistaken for regressionsRepeat runs; compare rates

What to do next

  1. Split each tool into a thin MCP adapter and a plain function, and unit-test the function.
  2. Add an in-memory Client(mcp) fixture and one test per outcome: success, tool error, crash.
  3. Write one test without raise_exceptions that plants a secret and proves it never reaches the result.
  4. Parametrize a table of bad inputs and assert each gives a helpful, correctly flagged error.
  5. Commit a tools/list snapshot and require review of every diff to it.
  6. Add an Inspector CLI smoke test with --transport http against the built server, plus auth negative tests.
  7. Build a small model-in-the-loop eval set, run it before releases, and track the right-tool rate.
Key takeaway: Test an MCP server for both of its users. For client software, test in layers: plain unit tests for tool logic, fast in-memory protocol tests through the SDK's Client for successes, tool errors, protocol errors and sanitized crashes, a committed snapshot of the tools/list contract, and a smoke suite over the real HTTP transport with authentication. For the model, keep a small eval set that measures whether it picks the right tool with the right arguments, and treat the scores as rates. Most of the value comes from the cheap layers, so put them in CI first.