Most first MCP servers are written the same way: take an existing REST client, wrap every method in a tool, ship it. It works in a demo, and then the model starts calling the wrong endpoint, passing internal IDs it guessed, retrying a payment because a timeout looked like a failure, and burning half its context window on tool descriptions it never uses. None of that is a protocol problem. The protocol carried every message correctly. It is a design problem, and design problems have patterns.
This article is a catalogue of server-level design patterns, each with when to use it and when not to, measured against the MCP specification version 2025-11-25 as published on modelcontextprotocol.io and read on 2026-10-01. Tool naming, granularity, descriptions and the tool-count budget are covered in MCP tools; this page assumes them and works one level up: the shape of the whole server. Code uses plain JSON-RPC handling in Python's standard library, so it maps onto any SDK without relying on SDK-specific names.
The server is an interface for a probabilistic caller
A conventional API is designed for a programmer who reads documentation once, writes code, and gets the same behaviour on every run. An MCP server's main caller is a language model that rereads your descriptions on every turn, chooses among tools by meaning rather than by name, and recovers from errors only if the error tells it what to do. Three consequences drive every pattern below.
First, every tool, resource and prompt definition you expose costs context and attention on every turn, so the surface should be small and task-shaped. Second, the model will sometimes pick the wrong tool or pass a plausible but wrong argument, so the server, not the model, must enforce invariants: validation, authorization, idempotency and limits. The tools page says servers MUST validate all tool inputs, implement proper access controls, rate limit tool invocations and sanitize tool outputs. Third, results are read by the model as well as by code, so they need both a machine contract and a readable rendering.
Pattern 1: workflow facade, not API mirror
An API mirror exposes one tool per backend endpoint: get_ticket, list_comments, update_ticket_field, get_user_by_email and thirty more. A workflow facade exposes the jobs users actually ask for: find_tickets, triage_ticket, reply_to_ticket. Each facade tool may call several endpoints, resolve names to IDs internally, and apply business rules before writing anything.
Use a facade when the backend's natural unit of work is smaller than the user's, when correct use needs a specific call order, or when IDs and enums are opaque. The model never has to learn that closing a ticket requires setting a resolution code first; the facade does it. Use a mirror only when the server is a thin developer tool for a general API whose callers really do want every endpoint, and even then group by resource and paginate the list.
The cost of a facade is that you own more logic and you must version it. The benefit is fewer, more reliable calls and a surface you can test with scripted conversations rather than hoping the model composes primitives correctly.
Pattern 2: choose the primitive by who is in control
MCP gives a server three primitives, and the specification frames them by control. Tools are model-controlled: the model discovers and invokes them. Resources are application-driven: the host decides what context to attach. Prompts are user-controlled templates, typically surfaced as commands. Picking the wrong one is a common, quiet design error.
| Need | Primitive | Why |
|---|---|---|
| Take an action or run a query the model decides on | Tool | Only tools are invoked by the model on its own judgement |
| Expose a document, schema or file the user or host may attach | Resource | Readable by URI, subscribable, no model decision needed |
| Offer a reusable starting instruction, such as a triage checklist | Prompt | The user picks it explicitly; arguments fill the template |
| Return a large artefact from a tool | Tool returning a resource_link | The model gets a URI and summary instead of 200 KB inline |
The last row is worth adopting early. A tool MAY return links to resources, and a resource link returned by a tool is not guaranteed to appear in resources/list, so you can mint per-call artefacts, such as an export or a diff, without polluting the global listing.
Pattern 3: errors are results the model can act on
The 2025-11-25 tools page defines two error channels. Protocol errors are JSON-RPC errors for unknown tools, malformed requests and server errors. Tool execution errors are normal results with isError: true and cover API failures, input validation errors and business logic errors. Clients SHOULD pass execution errors to the model so it can self-correct, and MAY pass protocol errors, which are less likely to help.
The design rule follows: anything the model could fix by changing its arguments or its plan belongs in an isError result, written as an instruction. Say which argument was wrong, what the allowed values are, and what to do next. Reserve protocol errors for bugs and transport faults. A deeper treatment of codes and retry classes is in MCP error handling.
import json
TOOLS = {}
class ToolError(Exception):
'''An error the model can fix; becomes a result with isError: true.'''
def tool(name, description, input_schema, output_schema=None):
def register(fn):
spec = {"name": name, "description": description, "inputSchema": input_schema}
if output_schema is not None:
spec["outputSchema"] = output_schema
TOOLS[name] = {"fn": fn, "spec": spec}
return fn
return register
def call_tool(params, validate):
entry = TOOLS.get(params.get("name"))
if entry is None: # protocol error: JSON-RPC -32602
raise LookupError("Unknown tool: %s" % params.get("name"))
args = params.get("arguments") or {}
try:
problems = validate(args, entry["spec"]["inputSchema"]) # any JSON Schema 2020-12 validator
if problems:
raise ToolError("Invalid arguments: " + "; ".join(problems))
data = entry["fn"](**args)
except ToolError as exc:
return {"content": [{"type": "text", "text": str(exc)}], "isError": True}
result = {"content": [{"type": "text", "text": json.dumps(data)}]}
if "outputSchema" in entry["spec"]:
result["structuredContent"] = data # MUST conform to outputSchema
return resultHandlers raise ToolError for upstream failures too, after translating them: "The ticket system is rate limiting this account; retry after 30 seconds or narrow the query" is useful, a stack trace is not and may leak internals.
Pattern 4: a typed output contract with a readable twin
If a tool declares an outputSchema, the specification says the server MUST return structured results that conform to it, and clients SHOULD validate them. For backwards compatibility a tool returning structured content SHOULD also return the serialized JSON in a text block, which is why the code above sets both. Declare output schemas for tools whose results feed other code or later tool calls, such as IDs, counts and states, and keep them stable: removing or renaming a field is a breaking change for every client that validates. Structured content covers schema evolution in detail.
Pattern 5: stateless core, explicit handles
Session state is tempting: remember the last search, let next_page continue it, keep a draft in memory. It breaks the moment the server runs as more than one replica behind Streamable HTTP, restarts during a deploy, or serves a client that reconnects. Prefer a stateless core in which every tool call carries what it needs, and when state is unavoidable, make it an explicit handle the model passes back: a cursor, a draft ID or a job ID stored in a shared store with a TTL.
Handles have three properties worth enforcing. They are opaque, so the model cannot construct them. They are bound to the caller's authorization context, so one user's handle is useless to another. They expire, so abandoned drafts do not accumulate. Where a transport-level session exists, treat it as a routing and lifecycle hint rather than your data store; MCP sessions explains what the session ID does and does not guarantee.
Pattern 6: long-running work as tasks, with a fallback
A report that takes four minutes should not hold a tools/call open. The 2025-11-25 specification adds tasks, explicitly marked experimental. A server that declares the tasks.requests.tools.call capability can mark a tool with execution.taskSupport set to forbidden (the default), optional or required. The client adds a task field to its tools/call params, the server replies immediately with a task carrying a taskId and status: working, and the client polls tasks/get and fetches the final result with tasks/result, which returns exactly what the tool call would have returned. Statuses are working, input_required, completed, failed and cancelled.
Two requirements matter for design. When an authorization context exists, the server MUST bind tasks to it and reject access from other contexts; without one, task IDs MUST be unguessable. And the receiver may delete a task after its TTL, so results need to be fetched promptly. Because tasks are experimental and client support varies, keep a fallback: a start_report tool that returns a job handle and a get_report_status tool, backed by the same job store. Both paths should share one implementation so they cannot drift.
Pattern 7: progressive disclosure with list_changed
A server with sixty capabilities does not need to show sixty tools on turn one. Declare listChanged: true under the tools capability, start with a small core set, and expand when the conversation needs more: after the user connects a project, or when an enable_admin_tools tool is called and authorized. The server SHOULD then send notifications/tools/list_changed, and the client re-lists. The trade-off is predictability: some clients refresh lazily, and a tool that appears mid-conversation is one the model has not seen described before. Keep the core set sufficient on its own, and never hide a tool to enforce security; authorization belongs in the handler.
Pattern 8: compose behind a gateway, namespace the names
Organisations end up with many servers. Rather than asking every host to configure fifteen of them, put an aggregating gateway in front that handles authentication once, applies policy, and re-exports tools under namespaced names. The specification's recommended tool-name characters include the dot, so tickets.find and billing.find stay unique. The gateway owns cross-cutting concerns; the backend servers stay small and single-purpose. MCP gateway covers routing, policy and failure isolation. Avoid composition when two backends' tools overlap semantically; merging them just moves the model's confusion upstream.
Worked example: redesigning a ticketing server
Suppose the first version mirrored a ticket API with 34 tools and showed three typical failure patterns: the model called update_ticket_field with field names it guessed, it searched by email before looking up user IDs and got the order wrong, and it retried add_comment after timeouts, posting duplicate replies to customers.
The redesign exposes five tools. find_tickets takes free text plus optional status and assignee names and returns a cursor and an output schema with ticket IDs and summaries. get_ticket returns the ticket and a resource_link to the full thread. triage_ticket takes a ticket ID, priority and team name, resolves the team internally, and returns isError with the list of valid teams when the name does not match. reply_to_ticket requires a client-supplied idempotency key, so a retried call returns the first reply instead of posting a second. export_tickets is task-capable with a start/poll fallback. Prompts provide a weekly triage checklist; a resource exposes the team directory for hosts that want to attach it.
Expect a much smaller tool-definition footprint, no duplicate replies, and wrong-field errors that turn into self-correcting retries because the error text lists valid values. None of this needs protocol changes.
Failure modes
| Failure | Symptom | Defence |
|---|---|---|
| API mirror | Wrong call order, guessed IDs, huge tool list | Workflow facade; resolve names server-side |
| Errors as protocol errors | Model gives up or loops on opaque failures | Fixable errors as isError results with instructions |
| Hidden session state | Works on one replica, fails after deploy | Stateless core, explicit opaque handles in a shared store |
| Blocking long calls | Timeouts, duplicated work on retry | Tasks where supported, start/poll pair otherwise |
| Unscoped handles or task IDs | One user reads another's results | Bind to authorization context; unguessable IDs |
| Non-idempotent writes | Duplicate side effects after retries | Idempotency keys on every write tool |
| Visibility as security | Hidden tool still callable | Authorize in the handler, always |
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| Surface | Mirror: fast to build, flexible, error-prone for models | Facade: more server logic, fewer and safer calls |
| Long-running work | Tasks: protocol-native, experimental, client support varies | Start/poll tools: works everywhere, costs the model extra turns |
| Tool set | Static: predictable, larger context cost | Progressive: smaller context, depends on client refresh |
| Topology | Many small servers: isolation, more host configuration | Gateway: one policy point, one more hop and failure domain |
What to do next
- List the ten requests users actually make of your server and check that each maps to one or two tool calls; collapse mirrored endpoints into facade tools where it does not.
- Classify every capability as tool, resource or prompt by who controls it, and return large artefacts as resource links.
- Audit error paths: every fixable error becomes an
isErrorresult that names the argument and the valid values. - Add output schemas to tools whose results feed later calls, and return a text twin.
- Remove in-memory session state; replace it with opaque, scoped, expiring handles.
- Move any call that can exceed your client timeout to a task-capable tool with a start/poll fallback over one job store.
- Add idempotency keys to every write tool, and replay a recorded conversation against each release.