Most first MCP servers are written the same way: take an existing REST client, wrap every method in a tool, ship it. It works in a demo, and then the model starts calling the wrong endpoint, passing internal IDs it guessed, retrying a payment because a timeout looked like a failure, and burning half its context window on tool descriptions it never uses. None of that is a protocol problem. The protocol carried every message correctly. It is a design problem, and design problems have patterns.

This article is a catalogue of server-level design patterns, each with when to use it and when not to, measured against the MCP specification version 2025-11-25 as published on modelcontextprotocol.io and read on 2026-10-01. Tool naming, granularity, descriptions and the tool-count budget are covered in MCP tools; this page assumes them and works one level up: the shape of the whole server. Code uses plain JSON-RPC handling in Python's standard library, so it maps onto any SDK without relying on SDK-specific names.

Advertisement

The server is an interface for a probabilistic caller

A conventional API is designed for a programmer who reads documentation once, writes code, and gets the same behaviour on every run. An MCP server's main caller is a language model that rereads your descriptions on every turn, chooses among tools by meaning rather than by name, and recovers from errors only if the error tells it what to do. Three consequences drive every pattern below.

First, every tool, resource and prompt definition you expose costs context and attention on every turn, so the surface should be small and task-shaped. Second, the model will sometimes pick the wrong tool or pass a plausible but wrong argument, so the server, not the model, must enforce invariants: validation, authorization, idempotency and limits. The tools page says servers MUST validate all tool inputs, implement proper access controls, rate limit tool invocations and sanitize tool outputs. Third, results are read by the model as well as by code, so they need both a machine contract and a readable rendering.

An MCP server in layers: the model sees only the top surface; everything below is yours to shapeClient / host applicationtools/list, tools/call, resources/read, prompts/get, tasks/get ...JSON-RPC over stdio or Streamable HTTPProtocol layer: lifecycle, capability negotiation, dispatch, paginationToolsmodel-controlled actionsResourcesapplication-driven contextPromptsuser-invoked templatesWorkflow facade (domain layer)validation, idempotency keys, business rules, error translationTicket APISearch indexJob storelong-running workCross-cutting on every call: authorization, rate limits, audit log, output sanitisation (server MUSTs in the tools spec).
The layering this article recommends. The protocol layer is generic; the workflow facade is where design decisions live; backends never leak their shape directly to the model.

Pattern 1: workflow facade, not API mirror

An API mirror exposes one tool per backend endpoint: get_ticket, list_comments, update_ticket_field, get_user_by_email and thirty more. A workflow facade exposes the jobs users actually ask for: find_tickets, triage_ticket, reply_to_ticket. Each facade tool may call several endpoints, resolve names to IDs internally, and apply business rules before writing anything.

Use a facade when the backend's natural unit of work is smaller than the user's, when correct use needs a specific call order, or when IDs and enums are opaque. The model never has to learn that closing a ticket requires setting a resolution code first; the facade does it. Use a mirror only when the server is a thin developer tool for a general API whose callers really do want every endpoint, and even then group by resource and paginate the list.

The cost of a facade is that you own more logic and you must version it. The benefit is fewer, more reliable calls and a surface you can test with scripted conversations rather than hoping the model composes primitives correctly.

Advertisement

Pattern 2: choose the primitive by who is in control

MCP gives a server three primitives, and the specification frames them by control. Tools are model-controlled: the model discovers and invokes them. Resources are application-driven: the host decides what context to attach. Prompts are user-controlled templates, typically surfaced as commands. Picking the wrong one is a common, quiet design error.

NeedPrimitiveWhy
Take an action or run a query the model decides onToolOnly tools are invoked by the model on its own judgement
Expose a document, schema or file the user or host may attachResourceReadable by URI, subscribable, no model decision needed
Offer a reusable starting instruction, such as a triage checklistPromptThe user picks it explicitly; arguments fill the template
Return a large artefact from a toolTool returning a resource_linkThe model gets a URI and summary instead of 200 KB inline

The last row is worth adopting early. A tool MAY return links to resources, and a resource link returned by a tool is not guaranteed to appear in resources/list, so you can mint per-call artefacts, such as an export or a diff, without polluting the global listing.

Pattern 3: errors are results the model can act on

The 2025-11-25 tools page defines two error channels. Protocol errors are JSON-RPC errors for unknown tools, malformed requests and server errors. Tool execution errors are normal results with isError: true and cover API failures, input validation errors and business logic errors. Clients SHOULD pass execution errors to the model so it can self-correct, and MAY pass protocol errors, which are less likely to help.

The design rule follows: anything the model could fix by changing its arguments or its plan belongs in an isError result, written as an instruction. Say which argument was wrong, what the allowed values are, and what to do next. Reserve protocol errors for bugs and transport faults. A deeper treatment of codes and retry classes is in MCP error handling.

import json

TOOLS = {}


class ToolError(Exception):
    '''An error the model can fix; becomes a result with isError: true.'''


def tool(name, description, input_schema, output_schema=None):
    def register(fn):
        spec = {"name": name, "description": description, "inputSchema": input_schema}
        if output_schema is not None:
            spec["outputSchema"] = output_schema
        TOOLS[name] = {"fn": fn, "spec": spec}
        return fn
    return register


def call_tool(params, validate):
    entry = TOOLS.get(params.get("name"))
    if entry is None:                       # protocol error: JSON-RPC -32602
        raise LookupError("Unknown tool: %s" % params.get("name"))
    args = params.get("arguments") or {}
    try:
        problems = validate(args, entry["spec"]["inputSchema"])   # any JSON Schema 2020-12 validator
        if problems:
            raise ToolError("Invalid arguments: " + "; ".join(problems))
        data = entry["fn"](**args)
    except ToolError as exc:
        return {"content": [{"type": "text", "text": str(exc)}], "isError": True}
    result = {"content": [{"type": "text", "text": json.dumps(data)}]}
    if "outputSchema" in entry["spec"]:
        result["structuredContent"] = data  # MUST conform to outputSchema
    return result

Handlers raise ToolError for upstream failures too, after translating them: "The ticket system is rate limiting this account; retry after 30 seconds or narrow the query" is useful, a stack trace is not and may leak internals.

Pattern 4: a typed output contract with a readable twin

If a tool declares an outputSchema, the specification says the server MUST return structured results that conform to it, and clients SHOULD validate them. For backwards compatibility a tool returning structured content SHOULD also return the serialized JSON in a text block, which is why the code above sets both. Declare output schemas for tools whose results feed other code or later tool calls, such as IDs, counts and states, and keep them stable: removing or renaming a field is a breaking change for every client that validates. Structured content covers schema evolution in detail.

Pattern 5: stateless core, explicit handles

Session state is tempting: remember the last search, let next_page continue it, keep a draft in memory. It breaks the moment the server runs as more than one replica behind Streamable HTTP, restarts during a deploy, or serves a client that reconnects. Prefer a stateless core in which every tool call carries what it needs, and when state is unavoidable, make it an explicit handle the model passes back: a cursor, a draft ID or a job ID stored in a shared store with a TTL.

Handles have three properties worth enforcing. They are opaque, so the model cannot construct them. They are bound to the caller's authorization context, so one user's handle is useless to another. They expire, so abandoned drafts do not accumulate. Where a transport-level session exists, treat it as a routing and lifecycle hint rather than your data store; MCP sessions explains what the session ID does and does not guarantee.

Pattern 6: long-running work as tasks, with a fallback

A report that takes four minutes should not hold a tools/call open. The 2025-11-25 specification adds tasks, explicitly marked experimental. A server that declares the tasks.requests.tools.call capability can mark a tool with execution.taskSupport set to forbidden (the default), optional or required. The client adds a task field to its tools/call params, the server replies immediately with a task carrying a taskId and status: working, and the client polls tasks/get and fetches the final result with tasks/result, which returns exactly what the tool call would have returned. Statuses are working, input_required, completed, failed and cancelled.

Two requirements matter for design. When an authorization context exists, the server MUST bind tasks to it and reject access from other contexts; without one, task IDs MUST be unguessable. And the receiver may delete a task after its TTL, so results need to be fetched promptly. Because tasks are experimental and client support varies, keep a fallback: a start_report tool that returns a job handle and a get_report_status tool, backed by the same job store. Both paths should share one implementation so they cannot drift.

Pattern 7: progressive disclosure with list_changed

A server with sixty capabilities does not need to show sixty tools on turn one. Declare listChanged: true under the tools capability, start with a small core set, and expand when the conversation needs more: after the user connects a project, or when an enable_admin_tools tool is called and authorized. The server SHOULD then send notifications/tools/list_changed, and the client re-lists. The trade-off is predictability: some clients refresh lazily, and a tool that appears mid-conversation is one the model has not seen described before. Keep the core set sufficient on its own, and never hide a tool to enforce security; authorization belongs in the handler.

Pattern 8: compose behind a gateway, namespace the names

Organisations end up with many servers. Rather than asking every host to configure fifteen of them, put an aggregating gateway in front that handles authentication once, applies policy, and re-exports tools under namespaced names. The specification's recommended tool-name characters include the dot, so tickets.find and billing.find stay unique. The gateway owns cross-cutting concerns; the backend servers stay small and single-purpose. MCP gateway covers routing, policy and failure isolation. Avoid composition when two backends' tools overlap semantically; merging them just moves the model's confusion upstream.

Worked example: redesigning a ticketing server

Suppose the first version mirrored a ticket API with 34 tools and showed three typical failure patterns: the model called update_ticket_field with field names it guessed, it searched by email before looking up user IDs and got the order wrong, and it retried add_comment after timeouts, posting duplicate replies to customers.

The redesign exposes five tools. find_tickets takes free text plus optional status and assignee names and returns a cursor and an output schema with ticket IDs and summaries. get_ticket returns the ticket and a resource_link to the full thread. triage_ticket takes a ticket ID, priority and team name, resolves the team internally, and returns isError with the list of valid teams when the name does not match. reply_to_ticket requires a client-supplied idempotency key, so a retried call returns the first reply instead of posting a second. export_tickets is task-capable with a start/poll fallback. Prompts provide a weekly triage checklist; a resource exposes the team directory for hosts that want to attach it.

Expect a much smaller tool-definition footprint, no duplicate replies, and wrong-field errors that turn into self-correcting retries because the error text lists valid values. None of this needs protocol changes.

Failure modes

FailureSymptomDefence
API mirrorWrong call order, guessed IDs, huge tool listWorkflow facade; resolve names server-side
Errors as protocol errorsModel gives up or loops on opaque failuresFixable errors as isError results with instructions
Hidden session stateWorks on one replica, fails after deployStateless core, explicit opaque handles in a shared store
Blocking long callsTimeouts, duplicated work on retryTasks where supported, start/poll pair otherwise
Unscoped handles or task IDsOne user reads another's resultsBind to authorization context; unguessable IDs
Non-idempotent writesDuplicate side effects after retriesIdempotency keys on every write tool
Visibility as securityHidden tool still callableAuthorize in the handler, always

Trade-offs

DecisionOption AOption B
SurfaceMirror: fast to build, flexible, error-prone for modelsFacade: more server logic, fewer and safer calls
Long-running workTasks: protocol-native, experimental, client support variesStart/poll tools: works everywhere, costs the model extra turns
Tool setStatic: predictable, larger context costProgressive: smaller context, depends on client refresh
TopologyMany small servers: isolation, more host configurationGateway: one policy point, one more hop and failure domain

What to do next

  1. List the ten requests users actually make of your server and check that each maps to one or two tool calls; collapse mirrored endpoints into facade tools where it does not.
  2. Classify every capability as tool, resource or prompt by who controls it, and return large artefacts as resource links.
  3. Audit error paths: every fixable error becomes an isError result that names the argument and the valid values.
  4. Add output schemas to tools whose results feed later calls, and return a text twin.
  5. Remove in-memory session state; replace it with opaque, scoped, expiring handles.
  6. Move any call that can exceed your client timeout to a task-capable tool with a start/poll fallback over one job store.
  7. Add idempotency keys to every write tool, and replay a recorded conversation against each release.
Key takeaway: An MCP server is an interface for a caller that reasons by meaning and recovers only from errors it can read. Design it as a small workflow facade, pick tools, resources and prompts by who controls them, return fixable errors as results with instructions, give outputs a typed contract, keep the core stateless with explicit scoped handles, move long work to tasks with a fallback, and compose servers behind a gateway with namespaced names. The protocol carries the messages; these patterns decide whether the model uses them well.