Agents that operate websites today mostly do it the hard way. They read a screenshot or an accessibility tree, guess which element is the search box, type, click and look again. A redesigned button, a modal or a localized label can derail a run, and every step costs a model call. The site, meanwhile, already has clean functions for everything the agent is fumbling towards. WebMCP is a proposal to let the page hand those functions to the agent directly.

WebMCP is being incubated in the W3C Web Machine Learning Community Group. A page registers tools: JavaScript functions with a name, a natural-language description and a JSON Schema for their input. An agent in or attached to the browser discovers the tools and calls them, and the function runs in the page with the user's existing session. The spec describes such a page as a Model Context Protocol server implemented in client-side script instead of on a backend. This article follows the Community Group draft as of late September 2026; it has already changed shape once, so check the current spec before shipping.

Advertisement

Why tools beat pixels

The case for WebMCP rests on three properties that screen automation cannot offer. First, intent is explicit. A tool named add_to_cart with a schema that requires a product id and a quantity tells the agent exactly what the operation is and what it needs. Second, the call is cheap and deterministic: one structured call replaces a loop of perceive, act and verify. Third, the page stays in charge. The tool runs the site's own code, updates the site's own UI, and can refuse or ask for confirmation in the site's own way.

WebMCP gives the agent no new powers: a tool can only do what the page's JavaScript could already do with the user's session, and its requests are ordinary page requests that the backend still authorizes. Sites that register no tools still need DOM automation. For how small models handle the tool-calling side of this exchange, see Small Models for Tool Calling.

The API surface in the current draft

Where WebMCP sits: the page registers tools, the browser brokers callsPage scriptregisterTool(name, schema, execute)Browser (user agent)tool map per documentAgentbrowser agent or extensiontool call: JSON input -- execute() runs in the page -- JSON-serialized resultApp state + sessioncookies, auth, UIGatesPermissions-Policy tools, same originExtension bridgerelay to desktop MCP clientDesktop MCP clientspeaks MCP, not WebMCPYour backendstill authorizes every write; WebMCP is not an auth layerThe tool runs as page JavaScript with the user's session. Everything it can do, the page could already do.
A page registers tools with the browser; a browser agent, or an extension that relays to a desktop MCP client, calls them. The tool executes as page JavaScript under the user's session, and the backend still authorizes each request.

In the draft, every Document has a modelContext attribute, available only in secure contexts. Earlier Chrome previews exposed it as navigator.modelContext, and earlier revisions had methods such as provideContext() that are now gone. Feature-detect both locations and keep the integration in one small module.

  • registerTool(tool, options) returns a promise. tool carries name, optional title, description, optional inputSchema, execute and optional annotations. options may carry an AbortSignal and an exposedTo list of origins.
  • There is no unregister method: aborting the registration signal removes the tool.
  • execute(inputObject, {signal}) receives the parsed JSON input and a signal that aborts if the caller cancels. Whatever it resolves with is serialized to JSON.
  • getTools(options) and executeTool(tool, input, options) let a document enumerate and call tools registered in the tab's frame tree, same-origin by default and cross-origin only when both sides opt in with fromOrigins and exposedTo.
  • Events: toolchange when the set of tools changes, toolactivated when one of your tools is invoked, and toolcancel when a pending call is cancelled.
  • Annotations: readOnlyHint, untrustedContentHint, consequentialHint and debugging, all booleans defaulting to false.

Registration is refused for a tool name that is empty, longer than 128 characters or contains anything other than ASCII letters, digits, underscore, hyphen and period; an empty description; a name already registered in the same document; an input schema that cannot be serialized to JSON; a document that is not allowed the tools permissions-policy feature; and a page whose agent cluster is not origin-keyed. That last one bites sites that still set document.domain or send Origin-Agent-Cluster: ?0.

Advertisement

Registering tools: a worked example

Consider a shop that wants agents to search the catalog and add items to the cart. The example registers two tools under one AbortController, validates inputs inside execute, renders the effect in the page, and returns a small structured result.

// Feature-detect: the current draft puts the API on document; early Chrome
// previews exposed it as navigator.modelContext. Treat both as provisional.
const mc = document.modelContext ?? navigator.modelContext;

function registerShopTools(store) {
  if (!mc) return null;                       // no WebMCP: the page works as before
  const ac = new AbortController();           // aborting unregisters every tool

  mc.registerTool({
    name: "search_products",
    description: "Search the catalog. Returns items with id, name, " +
                 "price_cents and in_stock. Does not change the cart.",
    inputSchema: {
      type: "object",
      properties: {
        query: { type: "string", minLength: 1, maxLength: 200 },
        limit: { type: "integer", minimum: 1, maximum: 20, default: 10 }
      },
      required: ["query"],
      additionalProperties: false
    },
    annotations: { readOnlyHint: true, untrustedContentHint: true },
    async execute({ query, limit = 10 }, { signal }) {
      const items = await store.search(query, { limit, signal });
      return { items: items.map(i => ({ id: i.id, name: i.name,
               price_cents: i.priceCents, in_stock: i.stock > 0 })) };
    }
  }, { signal: ac.signal }).catch(reportRegistrationError);  // e.g. NotAllowedError

  mc.registerTool({
    name: "add_to_cart",
    description: "Add a product to the signed-in user's cart. " +
                 "Returns the cart line count and subtotal.",
    inputSchema: {
      type: "object",
      properties: {
        product_id: { type: "string", pattern: "^[A-Z0-9-]{4,32}$" },
        quantity:   { type: "integer", minimum: 1, maximum: 10 }
      },
      required: ["product_id", "quantity"],
      additionalProperties: false
    },
    async execute(input, { signal }) {
      const q = validateAddToCart(input);     // do not rely on the schema
      const cart = await store.addToCart(q.product_id, q.quantity, { signal });
      renderCart(cart);                       // show the user the change
      return { lines: cart.lines.length, subtotal_cents: cart.subtotalCents };
    }
  }, { signal: ac.signal }).catch(reportRegistrationError);

  return ac;
}

The search tool is marked readOnlyHint so an agent can call it without asking, and untrustedContentHint because product text can be seller-written and may try to instruct the model. The cart tool validates even though the schema constrains the input: the draft stores the schema as a string for the agent to read and does not say the browser enforces it. Results are plain objects; a DOM node, a BigInt or a cyclic structure cannot be serialized, and the call fails.

Designing tool definitions that agents use well

The description and schema are a prompt, and the usual prompt discipline applies. Say what the tool does, what it returns and what it does not do. Name units in field names (price_cents, not price). Prefer enums to free text wherever the value set is closed. Keep tools few and close to user intent: add_to_cart, not set_cart_state.

Return what the agent needs for its next decision and nothing else. A search result with ids, names, prices and availability lets the agent pick an item and call the next tool; a result with full HTML descriptions wastes context and widens the injection surface. When the agent can fix a failure, such as an out-of-stock item, return a structured error with a code instead of throwing. When semantics change, use a new name.

Lifecycle: tools follow the UI state

A tool should exist only when calling it makes sense. add_to_cart belongs on shop pages while a user is signed in; checkout belongs on the cart page when the cart is not empty. Because a name can be registered only once per document, updating a tool means aborting the old registration and registering again, and single-page apps must do this on route changes and session changes.

// Tools follow UI state. A name can be registered once per document, so a
// schema change is "abort the old registration, then register the new one".
let shopTools = null;

router.on("enter", route => {
  if (route.name === "shop" && session.signedIn) {
    shopTools ??= registerShopTools(store);
  } else {
    shopTools?.abort();                        // unregisters search_products and add_to_cart
    shopTools = null;
  }
});

mc?.addEventListener("toolactivated", e => {
  metrics.count("webmcp.tool_activated", { tool: e.toolName });
  ui.showAgentBanner(e.toolName);              // tell the user an agent is acting
});
mc?.addEventListener("toolcancel", e => {
  metrics.count("webmcp.tool_cancelled", { tool: e.toolName });
});

The toolactivated listener is the place for telemetry and for making agent activity visible, which is the collaborative mode WebMCP is designed for: user and agent in the same interface, with the user able to see and interrupt.

Declarative tools from forms

Many useful actions are already HTML forms. A companion explainer, not yet specified in the main draft, proposes that forms become tools through attributes: toolname and tooldescription on the form, toolparamdescription on each field, and an optional toolautosubmit that lets the agent submit without a user click. The browser would synthesize the input schema from the form controls.

<form toolname="search_flights"
      tooldescription="Search one-way flights between two airports on a date."
      action="/flights/search">
  <input name="origin" required pattern="[A-Z]{3}"
         toolparamdescription="IATA code of the departure airport, e.g. SFO">
  <input name="destination" required pattern="[A-Z]{3}"
         toolparamdescription="IATA code of the arrival airport, e.g. JFK">
  <input name="date" type="date" required
         toolparamdescription="Departure date">
  <button>Search</button>
</form>

The same explainer proposes SubmitEvent.agentInvoked so a submit handler can tell agent submissions apart, and SubmitEvent.respondWith() so script can return a structured result instead of navigating. All of this is explainer-stage; the imperative API is the one with normative text today.

The browser extension bridge

Native support serves agents built into the browser, but many agents are desktop applications that speak MCP. Community projects such as MCP-B and webmcp-bridge close that gap with a browser extension: a script in the page collects registered tools, either through the page's own API or through a polyfill that implements modelContext where the browser does not, and the extension relays them to a local MCP server process that desktop clients connect to. From the desktop client's point of view, each open tab becomes a set of MCP tools.

This architecture is useful, and it changes the security picture. The extension sees tools from every tab it is granted, which merges origins the browser would otherwise keep apart; the desktop agent can act on a banking tab while reading a forum tab. A bridge should namespace tools by origin, require the user to enable each site explicitly, show which tab a call targets, and pass the annotations through so the client can ask for confirmation on consequential calls.

Security: what the page must still own

The spec treats prompt injection as the central risk, from three directions. Tool poisoning is a malicious or compromised page writing descriptions that instruct the agent. Output injection is untrusted content, such as reviews, emails or listings, returned from a tool and read by the model as instructions. Attacks on the implementation are crafted arguments that exploit a tool's code. The spec also flags over-parameterization, where a tool asks for more personal data than it needs so a site can harvest what the agent knows about the user, and misrepresented intent, where an agent finalizes an action the user did not mean.

  • Send Permissions-Policy: tools=() on pages that should never expose tools, so injected scripts and compromised dependencies cannot register any. Cross-origin iframes get the feature only if the embedder delegates it.
  • Mark tools that return third-party text with untrustedContentHint, and keep those payloads short and structured.
  • Mark purchases, transfers, deletions and messages with consequentialHint, and still confirm them in your own UI; a hint is advice to the agent, not enforcement.
  • Keep server-side authorization, rate limits and CSRF protections unchanged. A tool call is just page JavaScript making requests.
  • Ask only for the inputs the operation needs. A schema field is a request for data the agent may hold about the user.
  • Leave exposedTo empty unless a specific partner origin needs your tools, and list it explicitly when one does.

Failure modes

  • Silent no-op in unsupported browsers. Without feature detection the integration throws on load; with detection it silently does nothing, so measure how often tools are actually registered.
  • Duplicate registration. Re-registering on every render rejects with InvalidStateError and leaves stale definitions in place.
  • Stale tools after navigation. Single-page apps that never abort keep offering checkout on pages where it is meaningless.
  • Invisible agent actions. State changes with no UI update leave the user unable to see or undo what happened.
  • Ignored cancellation. An execute that ignores its signal keeps working after the agent has moved on.

Testing and operating it

Because getTools and executeTool work same-origin, a page can exercise its own tools the way an agent would, in end-to-end tests with WebMCP enabled.

// Same-origin self-test: enumerate your own tools and call them the way an
// agent would. executeTool resolves with the JSON-serialized result string.
async function webmcpSelfTest() {
  const tools = await mc.getTools();
  for (const t of tools) {
    console.assert(t.description.length > 20, `${t.name}: description too thin`);
    console.assert(t.inputSchema, `${t.name}: no inputSchema`);
  }
  const search = tools.find(t => t.name === "search_products");
  const out = JSON.parse(await mc.executeTool(search, { query: "usb-c cable" }));
  console.assert(Array.isArray(out.items), "search_products: bad result shape");
}

In production, count registrations, activations, cancellations and failures per tool, and log argument shapes without values. A tool that is registered often but never activated has a description agents do not understand; a tool with a high failure rate usually has a schema looser than its validator.

Trade-offs against the alternatives

A backend MCP server reaches agents that never open a browser, but needs its own auth flow and cannot share the user's live UI state. WebMCP reuses the logged-in session and keeps the user in the loop, at the cost of requiring an open tab, a supporting browser or a bridge, and a draft API that is still moving. DOM automation needs no cooperation from the site and works everywhere, but it is slow, brittle and expensive per step. Many sites will end up offering both a backend MCP server for headless use and WebMCP tools for in-browser collaboration, sharing validation and business logic between them.

Key takeaway: WebMCP turns a page's own JavaScript into agent tools: register a small set of intent-level tools with precise descriptions and schemas on document.modelContext, tie their lifetime to an AbortController that follows UI state, validate inside execute and return small JSON results, annotate read-only, untrusted and consequential behavior, keep backend authorization unchanged, disable the feature where it is not wanted, and treat both the draft API and extension bridges as moving targets to re-check before each release.