A computer worm spreads by itself: it runs on one machine, copies itself to others, and needs no user to click anything. An AI worm does the same thing inside GenAI applications, except that its code is text and its processor is a language model. The attacker plants a prompt that, when a model reads it, makes the model do something harmful and also reproduce the prompt in its own output. If that output flows into another model-driven system, as email replies, chat messages, documents or agent-to-agent calls do, the cycle repeats without the attacker doing anything more.

This is not a new vulnerability class so much as prompt injection with a feedback loop. Indirect prompt injection covers a single hop: untrusted data steering one model. This article covers what changes when the output of one hop becomes the input of the next: the research that demonstrated it, a simple epidemic model that tells you which controls matter, the places in a modern agent stack where loops form, and the engineering that keeps an outbreak from spreading. No working payload appears here; the anatomy section uses placeholders.

Advertisement

Three ingredients: replication, propagation, payload

A prompt becomes a worm only when three things hold at once. Replication: the model, having read the prompt, emits text containing the prompt or an equivalent of it. Propagation: that output is delivered somewhere another model will later read it, without a human deciding to forward it. Payload: while replicating, the prompt also makes the model do something the attacker wants, such as including private data from its context in the output, sending spam, or steering users to a link.

Each ingredient maps to a property of real applications. Replication exploits the fact that instruction-tuned models follow instructions wherever they appear in the context and cannot reliably tell the developer's instructions from text inside an email. Propagation exploits automation: assistants that draft and send replies, summarise and post, or write into shared stores. The payload exploits whatever the model can see or do: retrieved documents, the user's contacts, tools. Remove any one of the three and there is no worm, only a one-hop injection. That observation is the basis for every defense below.

What Morris II demonstrated

The best-known demonstration is Morris II, named after the 1988 Morris worm, by Stav Cohen, Ron Bitton and Ben Nassi (arXiv 2403.02817, "Here Comes The AI Worm: Unleashing Zero-click Worms that Target GenAI-Powered Applications"). The researchers built an ecosystem of GenAI-powered email assistants that use retrieval-augmented generation: incoming mail is stored in a retrieval database, and when the assistant drafts a reply it retrieves related messages into the model's context. They crafted adversarial self-replicating prompts and showed two propagation classes. In RAG-based propagation, the malicious email is stored, later retrieved as context for a reply, and the model's reply both carries the prompt onward and leaks confidential data from other retrieved emails. In application-flow steering, the prompt's output steers what the application does next so that the worm moves along the application's own flow.

The paper evaluated several commercial models and embedding models; the versions of the paper differ in which, so check the current version rather than relying on a summary. It also proposed a guardrail, Virtual Donkey, aimed at detecting propagation, and reported a true-positive rate of 1.0 with a false-positive rate of 0.015 in its evaluation. Two lessons generalise beyond email. First, retrieval is an amplifier: a message that sits harmlessly in a store becomes active every time it is retrieved. Second, zero-click matters: the victim never opened anything; their assistant did.

How a self-replicating prompt moves between GenAI assistantsattacker messagepayload + replication textzero-clickvictim A inboxstored as ordinary dataindexedA's RAG storepersistenceretrieved laterA's assistant builds a replypayload now inside the model contextmodel outputcopy of payload + leaked dataauto-sendcontact Bnew RAG storecontact Cnew RAG storecontact Dnew RAG storeR0 = m x p_retrieve x p_replicate x p_delivercut mrecipient and rate capscut p_retrieveprovenance, quarantinecut p_replicateinput-output overlap checkcut p_deliverapproval before sendKeep the product of the four factors below 1 and an outbreak dies out on its own.
RAG-based propagation: the payload is stored as data, retrieved into a later context, copied into the output, and delivered to new stores. Each factor in R0 is a place to intervene.
Advertisement

An epidemic model you can reason with

Epidemiology gives a useful frame. Let R0 be the expected number of new systems one infected system infects. If R0 is below 1, an outbreak shrinks geometrically and dies out; if above 1, it grows. For an assistant that sends messages, R0 is roughly the product of four factors: m, the number of recipients an infected assistant's outputs reach while the payload stays retrievable; p_retrieve, the chance the malicious item is retrieved into a context that produces an outbound message; p_replicate, the chance the model copies the prompt faithfully enough to work again; and p_deliver, the chance that output leaves without a human stopping it.

Consider an illustrative configuration, not a measurement. An auto-replying assistant reaches m = 25 recipients a week, the poisoned email is retrieved for 30 percent of replies (p_retrieve = 0.3), the model replicates the prompt half the time (p_replicate = 0.5), and replies are sent automatically (p_deliver = 1). R0 = 25 x 0.3 x 0.5 x 1 = 3.75: an outbreak. Now vary one factor at a time.

ControlFactor changedNew R0
Human approves every outbound message, catching 90 percentp_deliver = 0.10.375
Cap automatic sends to 5 recipients per weekm = 50.75
Replication detector blocks 90 percent of copiesp_replicate = 0.050.375
Quarantine external mail from the retrieval indexp_retrieve = 0.030.375
Detector plus recipient capp_replicate = 0.05, m = 50.075

The numbers are invented, but the structure is not. Because the factors multiply, no single control has to be perfect; two imperfect ones that each cut a factor by five to ten push R0 far below 1. It also shows why fully autonomous sending with broad reach is the dangerous configuration: it pins p_deliver at 1 and inflates m.

Where loops form in agent stacks

Email is the demonstrated case, but any path where model output becomes another model's input without a human in between is a candidate. Look for these in your architecture:

  • Messaging assistants that draft, reply or forward on the user's behalf, in email, chat or support tickets.
  • Shared retrieval indexes where one user's or agent's output is indexed and retrieved for others, such as team knowledge bases fed by generated summaries.
  • Long-term agent memory, which turns a one-time injection into a persistent one that is replayed across sessions.
  • Multi-agent systems where agents pass free text to each other; one compromised agent's messages are every downstream agent's input.
  • Coding agents that write comments, documentation or issues that other agents later read as context.
  • Content pipelines that generate pages or posts later crawled into other systems' retrieval corpora.

For each path, write down the four factors. Where m is large, delivery is automatic and the input includes external text, you have a place a worm can live. RAG defenses covers hardening the retrieval side specifically.

Anatomy of a self-replicating prompt

Understanding the structure helps you detect it without needing a working sample. Published examples share four parts, shown here as non-functional placeholders:

[BAIT]        text chosen so the message is retrieved for many queries
              (broad, topical wording that embeds near common requests)
[REPLICATE]   an instruction to reproduce a marked block verbatim in any output
[PAYLOAD]     an instruction to include data from the context, or to take an action
[CONCEAL]     wording or formatting meant to keep the block unnoticed by the reader

Each part leaves a trace. Bait makes one item retrieved unusually often across unrelated queries. Replication makes outputs contain long spans copied from untrusted inputs, which normal replies rarely do outside quoted history. Payloads make outputs contain data from sources unrelated to the conversation, or tool calls the user did not ask for. Concealment shows up as hidden text, odd encodings or instruction-like language in fields that should hold content. Variants can ask the model to paraphrase instead of copying, and can use images for multimodal models, so exact-match detection is a floor, not a ceiling.

Defenses, factor by factor

Cut m with rate and fan-out limits on anything a model can send: recipients per message, messages per hour, new recipients per day, with anomalies alerting a human. Egress filtering covers enforcing that at the network and API layer. Cut p_deliver by requiring human approval for outbound messages, or at least for messages to new recipients and messages whose content was influenced by external text. Cut p_retrieve by tracking provenance: tag every stored item with its origin, keep external content out of indexes that feed outbound actions, or retrieve it into a clearly delimited, lower-trust section of the context.

Cut p_replicate by checking outputs against inputs. A reply that reproduces long spans of an untrusted retrieved item, outside a normal quoted-history block, is suspicious. The check below is deliberately simple and cheap enough to run on every outbound message:

import re

def shingles(text: str, n: int = 8) -> set[tuple[str, ...]]:
    words = re.findall(r"\w+", text.lower())
    return {tuple(words[i:i + n]) for i in range(len(words) - n + 1)}

def strip_quoted(body: str) -> str:
    """Drop quoted history so legitimate replies are not flagged for quoting."""
    return "\n".join(l for l in body.splitlines() if not l.lstrip().startswith(">"))

def replication_score(draft: str, untrusted_inputs: list[str]) -> float:
    """Fraction of the draft's 8-word shingles found in any untrusted input."""
    out = shingles(strip_quoted(draft))
    if not out:
        return 0.0
    seen = set().union(*(shingles(t) for t in untrusted_inputs)) if untrusted_inputs else set()
    return len(out & seen) / len(out)

def gate_outbound(draft, untrusted_inputs, recipients, sender_stats, policy):
    reasons = []
    if replication_score(draft, untrusted_inputs) > policy.max_copy_ratio:     # e.g. 0.2
        reasons.append("copies untrusted input")
    if len(recipients) > policy.max_recipients:
        reasons.append("fan-out")
    if sender_stats.new_recipients_today + len(recipients) > policy.max_new_per_day:
        reasons.append("rate")
    return ("hold_for_human", reasons) if reasons else ("send", [])

Shingle overlap catches verbatim replication and misses paraphrase, so pair it with a semantic similarity check and a classifier trained on injection-like text. More structural options remove the model's ability to act on untrusted text at all: the dual-LLM pattern keeps a quarantined model that reads untrusted content but has no tools, and designs such as CaMeL from Google DeepMind derive allowed data flows from the trusted user request alone. Output handling and agent permissions cover the surrounding controls.

Detection and incident response

Worms have an operational signature that single injections lack: the same content appears across many users in a short time. Fingerprint outbound messages and stored items with locality-sensitive hashes and alert when one fingerprint crosses tenant or user boundaries at an unusual rate. Watch per-sender fan-out, the ratio of automated to human-approved sends, and items retrieved far more often than their peers. Log, for each outbound message, which retrieved items were in the context, so you can trace an infection back to its first stored copy.

When an outbreak is suspected, the order matters. First stop propagation: switch automated sending to approval-required globally, which a well-designed system exposes as a single flag. Then find every stored copy by fingerprint and quarantine it from retrieval and memory, since leaving copies in stores lets the worm restart when automation resumes. Then assess the payload: what data could each infected context see, and where was it sent. Only then restore automation, ideally with the missing control added. Rehearse this with a harmless canary prompt that asks the model to include a marker string, in a sandboxed copy of your ecosystem, and measure R0 directly.

Trade-offs

Every control costs something. Approval gates cost user time and remove much of an assistant's value if applied to everything; scope them to external-influenced, new-recipient or high-fan-out messages. Overlap detection produces false positives on legitimate summaries and quotes, so tune thresholds on real traffic and route hits to review rather than silently dropping. Provenance tagging requires plumbing through ingestion, retrieval and prompting. The cheapest control is often a fan-out cap, because legitimate users rarely need an assistant to send hundreds of messages autonomously.

What to do next

  1. Inventory every path where model output can reach another model's input without a human, and estimate the four R0 factors for each.
  2. Put fan-out and rate caps on every model-initiated send, with alerts on anomalies.
  3. Require approval for outbound messages influenced by external content, at least to new recipients.
  4. Tag stored and retrieved items with provenance and keep external content out of indexes that drive actions.
  5. Run a replication check on outbound drafts against untrusted inputs, excluding quoted history.
  6. Log which retrieved items shaped each output, and fingerprint content to spot cross-user spread.
  7. Build a global switch that turns automation off, and a job that quarantines items by fingerprint.
  8. Rehearse with a harmless canary prompt in a sandbox and measure R0 before and after each control.
Key takeaway: An AI worm is prompt injection with a feedback loop: a prompt that makes a model copy it into outputs which reach other models, carrying a payload each time. Morris II showed this against RAG-based GenAI email assistants, with retrieval acting as the amplifier and no user interaction needed. Model the risk as R0, the product of reach, retrieval, replication and unreviewed delivery, and keep it below 1 with layered controls: fan-out caps, provenance and quarantine for external content, output-versus-input replication checks, and approval before sends. Then instrument for cross-user spread and rehearse the response.