A web application firewall sits in front of an application, inspects each HTTP request and decides whether it passes. Security teams already run one, so when an LLM feature ships, the natural question is whether the WAF can stop prompt injection too. Vendors now say yes. Cloudflare's AI Security for Apps (announced as Firewall for AI) exposes an injection score to its rule engine, and Microsoft's Prompt Shields classifies user prompts and documents for attacks.

The honest answer is that a WAF helps, but it is a different kind of control from the one that stops SQL injection. This article explains how a prompt injection WAF works inside: how it extracts prompts from API bodies, normalises them, turns scores into rule actions, and tracks conversations. It also covers what the WAF structurally cannot see, how to roll one out without breaking legitimate users, and where it fits next to the controls that actually carry the load. Detector internals and threshold maths are covered in prompt injection scanners; this article is about the gateway layer around them.

Why injection breaks the classic WAF model

A classic WAF works because classic injection has syntax. A SQL injection payload has to close a quote and add a clause that the database parser will accept, so signatures and parser-based detectors can match on structure. The OWASP Core Rule Set adds up the scores of many such matches and blocks above a threshold. It offers paranoia levels that trade false positives for coverage.

Prompt injection has no syntax. The target is a language model that will follow instructions in any language, encoding or tone. "Ignore the above" can be written in thousands of ways, split across turns, translated, or put in a picture. So every prompt injection WAF ends up running a classifier, a small model that outputs the probability that a text is trying to override instructions, with signatures as a cheap first pass. Two consequences follow. Detection is probabilistic, so you tune for a false-positive rate instead of writing a correct rule. And an attacker who can query the WAF can search for inputs that score as benign.

Where the WAF sits

There are three placements, each with a different view of the traffic.

Where a prompt injection WAF sits, and what it can seeClientchat UI, SDKEdge / gateway WAFextract, normalise, scoreLLM applicationprompt assemblyModelself-hosted or APIallow / tagRetrieval and toolsdocs, web, emailIn-app scannerdocuments, tool outputThe edge sees only what the client sent.Retrieved documents and tool results never cross it.Decision logrule id, score, actionRate limiterper key, per session
An edge or gateway WAF inspects only the client's request. Instructions hidden in retrieved documents or tool output have to be scanned inside the application.

  • Edge (CDN or cloud WAF). It blocks before traffic reaches your servers and is run by the team that already runs the WAF. It sees only the raw HTTP request.
  • AI gateway or reverse proxy. It understands LLM API formats, can read the whole messages array, and can attach a score to the request for downstream code to use.
  • In-application middleware. It sees the assembled prompt, including retrieved context and tool results. That is the only place indirect injection can be inspected.

Most mature deployments use the edge or gateway for cheap, high-volume filtering and rate limits, and an in-app scanner for the content the edge never sees.

Finding the prompt in the request

A WAF cannot score what it cannot find. LLM requests are JSON, and the text to score is nested: a messages array with roles, content that can be a string or a list of parts, and tool definitions and tool results the client may supply. Cloudflare's documentation says its detection currently handles only requests with a JSON content type, and the product announcement described finding prompts inside JSON bodies. Check how your vendor handles your exact request shape. If you build your own, the extractor is the first thing to get right.

def extract_segments(body: dict) -> list[tuple[str, str]]:
    """Return (role, text) pairs to score from an OpenAI-style chat request."""
    out = []
    for i, msg in enumerate(body.get("messages", [])):
        role = msg.get("role", "unknown")
        content = msg.get("content")
        if isinstance(content, str):
            out.append((role, content))
        elif isinstance(content, list):
            for part in content:
                if part.get("type") == "text":
                    out.append((role, part.get("text", "")))
                elif part.get("type") in ("image_url", "input_audio", "file"):
                    out.append((role, "<non-text part: route to in-app scanner>"))
    # a client-supplied system message is a red flag on most public endpoints
    return out

Three rules make the extractor safe. Fail closed on bodies it cannot parse. Otherwise an attacker sends malformed JSON that the WAF skips but the lenient app parser accepts. Enforce a size limit, because a classifier with a 512-token window that sees only the first window lets a payload hide at the end; score overlapping windows instead. And treat roles the client should not control, such as a system message on a public endpoint, as an anomaly in their own right.

Normalising before scoring

Normalise before scoring, and score the normalised text. Common evasions are invisible characters and look-alike letters. Apply Unicode NFKC, strip zero-width characters and the Unicode tag block, which can carry invisible ASCII, and decode obvious base64 or percent-encoded runs once.

import base64, re, unicodedata

INVISIBLE = re.compile(r"[​-‏⁠-⁤\U000E0000-\U000E007F]")
B64_RUN = re.compile(r"[A-Za-z0-9+/]{40,}={0,2}")

def normalise(text: str) -> tuple[str, list[str]]:
    flags = []
    if INVISIBLE.search(text):
        flags.append("invisible_chars")
        text = INVISIBLE.sub("", text)
    text = unicodedata.normalize("NFKC", text)
    for run in B64_RUN.findall(text)[:3]:
        try:
            decoded = base64.b64decode(run, validate=True).decode("utf-8")
            flags.append("base64_text")
            text += "\n" + decoded           # score the decoded form too
        except Exception:
            pass
    return text, flags

The flags are signals in their own right. Normal users rarely send tag-block characters.

From score to action

The detector produces a score. The rule engine turns that score, plus context, into an action. Keep the two separate, so security can change policy without retraining anything. On Cloudflare the injection score is the field cf.llm.prompt.injection_score, from 1 to 99, and the documentation states that lower scores indicate higher risk. That is the opposite of most in-house classifiers, so a copied threshold can be silently inverted. An illustrative rule expression is cf.llm.prompt.injection_score lt 20; the threshold is yours to tune. A self-hosted gateway can use a policy file like this one, where higher means riskier:

routes:
  /v1/chat/public:            # anonymous, tools disabled
    mode: enforce
    rules:
      - {id: PI-100, when: "score >= 0.95",                  action: block}
      - {id: PI-110, when: "score >= 0.80",                  action: challenge}
      - {id: PI-120, when: "flags has invisible_chars",      action: block}
      - {id: PI-130, when: "client_system_message",          action: block}
  /v1/chat/agent:             # authenticated, tools enabled
    mode: enforce
    rules:
      - {id: PI-200, when: "score >= 0.80",                  action: restrict_tools}
      - {id: PI-210, when: "session_hits_10m >= 3",          action: block_session}
  /v1/chat/internal:
    mode: log                 # shadow only
fail_policy: {public: closed, agent: closed, internal: open}

The useful actions go beyond block. Challenge adds friction, such as a CAPTCHA or re-authentication, for borderline scores. Restrict tools lets the request through but tells the application to run it without side-effecting tools, which turns a likely attack into a harmless chat. Tag passes the score downstream in a header so the application can require confirmation before acting. Every decision is logged with its rule ID, so false positives can be traced to the exact rule.

Conversations and rate limits

Per-request scoring misses attacks spread over several turns, and it ignores the most useful signal: attackers retry. Keep a small per-session and per-key state: hits in the last ten minutes, maximum score so far, and request rate. An attacker searching for a bypass sends many near-identical prompts with small changes, which looks nothing like a user. Rate-limit by API key and session on WAF hits, not only on requests, and escalate to a session block after repeated hits. Scoring the last few user turns concatenated, as well as each one alone, catches payloads split across messages, at the cost of more classifier tokens.

What a WAF cannot see

Be precise about the limits, because the WAF tends to be oversold internally.

  • Indirect injection. When the application retrieves a web page, an email or a document, the instructions in it never cross the edge. This is the dominant risk for agents; see indirect prompt injection. Prompt Shields has a separate document-attack mode for this reason, and you call it from inside the application.
  • Non-text inputs. Images and audio need a multimodal scanner in the application.
  • Adaptive attackers. A determined attacker with query access can usually find phrasings below any threshold. Measure this with adaptive attacks, as described in prompt injection evaluation.
  • Authorization. A WAF cannot know whether this user may delete that record. Only code that checks permissions on each tool call can, as covered in direct prompt injection.

So treat the WAF as a detection and friction layer that lowers attack volume and gives you telemetry. It is not the boundary. Design as if a crafted prompt will reach the model, and make sure the model cannot do anything dangerous with it.

Latency and failure policy

A classifier adds latency before the first token. Small encoder classifiers usually run in tens of milliseconds on modest hardware, but measure your own p99 under load, including the windowing for long prompts. Give the WAF a hard timeout and decide per route what happens on timeout or error. Public, tool-enabled routes should fail closed. Low-risk internal routes can fail open, with an alert, so a WAF outage does not become a product outage. Run scoring in parallel with the cheap parts of request handling, such as auth and quota checks, not after them.

Rolling out without breaking users

Never start in enforce mode. A safe rollout has four steps.

  1. Shadow. Log scores and would-be actions for two weeks of real traffic. Sample the top-scoring requests and label them by hand.
  2. Measure. Estimate the false-positive rate per route from the labels. Benign prompts about security, role-play and code often score high.
  3. Enforce narrowly. Turn on block only for the highest-confidence rules on the riskiest route, with challenge or restrict-tools below that.
  4. Operate. Give users a way to report a block, review it weekly, add exceptions with expiry dates, and re-run the evaluation set whenever the classifier or a rule changes.

Worked example: a support assistant with refunds

Take a support assistant with a public chat route that has no tools and an authenticated agent route that can issue refunds of up to 50 dollars. The numbers in this example are illustrative. In shadow mode over two weeks, the public route sees 400,000 requests and 1,200 score above 0.8. Hand-labelling 300 of those finds 180 genuine injection attempts and 120 benign, mostly customers pasting error messages and asking about jailbreaks. Blocking at 0.8 would therefore hit hundreds of legitimate users. Instead the team blocks above 0.95, where the sample held almost no benign prompts, and challenges between 0.8 and 0.95.

On the agent route the team never blocks on score alone. A score above 0.8 sets the restrict-tools action: the conversation continues, but the refund tool is unavailable for that turn and the user is asked to confirm in the UI. Three hits in ten minutes block the session. Meanwhile the refund tool itself checks the order owner and the amount in code. When a later red-team run gets an instruction through the WAF via a pasted screenshot, the refund still fails the ownership check. That is the layering working as intended.

Failure modes

  • Inverted thresholds. A rule copied between products where one treats low scores as risky and the other high.
  • Parser differentials. The WAF and the app parse JSON differently, for example with duplicate keys, so the WAF scores one string and the model sees another.
  • Truncated scoring. Only the first window of a long prompt is scored.
  • Silent fail-open. The classifier times out under load and every request passes with no alert.
  • False confidence. Teams skip tool authorization because the WAF exists.

What to do next

  1. List your LLM routes and mark which have tools, which are public, and which process retrieved content.
  2. Put a gateway extractor in front of them that understands the messages format and fails closed on bad JSON.
  3. Normalise text, flag invisible characters, and score overlapping windows.
  4. Write a per-route policy with block, challenge, restrict-tools and tag, and confirm each score's direction.
  5. Run shadow mode, label a sample, and enforce only where the false-positive rate is acceptable.
  6. Add an in-app scanner for documents and tool output, and check authorization on every tool call.
Key takeaway: A prompt injection WAF extracts prompts from API bodies, normalises them, scores them with a classifier and turns scores into per-route actions. It lowers attack volume and gives you telemetry, but it cannot see retrieved content and it can be evaded adaptively. Roll it out in shadow mode, check each score's direction, and keep authorization in the code that runs tools.