LLM SECURITY · WHAT A PATTERN SCANNER CAN AND CANNOT SEE

Prompt Injection Scanner

Score text against eight visible injection rules, then see on eight labelled samples which attacks the rules catch, which they miss and which harmless text they flag.

Direct tool · updates as you edit

Your experiment

Start with the direct override at threshold 4. Step through the labelled samples, move the threshold from 1 to 8, turn on base64 decoding and turn off normalisation, and watch the caught, missed and false-alarm counts.

Every input recomputes the result immediately; there is no animation because nothing here unfolds over time. An input outside its allowed range is rejected with a message and the previous valid result stays on screen.

Rules, weights and matches

Computed data

Metrics

The scanner is eight regular expressions with weights, shown in the text panel; a text is flagged when its score reaches the threshold. It is a teaching tool, not a product: its detection rate on real attacks is not established, and the labelled samples only demonstrate the failure modes every pattern scanner shares. Treat retrieved text and tool output as data, give tools least privilege and confirm side effects with a person regardless of what a scanner says.

Eight visible rules

The scanner adds up weights for patterns that often appear in injected instructions: an override such as ignore all previous instructions (4), a request for the system prompt (4), a fake system or admin marker (3), a link that carries data in its query string (3), a request to send data somewhere (3), a persona switch (2), zero-width characters (2) and a long base64-like run (1). Whatever you put in Text to scan, a user message, a retrieved page or a tool output, is scored the same way, and it is flagged when the score reaches the Flag at score threshold. The direct override scores 8 at the opening threshold of 4: CAUGHT INJECTION.

What it catches and misses

At threshold 4 the rules catch 3 of 6 labelled injections and raise 1 false alarm. The product-page attack scores 6 from its [system] marker and its request to forward the conversation history. The data-carrying image link scores only 3 and is missed. The paraphrase, the guidance you were given earlier no longer applies, matches nothing and scores 0: MISSED INJECTION. No pattern list can enumerate every way to say the same thing.

The false alarm

The security newsletter quotes the phrase ignore previous instructions in order to warn readers, and the override rule cannot tell quoting from instructing: score 4, FALSE ALARM. Its score is above the image link's 3, so no threshold catches the image link without also flagging the newsletter. Raising the threshold to 5 removes the false alarm and still catches 3; at 8 only 1 of 6 is caught. A threshold trades misses against false alarms; it never removes both.

Evasion and normalisation

The zero-width sample hides invisible characters inside the words of an override. With Normalise text and remove zero-width characters on it scores 6, the override plus the zero-width rule; turn normalisation off and the override no longer matches, the score falls to 2 and the scanner catches only 2 of 6. The base64 sample scores 1 until Decode base64 runs and scan them is turned on; then the decoded text matches the override and the request for the system prompt, it scores 9, and at threshold 3 the scanner catches 5 of 6.

Reading the tool and its limits

The bars score every labelled sample against the threshold, the lanes show the rules matched in your text, the table marks each sample right or wrong, and the text panel lists all rules with their weights. For your own text there is no label, so the result is only FLAGGED or NOT FLAGGED. A scanner is one layer: keep retrieved content separate from instructions, limit what tools can do and ask a person before irreversible actions.