An indicator of compromise (IoC) is an observable that, when seen, suggests an attack has happened or is under way: a file hash, a command-and-control domain, a registry key. Classic IoC feeds are built for malware and networks, and an LLM application rarely matches them, because the attack arrives as ordinary text and its effects are ordinary-looking tool calls. A poisoned web page carries no malicious binary; an exfiltration can be a markdown image whose URL contains a customer's email address.
This article builds a practical IoC programme for LLM and agent systems: what counts as an indicator, where to observe it, a catalogue of indicators that have proven useful, how to normalise text so matching survives trivial obfuscation, how to store indicators as versioned data, canary tokens, a worked investigation and the lifecycle that keeps the list from rotting. It complements LLM incident response, which decides what to do once an indicator fires, and AI forensics, which reconstructs what happened.
Compromise, attack and anomaly
It helps to separate three ideas that are often blurred. An indicator of compromise is evidence that something bad already happened: a known injection payload found in a retrieved document, a system-prompt canary appearing in an answer. An indicator of attack describes behaviour in progress, such as an agent calling a tool it has never called before, immediately after reading external content. An anomaly is merely unusual, such as token volume tripling, and needs context before it means anything. A good programme uses all three, but only the first two belong in a matching engine with automatic alerts; anomalies feed triage.
Indicators also differ in form. Atomic indicators are single values: a domain, an API key identifier, a hash of a normalised payload. Computed indicators are derived: a shingle signature that matches paraphrases of a known payload, or an entropy score for a URL query string. Behavioural indicators are sequences: read untrusted content, then call a send or fetch tool with arguments that contain data from a private source. Atomic indicators are cheap and precise but easy to evade; behavioural ones are robust but noisier. In LLM systems the behavioural class carries more weight than in classic endpoint security, because the payload text can be rewritten endlessly while the harmful effect cannot.
Observation points
Indicators are only as good as the telemetry they run over. Five observation points cover most incidents. The API gateway sees keys, source addresses and volumes. Context assembly sees what the model will read: retrieved chunks, memory entries, tool results and their provenance, which is why a per-request context manifest is the single most valuable log you can keep. The model call sees prompt and output. The tool layer sees every call, its arguments and any network egress. Offline, the corpus, the agent memory store and the model artefacts can be scanned at ingest and on a schedule. The logging requirements are set out in audit logging for LLM apps.
A starting catalogue
The following catalogue is a starting set. Each row names the observation point and the main source of false positives, because an indicator with no stated false-positive story will eventually be ignored.
| Indicator | Class | Observed at | False positives |
|---|---|---|---|
| Hash of a known injection payload after normalisation | Atomic | Ingest, context assembly | Security articles that quote the payload |
| Unicode Tags block (U+E0000 to U+E007F) or long zero-width runs in content | Atomic | Ingest, context assembly | Some emoji flag sequences use tag characters |
| System-prompt canary token in model output | Atomic | Output filter | Almost none |
| Corpus canary document retrieved for an unrelated query | Atomic | Retrieval log | Overly broad embedding matches |
| Markdown image or link to a non-allowlisted domain with a long query string | Computed | Output, renderer | Legitimate links users asked for |
| First-seen egress domain in a tool call after reading external content | Behavioural | Tool layer | New but valid sites in browsing agents |
| Private-source data (PII, secrets) in arguments of an outbound tool | Behavioural | Tool layer | User-requested sends |
| Instruction-like text written to long-term agent memory | Computed | Memory writes | Users storing their own preferences |
| API key used from a new network with a sharp volume rise | Anomaly plus atomic | Gateway | New deployments, CI jobs |
| Model or adapter file hash not in the signed release list | Atomic | Model loader | Unregistered internal experiments |
Two rows deserve emphasis. Canaries are the highest-precision indicators available, because no legitimate process should ever emit them. And the egress rows catch the outcome that matters most, data leaving, regardless of how the injection was phrased.
Normalise before you match
Text indicators fail when an attacker inserts a zero-width space or swaps a Latin letter for a full-width one. Normalise before hashing so that trivially different variants of a payload collapse to one value, and use shingles to catch paraphrases. The same normalisation must run at every observation point, or a payload hashed at ingest will not match the same payload at context assembly.
import hashlib
import re
import unicodedata
ZERO_WIDTH = dict.fromkeys(map(ord, ""))
TAGS = re.compile("[\U000E0000-\U000E007F]")
def normalise(text: str) -> str:
t = unicodedata.normalize("NFKC", text) # full-width and compatibility forms
t = TAGS.sub("", t).translate(ZERO_WIDTH) # invisible carriers
t = t.casefold()
return re.sub(r"\s+", " ", t).strip()
def payload_hash(text: str) -> str:
return hashlib.sha256(normalise(text).encode("utf-8")).hexdigest()
def shingles(text: str, k: int = 5) -> set:
words = normalise(text).split()
return {" ".join(words[i:i + k]) for i in range(max(1, len(words) - k + 1))}
def similarity(a: set, b: set) -> float:
return len(a & b) / max(1, len(a | b)) # Jaccard; alert above a tuned threshold
def hidden_char_count(text: str) -> int:
return len(TAGS.findall(text)) + sum(text.count(chr(c)) for c in ZERO_WIDTH)Record the hidden-character count as its own signal before normalising it away, since the presence of invisible characters in a retrieved support ticket is itself suspicious. Signature-based detection of indirect injection goes further into scoring and its limits.
Indicators as versioned data
Indicators belong in a version-controlled repository, not in code. Each record needs enough metadata to decide what to do when it fires and when to delete it. If you exchange indicators with partners, STIX 2.1 Indicator objects can carry them, with custom properties for LLM-specific observables; internally, a simple schema is enough.
- id: ioc-2026-0142
type: payload_hash # payload_hash | shingle_set | domain | canary | tool_sequence | artefact_hash
value: "9c1f...e07a"
scope: [ingest, context] # where the matcher evaluates it
confidence: high
action: quarantine_doc # alert | quarantine_doc | block_tool_call | revoke_key
source: case-2026-031 # the investigation that produced it
owner: appsec-oncall
created: 2026-09-14
expires: 2027-03-14 # forces review; expired indicators stop matching
notes: "Exfil via markdown image to attacker domain; see case timeline"The matcher loads this set, compiles each type into a fast structure, such as a hash set for atomic values, a MinHash index for shingles, and a small state machine per session for tool sequences, and stamps every alert with the indicator id and rule-set version. That stamp is what lets you answer, months later, why a document was quarantined.
def on_tool_call(session, call, iocs):
hits = []
domain = egress_domain(call)
if domain and domain in iocs.bad_domains:
hits.append(("domain", domain))
if domain and domain not in session.seen_domains and session.read_untrusted:
hits.append(("first_seen_egress_after_untrusted", domain))
if call.tool in OUTBOUND_TOOLS and contains_private_data(call.args, session.private_values):
hits.append(("private_data_outbound", call.tool))
for kind, val in hits:
emit_alert(session.id, kind, val, ruleset=iocs.version)
return "block" if any(k == "domain" for k, _ in hits) else "allow"
Sources and sharing
Your own incidents and red-team exercises are the best source of indicators, because they reflect your tools, your corpus and your users. External sources add breadth. MITRE ATLAS catalogues adversary techniques against AI systems and is useful for checking that each technique relevant to you maps to at least one indicator. The OWASP Top 10 for LLM Applications gives a coarser checklist. Public write-ups of prompt-injection and data-exfiltration research supply concrete payload patterns and exfiltration channels, such as image URLs, link previews and tool parameters. Treat imported payload text as untrusted data: store it hashed or in isolated fixtures, and never let it reach a model prompt in production. When you share indicators outward, strip customer data and internal hostnames, and share behavioural descriptions alongside hashes, since a hash of your normalised text is only useful to someone running the same normalisation.
Canary tokens
A canary is a random token planted where only a compromise would expose it. Put a unique token in each system prompt and alert if it appears in any output: that is a reliable sign of prompt extraction. Plant a few canary documents in the retrieval corpus with distinctive but harmless content; if one appears in a context for an unrelated query, or its token turns up in egress, retrieval or access control has been abused. Plant canary credentials in places an agent should never read, such as a decoy secrets file, and alert on any use. Rotate canaries per environment so that a hit tells you which deployment leaked. Keep them random and unguessable, and never reuse a token that has appeared in an incident report.
Worked example: from one alert to four documents
A support agent can read tickets, search the knowledge base and send email. At 10:42 the output filter flags a reply containing a markdown image whose URL points to an unknown domain with a 900-character query string: the computed indicator. The tool layer has logged no email, but the chat renderer would have fetched the image, so the data would have left anyway. Containment strips external images from rendered output.
Pivoting is where indicators pay off. The context manifest for that request lists a ticket received at 10:39; its body contains 214 tag-block characters, and decoding them reveals an instruction to summarise the customer's previous orders into an image URL. The team creates three indicators: the normalised payload hash, its shingle set and the domain. Back-searching the corpus and 30 days of manifests with them finds two more tickets carrying paraphrases of the payload (shingle similarity 0.6 and 0.7, no hash match) and one earlier session where the image was rendered. That session becomes a disclosure case; the three tickets are quarantined; and the indicators are added to ingest scanning with a six-month expiry. Without the manifest and the shingle index, the investigation would have stopped at one alert.
Keeping indicators alive
Indicator sets rot in two directions. Stale atomic indicators, such as a domain the attacker abandoned months ago, cost little but produce noise when the domain is re-registered by someone else. Unreviewed noisy indicators teach responders to ignore alerts. Give every indicator an expiry and an owner, measure precision per indicator from the triage outcome of each alert, and retire or tighten anything whose precision falls below an agreed bar. Track coverage too: for every incident class in your playbooks, list the indicators that would have fired. Gaps in that table are the backlog. AI incident playbooks shows how to keep playbooks as linted data that can reference indicator ids directly.
Failure modes
- Matching raw text. One zero-width space defeats the hash. Normalise identically everywhere.
- Only atomic indicators. Payloads mutate freely; pair them with shingles and behavioural rules on egress.
- No context manifest. An output alert with no record of what the model read cannot be traced to its source.
- Indicators leaking into prompts. Do not paste payload text into an LLM-based classifier prompt without isolation; the indicator can inject the detector.
- Alert without action. Each indicator needs a defined response, or high-confidence hits wait in a queue while data leaves.
- Logging secrets to match them. Store hashes or tokenised forms of private values for the private-data rule, not the values themselves.
What to do next
- Confirm you log a context manifest per request: every chunk, its source and its trust level.
- Add a unique canary to each system prompt and an output filter that alerts on it.
- Implement the normalise function above and use it at ingest, context assembly and output.
- Write a first-seen-egress-after-untrusted-content rule for every outbound tool.
- Move indicators into a versioned YAML repository with owner, action and expiry fields.
- After the next incident or red-team exercise, convert its findings into indicators and back-search 30 days of logs with them.
- Review per-indicator precision monthly and retire anything nobody trusts.