Security operations centres know how to handle a phishing alert or a suspicious login. An alert that says "assistant output contained an internal hostname after a retrieved web page" is different. The analyst on shift may not know what a system prompt is, the logs they need live in an application team's gateway rather than the SIEM, and the right first action depends on whether the agent behind the alert can send email or only chat.
This playbook is for the SOC layer of LLM security: getting the right telemetry into the SIEM, writing detections that fire on real attacks rather than every frustrated user, giving tier-one analysts triage runbooks they can follow at three in the morning, automating the safe parts, and handing confirmed incidents to the response team. Incident classes, severity, the containment ladder and the first hour of a declared incident are covered in LLM incident response in depth, and evidence handling in AI forensics. This page stops where those begin: at the moment an alert becomes an incident.
What the SOC owns for LLM systems
For conventional systems, the SOC watches infrastructure and identity signals and pages an application owner when something looks wrong. LLM applications add an attack surface that lives in content: the prompt, retrieved documents, model output and the tool calls the model decides to make. Most of it is invisible to network and endpoint tooling. If the SOC is to own first response, three things have to be true. The application must emit security events in a stable schema. Detections must be written against those events and mapped to known risk classes. And analysts need runbooks that translate an LLM-specific alert into actions they are allowed to take.
A useful division of labour: the SOC owns detection, triage, enrichment and pre-approved containment actions; the application team owns fixes to prompts, guardrails and tools; the incident response lead owns declared incidents. The playbook below assumes that split and is explicit about where each hand-off happens.
The telemetry contract
Everything starts at the LLM gateway or the agent runtime, the one place that sees every request, response and tool call. Emit one structured event per model call and one per tool call, and correlate them with a trace id. The annotated shape below is a practical minimum (strip the comments for real JSON); name the fields to match your SIEM's common schema.
{
"ts": "2026-10-03T02:14:07.412Z",
"event_type": "llm.tool_call",
"trace_id": "9f2c1e7a", // ties prompt, retrievals, output and tools together
"session_id": "s-4410", "user_id": "u-1832", "tenant": "acme",
"app": "support-agent", "model": "provider/model-name", "model_version": "2026-08-15",
"prompt_sha256": "4be1...", "output_sha256": "a03d...",
"input_sources": ["user", "retrieval:kb", "retrieval:web"],
"guardrail": {"injection_score": 0.91, "pii_types": ["email"], "blocked": false},
"tool": {"name": "send_email", "args_summary": "to=external domain, 1 attachment",
"outcome": "executed", "approval": "none"},
"tokens": {"in": 6120, "out": 410}, "cost_usd": 0.021, "latency_ms": 2380
}Three design choices make this usable. Hash full prompts and outputs in the SIEM event and keep the full text in a separate, access-controlled store keyed by trace id; analysts pull it during triage, and the SIEM does not become a second copy of every customer conversation. Record where each input came from, because the most important question in an injection alert is whether the instruction arrived from the user or from retrieved content. And log tool calls with an argument summary and outcome even when a guardrail blocked them, since blocked attempts are often the earliest signal. How to make these logs tamper-evident is covered in LLM audit logging architecture.
A detection catalogue
Start with a small catalogue of detections, each tied to a risk class so coverage can be reasoned about. The OWASP Top 10 for LLM Applications (2025 edition) is a convenient taxonomy; the rows below map to it.
| Detection | Signal | OWASP 2025 | Default severity |
|---|---|---|---|
| Injected instruction followed by a tool call | High injection score on retrieved content, then a side-effecting tool call in the same trace | LLM01 Prompt Injection, LLM06 Excessive Agency | High |
| Sensitive data in output | PII or secret detector fires on output, especially other customers' identifiers | LLM02 Sensitive Information Disclosure | High |
| System prompt disclosure | Output contains canary token or long overlap with the system prompt | LLM07 System Prompt Leakage | Medium |
| Jailbreak campaign | Many refusals or guardrail blocks for one account or IP in a short window | LLM01 Prompt Injection | Medium |
| Cost or volume spike | Tokens or spend per key far above its baseline | LLM10 Unbounded Consumption | Medium |
| Unusual tool destination | First-seen external domain or recipient in a tool call | LLM06 Excessive Agency | High |
Two cheap techniques give unusually precise signals. A canary token, a random string placed in the system prompt and never shown to users, turns system prompt leakage into an exact match. And correlating within a trace, injection-like content in a retrieved document followed by a tool call, is far more precise than alerting on injection scores alone, which fire on every user who pastes an article about prompt injection.
Writing the detections
Write detections as code, version them, and test them against recorded traffic. Sigma is a common vendor-neutral format; the rule below uses a custom logsource, because LLM gateways have no standard Sigma product category. Correlation across events, the "followed by" part, is expressed in your SIEM's own language or in Sigma's correlation extensions, depending on what your backend supports.
title: Side-effecting tool call in a trace with injected retrieved content
id: 6d0c2a8e-4b1f-4f7e-9a51-2f3c8b7e1d42
status: experimental
description: A tool that changes external state ran after retrieved content scored as likely injection.
logsource:
product: llm_gateway # custom logsource, mapped to your collector's index
category: tool_call
detection:
selection_tool:
event_type: llm.tool_call
tool.name:
- send_email
- create_ticket
- http_request
- write_file
tool.outcome: executed
selection_source:
input_sources|contains: 'retrieval:'
guardrail.injection_score|gte: 0.8
condition: selection_tool and selection_source
falsepositives:
- Security training content in the knowledge base
level: highVolume and cost detections are better as baselines than fixed thresholds. A sketch in SQL-like pseudocode:
-- per API key: tokens in the last 15 minutes versus that key's 7-day hourly profile
SELECT api_key, SUM(tokens_in + tokens_out) AS recent
FROM llm_events
WHERE ts > now() - INTERVAL '15 minutes'
GROUP BY api_key
HAVING SUM(tokens_in + tokens_out) > 6 * baseline_p95(api_key, hour_of_week(now()))
The first five minutes of any LLM alert
Every LLM alert gets the same first five minutes before any alert-specific step. Make the SOAR platform do as much of this as possible and attach the results to the case.
- Identify the actor. User, tenant, API key and authentication method. Is this a known customer, an internal tester, or an anonymous key?
- Identify the system. App, model and version, the guardrail configuration in force, and which tools the agent can call. An alert on a chat-only agent and one on an agent with email and file access are different severities.
- Pull the trace. Fetch the full prompt, retrieved documents, output and tool calls for the trace id from the content vault, under the analyst's own access.
- Locate the instruction source. Did the suspicious instruction come from the user, from retrieved content, or from a tool result?
- Check for impact. Did any tool call execute, did output reach a user or an external party, and did the same pattern occur in other sessions?
The output of triage is one of three outcomes: false positive with a tuning note, contained security event closed by the SOC, or escalation to a declared incident. Escalate whenever data left the trust boundary, a side-effecting tool executed on an attacker's instruction, or the same pattern appears across tenants.
Runbooks per alert type
Injected instruction with a tool call. Confirm from the trace that the instruction came from retrieved or tool-returned content. Check what the tool did: recipients, URLs, files. If it executed against an external destination, escalate immediately and ask the application owner to disable that tool for the app, which should be a pre-approved switch as described in agent kill switches. Quarantine the source document so it cannot be retrieved again, and search for other traces that retrieved it.
Sensitive data in output. Determine whose data it is. A user's own data echoed back is usually a tuning issue; another customer's data is a likely breach and goes straight to the incident lead with the trace attached, because notification clocks may start at discovery. Look for the retrieval or tool call that supplied it; the pattern is typically an authorisation gap, explained in data exfiltration through LLMs.
System prompt disclosure. A canary match is a confirmed disclosure. Check whether the prompt contains secrets, which it should not; if it does, rotate them now and treat it as a credential exposure. Otherwise record it, notify the owner, and block the account if this is part of a campaign.
Jailbreak campaign. Most attempts fail, so the question is whether any succeeded. Sample the outputs that were not blocked, look for policy-violating content, and rate-limit or suspend the account according to your terms of service. Feed successful prompts to the red team and the guardrail owners.
Cost spike. Distinguish a stolen key from a customer's runaway loop: new source IPs, unusual models, and prompts unrelated to the customer's product suggest theft. Revoke and reissue stolen keys; for loops, apply the per-key cap and tell the customer.
What to automate and what to leave to people
Automate enrichment fully and containment selectively. Safe automatic actions are those that are reversible, narrow and low impact: attaching traces and identity context to the case, adding a document to a retrieval blocklist pending review, tightening a single key's rate limit, or requiring human approval for one tool in one app. Actions with customer impact, such as suspending a tenant, disabling an agent, or revoking a production key, should be one-click for the analyst but not automatic, because detection precision for content-based attacks is rarely high enough to act unattended.
Every automated action should write back to the case what it did and how to undo it, and every one should expire unless a human confirms it. An automated blocklist that silently grows is how a SOC ends up breaking a product with no incident to show for it.
Worked example: an injected PDF that sent an email
At 02:14 the rule above fires on the support agent. SOAR enriches the case: the user is a paying customer, the agent can call send_email, and the trace shows a knowledge-base retrieval of a recently uploaded PDF scoring 0.91 for injection, followed by an email to an external domain with an attachment. The analyst pulls the trace and finds white-on-white text in the PDF instructing the assistant to email the conversation history to an outside address.
That is a side-effecting tool executed on an attacker's instruction, so the analyst escalates without further investigation and, using the pre-approved switch, puts send_email behind human approval for the app. SOAR adds the PDF to the retrieval blocklist and searches the last 30 days for other traces that retrieved it, finding two more sessions, neither of which executed a tool. The incident lead takes over at 02:31 with a case already holding the trace, the source document, the blast radius and the containment actions taken, each with its undo step. Seventeen minutes from alert to a well-formed hand-off is the goal; the runbook, not heroics, is what makes it repeatable.
Tuning and metrics
Content detections are noisy at first. Track, per rule, the alert volume, the share closed as false positive and the share escalated. A rule that is more than about nine in ten false positives after a month needs rework or retirement. Replay red-team transcripts and known attack corpora through the detection pipeline on every rule change, so tuning does not silently remove coverage. Measure mean time to detect from the event timestamp, mean time to acknowledge, and time from alert to hand-off, and keep a coverage table showing which risk classes have at least one tested detection.
Failure modes and trade-offs
- Blind SIEM. The gateway logs to an application store the SOC cannot query. Agree the schema and pipeline before the first incident.
- Content in the SIEM. Full prompts copied into a widely accessible SIEM create a new data exposure. Store hashes there and text in an access-controlled vault.
- Score-only alerts. Alerting on any high injection score floods analysts. Correlate with source and consequence.
- Missing source attribution. Without input provenance, analysts cannot tell a curious user from an injected document.
- Unbounded automation. Automatic tenant suspensions on noisy rules cause outages. Automate only reversible, narrow actions, with expiry.
- No path to the owner. The analyst confirms an issue at night and nobody can disable the tool. Pre-approve switches and keep the on-call rota current.
The core trade-off is precision against coverage: correlated rules miss novel attacks, broad rules exhaust analysts. Run both, with broad rules feeding a lower-priority queue reviewed in daylight. A second trade-off is privacy against investigability; the vault model keeps both by gating access rather than discarding content.
What to do next
- Define the LLM event schema with the application teams, including trace id, input sources, guardrail results and tool outcomes.
- Route events to the SIEM and full text to an access-controlled vault keyed by trace id.
- Place a canary token in each system prompt and alert on any match in output.
- Write the six detections in the catalogue as versioned rules and replay red-team traffic through them.
- Give tier-one analysts the five-step triage and the per-alert runbooks, with explicit escalation criteria.
- Pre-approve reversible containment switches per app, and make every automated action expire unless confirmed.
- Report per-rule false-positive rate, time to detect and time to hand-off monthly, and retire rules that do not earn their noise.