An AI blue team is the defensive function for systems built on language models: the people and the machinery that keep an inventory of what AI is running, collect the telemetry that makes attacks visible, write and test the detections, respond when one fires, and prove on a schedule that the controls still work. Most organisations already have a security operations centre; the AI blue team is not a replacement for it but the specialist capability that makes the SOC effective on a new class of system.

This article treats the blue team as a program. Alert handling is covered in the SOC playbook for LLM incidents and investigation in AI forensics; here the focus is what comes before and around them: charter and scope, inventory, mapping threats to coverage, detection engineering with tests, the loop with the red team, continuous control validation, hunting, and how to measure and staff the work.

Charter and boundaries

Write the charter down, because AI systems sit across team boundaries and unowned work falls through them. A workable charter gives the AI blue team four responsibilities: maintain the AI asset inventory and the telemetry contract every AI system must meet; own the AI-specific detection catalogue and its tests; support the SOC and incident response on AI alerts as the escalation tier; and run continuous validation of AI controls such as guardrails, tool permissions and output filters.

Equally important is what it does not own. Product teams own their agents and fix their bugs. The red team owns offensive testing, which the blue team consumes. The SOC owns first-line triage and the on-call rota. Model governance owns approval of new models and uses. Drawing these lines early prevents the most common failure, in which the blue team becomes the default owner of every AI risk and has time for none of them.

The AI asset inventory

You cannot defend what you do not know is running. The AI asset inventory lists every model endpoint, agent, tool, MCP server, vector store, prompt template and credential that grants model access, with an owner and the data classes each touches. It is the context every detection needs: an alert that an agent called a payment tool means little until you know which agent, who owns it and whether that tool is in its approved set.

# ai_inventory.yaml (one entry per agent; tools and data classes are the detection context)
- id: support-agent
  owner: team-support-eng
  model_endpoints: [gateway/primary-chat]
  tools: [lookup_order, issue_refund, escalate]
  untrusted_inputs: [customer_email, order_notes]
  data_classes: [customer_pii, order_history]
  egress_allowed: [api.internal/orders]
  keys: [vault:ai/support-agent/gateway-token]

Keep the inventory in version control and reconcile it against reality on a schedule. Three reconciliation checks find most gaps: gateway or provider billing keys that no inventory entry claims, outbound traffic to model provider domains from hosts that are not registered, and tool calls in telemetry naming tools absent from the agent's entry. Each mismatch is either an inventory fix or shadow AI to bring under management.

Mapping threats to coverage

Coverage mapping turns a threat list into a work list. Use a published taxonomy so gaps are visible and comparable: the OWASP Top 10 for LLM Applications (2025 edition) for application risks, and MITRE ATLAS for adversary techniques against AI systems. For each risk, record the telemetry required, the preventive control, the detection and the test that proves the detection works. An empty column is a gap.

OWASP LLM risk (2025)TelemetryPreventive controlDetection
LLM01 Prompt InjectionTool inputs and outputs with provenanceLeast-privilege tools, approval for writesAction following untrusted content
LLM02 Sensitive Information DisclosureOutputs, egressOutput filtering, scoped retrievalCanary or PII pattern in output
LLM05 Improper Output HandlingDownstream sink logsEncode and validate model outputModel output reaching a shell or query
LLM06 Excessive AgencyTool calls per sessionTool allowlists per agentTool outside the agent inventory
LLM07 System Prompt LeakageOutputsNo secrets in promptsSystem prompt canary in output
LLM10 Unbounded ConsumptionTokens and cost per keyQuotas and rate limitsSpend anomaly per key

The table shows six of the ten for space; the remaining risks (supply chain, data and model poisoning, vector and embedding weaknesses, and misinformation) lean more on preventive and governance controls than on runtime detection, but they belong in the same sheet. The threat modelling method in LLM-specific threat modelling produces the per-system inputs for this mapping.

The AI blue team loop: inventory, telemetry, detections, response, and validation by attack replayAI asset inventorymodels, agents, tools, keysGateway and runtimeprompts, calls, tool I/OTelemetry pipelinenormalised AI eventscontexteventsDetectionsrules as codeHuntshypothesesTriageSOC and IRRed team corpusattacks plus benign setDetection tests in CIrecall and false positivesgate changesControl canariesprobes and canary tokensschedulednew casesevery incident and every red team finding becomes a test
The program loop: inventory and runtime telemetry feed detections and hunts; the red team corpus and control canaries continuously test them; incidents feed new cases back.

Detection engineering with tests

Treat detections as code: each lives in version control with an owner, a mapped risk, a severity, a runbook link and, critically, tests. The test set has two halves. Attack examples come from the red team's corpus and from past incidents and must fire. Benign examples are sampled from real redacted traffic and must not. A change to the rule, to the guardrail or to the agent re-runs both, and the change is blocked if recall drops or the false-positive rate rises past its budget.

The example below is the detection that matters most for agents: an action-capable tool called after the session ingested untrusted content, with an argument that did not come from the user. It works on normalised session events from the telemetry pipeline.

WRITE_TOOLS = {"issue_refund", "send_email", "update_record"}

def tainted_action(session):
    # Fire when a write tool runs after untrusted content and uses a value the user never supplied.
    tainted, user_text, hits = False, "", []
    for ev in session["events"]:
        if ev["type"] == "user_message":
            user_text += " " + ev["text"].lower()
        elif ev["type"] == "tool_result" and ev.get("source") == "untrusted":
            tainted = True
        elif ev["type"] == "tool_call" and ev["name"] in WRITE_TOOLS and tainted:
            novel = [v for v in ev["args"].values()
                     if isinstance(v, (str, int)) and str(v).lower() not in user_text]
            if novel:
                hits.append({"tool": ev["name"], "novel_args": novel, "ts": ev["ts"]})
    return hits

def evaluate(rule, attacks, benign):
    tp = sum(1 for s in attacks if rule(s))
    fp = sum(1 for s in benign if rule(s))
    return {"recall": tp / len(attacks), "fp_rate": fp / len(benign)}

def test_tainted_action():
    m = evaluate(tainted_action, load("corpus/injection_actions/*.json"), load("corpus/benign_sample/*.json"))
    assert m["recall"] >= 0.9, m
    assert m["fp_rate"] <= 0.01, m

The rule is deliberately simple and will miss attacks that launder values through the user's own words; that is acceptable because it is one layer, and the recall number tells you honestly what it catches. The thresholds belong to you: set them from the first measured run, then ratchet. A detection without a measured false-positive rate is a source of alert fatigue waiting to happen.

The purple-team loop

The red team finds ways in; the blue team turns each one into permanent coverage. Run the loop as a scheduled exercise rather than a report hand-off. The red team executes a campaign against a staging agent with production telemetry enabled. The blue team watches in real time and records, for each attack step, whether it was prevented, detected, or invisible. Every invisible step becomes either a telemetry gap, a new detection or a new preventive control, and every attack transcript joins the detection test corpus. The red team architecture article covers the attack side and the corpus format.

The output that matters is the change in that three-way split over time. A program is working when the share of invisible steps shrinks quarter on quarter, not when the number of alerts grows.

Continuous control validation

Controls decay silently: a guardrail is disabled during an incident and never re-enabled, a model upgrade changes refusal behaviour, a new tool ships without an allowlist entry. Continuous validation catches this with two cheap mechanisms.

  • Control probes. A scheduled job sends a fixed set of known-bad requests through the production path with a test identity, and asserts each is blocked or flagged, and that the expected alert arrives in the SIEM within its time budget. A probe that gets through is an incident, whatever the reason.
  • Canary tokens. Unique, meaningless strings placed in system prompts, in restricted documents in the retrieval index, and in fixture records. They have no legitimate reason to appear in model output or outbound traffic, so a match is a high-confidence signal of prompt leakage or data exfiltration, with almost no false positives.

Run probes after every deployment of a gateway, guardrail or agent, as well as on a timer, so a regression is found in minutes rather than at the next audit.

Hunting hypotheses

Detections catch what you have already imagined. Hunting tests hypotheses against stored telemetry to find what you have not. Good AI hunting hypotheses are specific and falsifiable:

  • Some agent has called a tool that is not in its inventory entry in the last 30 days.
  • Some API key is used from more distinct hosts than the service it belongs to runs on.
  • Some sessions contain retrieved documents with instruction-like phrasing addressed to an assistant.
  • Some users receive outputs far longer than their prompts justify, consistent with data extraction.
  • Some agent sessions chain more tool calls than any legitimate workflow needs.

Each hunt ends in one of three outcomes: nothing found, which is still recorded with the query; an incident handed to response; or a pattern worth automating, which becomes a detection with tests.

Worked example: an email-triage agent

Consider a team protecting an email-triage agent that reads incoming mail, summarises it and can create tickets and draft replies. The inventory shows its untrusted input is the mail body and its write tools are ticket creation and reply drafting. Coverage mapping shows prompt injection has telemetry and a preventive control (drafts require human send) but no detection, and system prompt leakage has neither.

In the first two weeks the team adds tool-result provenance to the telemetry, deploys the tainted-action rule scoped to the two write tools, and places a canary token in the system prompt. A purple-team session then replays 40 injection emails from the red team corpus. Suppose the result is that 31 are detected, 6 are prevented by the human-send step without detection, and 3 are invisible because the injected instruction changed only the ticket priority, an argument the user never mentions in any case. Those three are the finding: the team adds a rule for priority changes that are not justified by the summary and adds the three transcripts to the test corpus. The numbers are illustrative; the shape of the outcome is typical.

Metrics and staffing

Measure the program with a few numbers that drive decisions: coverage, as the share of mapped risks with telemetry, control and tested detection for each inventoried system; detection quality, as recall and false-positive rate per rule from its test set; time to detect and time to contain from incidents and purple exercises; inventory drift, as unregistered assets found per reconciliation; and probe pass rate.

On staffing, a small program can start with two or three engineers who combine detection engineering with enough model and agent knowledge to read a trajectory, embedded alongside the SOC rather than separate from it. Rotate SOC analysts through the team so AI literacy spreads to first-line triage.

Failure modes

  • Detections without tests. Rules that have never fired on a real attack. Every rule needs attack and benign examples before it ships.
  • Inventory as a one-off. Accurate in the quarter it was written. Reconcile automatically.
  • Prompt-only telemetry. Logging prompts and responses but not tool calls and their provenance, which is where agent attacks become visible.
  • Silent control drift. No probes, so a disabled guardrail is found by an attacker.
  • Ownership creep. The blue team becomes the fixer of every agent bug. Hold the charter.
  • Alert volume as success. More alerts usually means worse precision, not better defence.

What to do next

  1. Write the charter, including what the team does not own, and agree it with the SOC, the red team and product.
  2. Build the AI asset inventory in version control and schedule the three reconciliation checks.
  3. Fill in the coverage sheet against the OWASP LLM Top 10 for every inventoried system and rank the gaps.
  4. Add tool-call provenance to telemetry, then ship the tainted-action detection with an attack and benign test set.
  5. Place canary tokens in system prompts and restricted documents, and alert on any match.
  6. Schedule control probes after every relevant deployment and hourly.
  7. Run a purple-team exercise this quarter and track the prevented, detected and invisible split over time.
Key takeaway: An AI blue team is a program, not a set of alerts. Charter it with clear boundaries, keep a reconciled inventory of every model, agent, tool and key, map threats to telemetry, controls and tested detections, and prove the controls work continuously with probes and canary tokens. Turn every red team finding and incident into a permanent test, and judge progress by how few attack steps remain invisible.