Securing an LLM application is not done alone. The models come from labs that publish system cards, capability policies and bug bounty rules. The attack vocabulary comes from shared catalogues such as MITRE ATLAS and the OWASP GenAI Security Project. Much of the testing tooling is open source and maintained by labs and vendors. Government institutes evaluate frontier models, and a research community finds new attack classes every month. Together these form an ecosystem, and teams that know how to plug into it spend their effort on their own risks instead of rediscovering known ones.

This article is a map of that ecosystem from the point of view of an application team: what each part produces, what to take from it, and what to give back. It does not re-explain governance frameworks or the lab scaling policies; those are covered in AI safety frameworks, in depth. The ecosystem changes quickly, so treat the specific names here as a starting list and check each source's current state before relying on it.

Advertisement

The map: six kinds of participant

ParticipantExamplesWhat they produceWhat you take
Model labsAnthropic, Google DeepMind, Meta, OpenAI, MicrosoftSystem and model cards, usage policies, capability policies, bug bounty and vulnerability reward programmesModel limits and known risks; disclosure routes
Threat knowledge basesMITRE ATLAS, OWASP GenAI Security Project, AVIDTechnique catalogues, top-risk lists, case studies, vulnerability recordsShared vocabulary for threat models and findings
Open toolinggarak (NVIDIA), PyRIT (Microsoft), Purple Llama (Meta), promptfooScanners, red-team orchestrators, safety classifiers, benchmarksRepeatable tests in CI; input and output filters
Industry coalitionsCoalition for Secure AI (an OASIS project), Frontier Model Forum, Google's Secure AI FrameworkControl frameworks, best-practice papersControl checklists and reference architectures
Government institutesUK AI Security Institute, US Center for AI Standards and InnovationPre-deployment model evaluations, evaluation tools and researchEvaluation methods and independent results
CommunityDEF CON AI Village, academic and independent researchersNew attack classes, public red-team events, papersEarly warning of techniques that will reach you
Who produces what in AI security, and where an application team plugs inModel labssystem cards, policies, bountiesThreat knowledgeMITRE ATLAS, OWASP, AVIDOpen toolinggarak, PyRIT, Purple LlamaGovernment institutesUK AISI, US CAISICommunityAI Village, researchersYour application teamthreat model, CI scans, registerFindings tagged by IDATLAS technique per riskCoordinated disclosuremodel issue -> lab programmemodel factsvocabularyscannerseval methodstechniquesConsume vocabulary, tools and model facts; contribute findings back through disclosure and case studies.
Labs, knowledge bases, tooling, institutes and the community each feed an application team different inputs. The team's findings flow back as tagged records and coordinated disclosures.

The arrows show the one principle worth keeping: consume shared artefacts instead of inventing private ones, and send what you learn back through the channels the ecosystem already has. A private attack taxonomy cannot be compared with anyone else's; a finding tagged with an ATLAS technique can.

MITRE ATLAS: the shared vocabulary

ATLAS (Adversarial Threat Landscape for Artificial-Intelligence Systems) is MITRE's knowledge base of adversary tactics and techniques against AI-enabled systems, modelled on ATT&CK. Tactics are the adversary's goals, such as initial access or exfiltration; techniques are how they achieve them, each with a stable identifier. Two that every LLM team needs are AML.T0051, LLM Prompt Injection, with sub-techniques .000 for direct and .001 for indirect injection through retrieved content, and AML.T0054, LLM Jailbreak. ATLAS also publishes mitigations and case studies of real incidents.

The data is published as YAML in the mitre-atlas/atlas-data repository, with monthly content releases. At the time of writing the current layout is format version 6, with top-level keys including tactics, techniques, mitigations and case-studies, and dist/ATLAS-latest.yaml tracks the newest release. Loading it lets you validate that every finding in your register cites a real technique:

import urllib.request
import yaml   # pip install pyyaml

# Pin a dated release in CI (dist/v6/ATLAS-<version>.yaml); "latest" moves monthly.
URL = "https://raw.githubusercontent.com/mitre-atlas/atlas-data/main/dist/ATLAS-latest.yaml"

def load_atlas(url=URL):
    with urllib.request.urlopen(url, timeout=30) as r:
        doc = yaml.safe_load(r)
    techs = doc.get("techniques")
    if techs is None:                       # older layout: lists nested under matrices
        techs = [t for m in doc.get("matrices", []) for t in m.get("techniques", [])]
    print("ATLAS format", doc.get("format-version"), "techniques", len(techs))
    return {t["id"]: t.get("name", "") for t in techs}

atlas = load_atlas()

findings = [
    {"id": "F-12", "title": "Ticket body can steer summariser", "atlas": ["AML.T0051.001"]},
    {"id": "F-13", "title": "Role-play prompt bypasses refusal", "atlas": ["AML.T0054"]},
]
for f in findings:
    unknown = [a for a in f["atlas"] if a not in atlas]
    if unknown:
        raise SystemExit(f"{f['id']}: unknown ATLAS ids {unknown}")   # typo or retired id
    print(f["id"], [(a, atlas[a]) for a in f["atlas"]])

The loader tolerates the older nested layout, pins nothing by default, and fails loudly on unknown identifiers, which catches typos and techniques that were renamed. Pair ATLAS with the OWASP Top 10 for LLM Applications: OWASP ranks risks by prevalence in applications and is good for prioritising, while ATLAS describes adversary behaviour precisely and is good for tagging. AVID, the AI Vulnerability Database, is a community effort to record concrete failures of specific models in a structured way.

Advertisement

Open tooling: scanners, orchestrators and classifiers

Three kinds of open tool matter. Scanners send large batteries of known attack prompts and score responses automatically. NVIDIA's garak is the best known: generators connect it to a model, probes send attack families such as prompt injection, encoding tricks and DAN-style jailbreaks, and detectors judge the outputs. Red-team orchestrators such as Microsoft's PyRIT script multi-turn attacks, where one model attacks, another scores, and the conversation adapts. Classifiers and benchmarks such as Meta's Purple Llama project, which includes the Llama Guard and Prompt Guard classifiers and the CyberSecEval benchmarks, can be deployed as filters or used to measure a model.

# Nightly scan of the model endpoint your app uses (never run against production keys you
# cannot rate-limit). garak writes report files and prints their location when it finishes.
python -m pip install -U garak
python -m garak --list_probes > probes.txt          # the probe catalogue changes between versions

python -m garak --target_type openai --target_name "$MODEL_ID" --spec probes.promptinject
python -m garak --target_type openai --target_name "$MODEL_ID" --spec probes.encoding
python -m garak --target_type openai --target_name "$MODEL_ID" --spec probes.dan

Read scanner results as signals, not grades. A probe's hit rate depends on its detector, and detectors have false positives and negatives; a drop from 40 percent to 5 percent after a prompt change is meaningful, a single number in isolation is not. Scan the configuration you actually ship, with your system prompt and filters in the path, because a raw model and your application have different attack surfaces. Where the target is your own HTTP endpoint rather than a hosted model, check garak's documentation for its REST generator configuration. How these runs fit into a wider programme is covered in LLM red team architecture and LLM safety evals architecture.

What the labs publish, and how to read it

Model labs publish three things an application team should read for every model it adopts. System or model cards describe evaluations run before release, known limitations and the safeguards the provider applies. Usage policies define what your application may do with the model and are part of your compliance surface. Capability policies, such as responsible scaling or preparedness frameworks, describe thresholds for dangerous capabilities and the safeguards they trigger; the evaluation discipline behind them is explained in dangerous capability evaluations.

Extract, per model: the prompt-injection and jailbreak results the card reports and under what conditions; the provider-side filters that are on by default and which ones you can configure; data retention and training-use terms; and the disclosure route for model-level issues. Put these into your risk register next to the model version, because a model upgrade silently changes all of them.

Coordinated disclosure: sending findings back

Most things you find are yours to fix: an over-privileged tool, a retrieval source that lets attacker text in, an output that is rendered unescaped. Some are not. If a technique defeats the provider's own safeguards across applications, or a model reliably leaks something it should not, the lab needs to know. The major labs run bug bounty or vulnerability reward programmes, and several state that some model-behaviour issues, such as ordinary jailbreaks, are out of scope for security rewards and go through separate safety reporting channels instead. Read the current scope page before filing, because scopes change.

Title:      Indirect prompt injection via ticket body causes tool call (summariser agent)
Component:  model behaviour | our app | third-party plugin   <- decide before sending
Technique:  AML.T0051.001 (LLM Prompt Injection: Indirect)
Model:      provider, model id, date/time (UTC), API region if known
Setup:      system prompt summary, tools enabled, temperature, retrieval source
Repro:      minimal input + exact steps; success rate over N attempts (e.g. 7 of 20)
Impact:     what an attacker gains in THIS deployment; data or actions reachable
Mitigation: what we deployed on our side; what we believe the model side could change
Contact:    security@ address; disclosure timeline we propose

The component line is the important decision: a report that blames the model for a bug in your tool permissions wastes everyone's time, and one that blames your app for a cross-provider model weakness leaves other teams exposed. Report success rates over repeated attempts rather than a single transcript, since outputs are stochastic. Do not publish working attacks against a provider before its programme's disclosure window has passed.

Worked example: a support copilot plugs in

A team ships a copilot that summarises customer tickets and can draft refunds. Week one, they threat-model with ATLAS: ticket text is attacker-controlled, so indirect prompt injection (AML.T0051.001) is the headline risk, with jailbreaks (AML.T0054) second. Each risk gets an owner, a mitigation and an ATLAS tag in the register, validated by the loader above.

Week two, they add a nightly garak run of the prompt-injection and encoding probe families against a staging endpoint with the production system prompt, and store results so trends are visible. They read the model card for their provider, note its reported injection robustness and default filters, and record the model version. Week four, a red-team exercise finds that a ticket containing a fake refund policy causes a draft refund. Root cause: the refund tool accepted drafts without human approval. That is their bug, fixed with an approval step, and recorded as a case with its technique ID. A second finding, a role-play prompt that bypasses the provider's refusal for a harmful category across several apps, is filed with the provider's programme using the template.

Failure modes and trade-offs

FailureWhy it happensBetter practice
Private taxonomy nobody else can readTeam invents its own labelsTag findings with ATLAS IDs; prioritise with OWASP
Green scanner run treated as secureProbes cover known attacks onlyScanner plus human red teaming plus runtime controls
Scanner run against the bare modelEasier to set upScan the shipped configuration
Model upgrade changes risk silentlyCards and filters tied to versionsRegister records model version; rescan on change
Report rejected as out of scopeScope page not read, or app bug filed as model bugDecide the component first; read the current scope
Ecosystem noise consumes the teamFollowing every paper and toolQuarterly review of new ATLAS techniques relevant to your surface

The central trade-off is breadth against depth. Shared tools and catalogues give broad coverage cheaply but only of known attacks; your own red teaming is expensive but finds the flaws specific to your tools and data. Use the ecosystem for the first so your people can spend their time on the second.

What to do next

  1. List every model you use with its version, card link, default filters and disclosure route.
  2. Threat-model each application with ATLAS techniques and add a tag to every register entry.
  3. Load the ATLAS YAML in CI and fail on unknown technique IDs.
  4. Schedule a garak scan of your shipped configuration and keep the results as a trend.
  5. Choose one orchestrated multi-turn red-team tool and run it before each major prompt or model change.
  6. Adopt the disclosure template, name an owner, and read each provider's current programme scope.
  7. Review new ATLAS techniques and OWASP updates quarterly against your attack surface.
Key takeaway: The AI security ecosystem already provides most of what an application team needs to start: labs publish model cards, policies and disclosure programmes; MITRE ATLAS and OWASP provide a shared vocabulary; garak, PyRIT and Purple Llama provide repeatable tests and filters; institutes and the community push evaluation methods and new attacks forward. Tag every risk with an ATLAS technique, scan the configuration you ship, record model versions, decide whether a finding is yours or the model's, and send model-level issues back through the provider's programme.