No single company sees every way an AI system fails. A jailbreak that works on one model family, a prompt-injection path through a document parser, or a benchmark that rewards refusal over helpfulness tends to surface in many places at once. Open communities exist so that these observations become shared vocabulary, shared tests and shared records instead of private anecdotes. For a team shipping an LLM product they are a source of free, peer-reviewed defensive material and a place where its own lessons can help others.
The AI labs security ecosystem article maps who the participants are and how coordinated disclosure works. This article is the practitioner's companion. It shows how to consume community artefacts safely, as a software supply chain with pinned versions and reviews. It shows how to map your findings onto shared vocabularies, how to contribute back without leaking data or publishing a working attack, and how to run participation as a small, measurable programme rather than an occasional favour.
The landscape, organised by what each community produces
Open AI-safety communities are easiest to understand by what they produce. The table lists communities whose outputs are public and that accept outside contributions. Membership and governance differ: some are foundation projects, some are industry consortia, some are research groups that publish open tools. Check each project's current contribution guide before relying on the details below.
| Community | What it produces | How a team uses it | How outsiders contribute |
|---|---|---|---|
| OWASP GenAI Security Project (the former Top 10 for LLM project, an OWASP flagship since 2025) | Top 10 lists for LLM and agentic applications, checklists, guides | Shared risk vocabulary for threat models and audits | Working groups, document reviews, public comment |
| OWASP AI Exchange | Broad open guidance on AI threats and controls | Control catalogue beyond LLM-specific risks | Edits and reviews through the project |
| MITRE ATLAS | Matrix of adversary tactics and techniques against AI, case studies | Technique ids for findings and coverage maps | Case-study and technique submissions |
| AVID (AI Vulnerability Database) | Structured reports and vulnerability records with a taxonomy | Lookup of known model behaviours; a home for reports | Report submissions; open tooling |
| AI Incident Database (Responsible AI Collaborative) | Records of real-world AI harms | Base rates and precedents for risk reviews | Incident submissions with sources |
| MLCommons AI Risk and Reliability working group | AILuminate safety benchmark across 12 hazard categories | Comparable safety measurement of chat models | Working-group membership and review |
| Coalition for Secure AI (CoSAI, hosted by OASIS since July 2024) | Open guidance on AI supply chain security and risk governance | Reference controls, for example provenance for models | Workstream participation |
| DEF CON AI Village | Talks, public red-teaming events such as the Generative Red Team at DEF CON 31 | Training, recruiting, early attack ideas | Talks, event volunteering |
| Open evaluation tools: garak (NVIDIA), PyRIT (Microsoft), Inspect (UK AI Security Institute) | Scanners, red-team orchestration, evaluation frameworks | Automated probes in CI and red-team runs | Pull requests: probes, detectors, fixes |
These are complements, not substitutes. A typical mature team uses the OWASP and ATLAS taxonomies to label its threat model and findings, runs two open tools in CI, tracks AVID and incident records for the model families it ships, and contributes a few probes or reports a year. The OWASP LLM Top 10 article and the MITRE ATLAS article cover the two taxonomies in detail.
Two directions of flow
The relationship runs in two directions, and each direction has its own discipline. Inbound artefacts are code and data that will run against your systems or shape your decisions. Treat them like any third-party dependency. Outbound contributions describe weaknesses. Treat them like vulnerability disclosure, because that is what they are.
Consuming community artefacts as a supply chain
A probe set pulled from a community repository can change between runs. A new version adds 300 prompts, your attack success rate moves, and nobody can tell whether the model got worse or the test got harder. Community datasets also carry licences, sometimes non-commercial ones, and red-team corpora contain genuinely harmful text that needs access control. Benchmarks published in the open can leak into training data, which inflates scores for any model trained after publication, which is why benchmark maintainers often keep official test prompts private.
The remedy is ordinary supply-chain hygiene: a manifest that pins each artefact to a version and a content hash, records its licence and an internal owner, and fails the build when anything drifts.
import hashlib, json, pathlib, sys
ALLOWED_LICENSES = {"Apache-2.0", "MIT", "CC-BY-4.0", "CC-BY-SA-4.0"}
def sha256(path):
h = hashlib.sha256()
with open(path, "rb") as f:
for chunk in iter(lambda: f.read(1 << 20), b""):
h.update(chunk)
return h.hexdigest()
def verify(manifest_path):
manifest = json.loads(pathlib.Path(manifest_path).read_text())
problems = []
for a in manifest["artefacts"]:
# each entry: name, version, path, sha256, license, owner, restricted
if a["license"] not in ALLOWED_LICENSES:
problems.append(f"{a['name']}: licence {a['license']} needs legal review")
if sha256(a["path"]) != a["sha256"]:
problems.append(f"{a['name']}: content changed since pinned {a['version']}")
if a.get("restricted") and not a["path"].startswith("restricted/"):
problems.append(f"{a['name']}: harmful-content corpus outside restricted storage")
if problems:
sys.exit("\n".join(problems))
print(f"{len(manifest['artefacts'])} community artefacts verified")Upgrades then become deliberate events. Bump the pin, run the old and new versions side by side once, and record how much of any score change comes from the test rather than the model. Pin the tools themselves the same way. A scanner release that changes a detector's threshold will move your results just as much as a new probe set.
Mapping findings onto shared vocabularies
Shared vocabularies only help if findings are labelled consistently. Label each internal finding with the closest OWASP GenAI Top 10 entry and the closest ATLAS technique. Add the affected component class (model, retrieval, tool, output handling) and, for model-behaviour issues, the model family and version. For example, an indirect prompt injection through retrieved documents maps to OWASP LLM01 Prompt Injection and ATLAS AML.T0051 LLM Prompt Injection. A jailbreak of the base model's refusal behaviour maps to AML.T0054 LLM Jailbreak.
Labelled this way, your findings can be aggregated into coverage maps, compared with public records, and handed to a community venue without rewriting. Look up ids in the current release of each taxonomy rather than copying them from older documents. Both projects revise their entries.
Contributing back without causing harm
Most findings should never leave the building in their raw form. A finding in your own glue code, such as a tool with excessive permissions, is fixed internally. The only thing worth sharing is the general lesson, perhaps as a comment on a guidance document. A finding in an upstream component, such as a model family that follows instructions hidden in image text or a parser that strips delimiters, belongs to its vendor or maintainer first. Report it through their security channel and give them time to fix it before anything public appears. The red team process article covers how to write the finding itself.
Sanitisation is where most of the risk sits. A shareable record describes the behaviour, the conditions and the impact class. It never includes customer data, internal hostnames, secrets or a copy-paste working payload against a live system. Keep the exact payload private and publish its hash, so a maintainer who later asks for it through a disclosure channel can confirm they received the same one.
import hashlib, re
REDACTIONS = [
(re.compile(r"[\w.+-]+@[\w-]+\.[\w.]+"), "<email>"),
(re.compile(r"\b(sk|key|token)[-_][A-Za-z0-9]{16,}\b"), "<secret>"),
(re.compile(r"\b[\w-]+\.(internal|corp|local)\b"), "<internal-host>"),
(re.compile(r"\b\d{1,3}(\.\d{1,3}){3}\b"), "<ip>"),
]
def scrub(text):
for pattern, token in REDACTIONS:
text = pattern.sub(token, text)
return text
def shareable_record(finding):
"""Turn an internal finding into a record safe to send to a community venue."""
return {
"summary": scrub(finding["summary"]),
"component_class": finding["component_class"], # model / retrieval / tool / output
"affected": finding["affected"], # e.g. model family and version
"owasp": finding["owasp"], "atlas": finding["atlas"],
"impact": finding["impact"],
"reproduction": scrub(finding["conditions"]), # conditions, not the payload
"payload_sha256": hashlib.sha256(finding["payload"].encode()).hexdigest(),
"vendor_notified": finding["vendor_notified_date"], # required before publishing
}Regular expressions catch the obvious leaks, not the subtle ones, so a human reviewer reads every record before it goes out. The function refuses nothing by itself. Make the publishing step check that vendor_notified is set and that an approver is recorded.
Worked example: one finding, three destinations
A team runs a customer-support copilot on an open-weights model. Their red team finds that instructions hidden in white-on-white text in uploaded PDFs are followed by the model. They also find that their pinned scanner version has no probe for instructions arriving through extracted document text. The finding splits three ways.
First, the product fix stays internal: the copilot now strips hidden text layers and marks retrieved content as data, and a regression test pins the behaviour. Second, the model behaviour is upstream. The team reports it through the model publisher's security contact with the private payload, waits out the agreed window, then submits a sanitised report to AVID. That report is labelled LLM01 and AML.T0051 and carries a payload hash instead of the payload. The AVID article shows the record format. Third, the scanner gap is a tooling contribution. After checking the project's contribution guide and the company's open-source policy, an engineer submits a probe that generates benign canary instructions, such as asking the model to output a fixed marker string. A detector then checks for that marker. The probe tests the channel without shipping a harmful instruction.
No customer was harmed, so there is no incident record. Had a customer been harmed, a factual incident submission with public sources would follow the internal incident review. The AI incident databases article covers that path. Total cost was about two engineer-days beyond the fix. The return was an upstream fix that protects every deployment of that model, and a probe the team no longer has to maintain alone.
Running participation as a programme
Participation that depends on one enthusiast disappears when that person changes jobs. Make it a small programme with an owner, a time budget and a few metrics.
- Policy. An open-source contribution policy that covers AI-safety artefacts: who approves, which licences outbound code may use, whether a contributor licence agreement is acceptable, and what can never be shared.
- Budget. A fixed share of red-team time for upstream work, often 5 to 10 percent, so contributions are not always the first thing cut.
- Intake. One reviewer per quarter who reads release notes for every pinned artefact and proposes upgrades.
- Metrics. Artefacts pinned and current, findings mapped to shared taxonomies, upstream reports filed and their time to fix, contributions merged.
- Safe harbour. Before testing someone else's hosted model, read its terms and any published vulnerability disclosure or safe-harbour policy. Community events often provide explicit authorisation, while ad-hoc testing may not.
Failure modes
- Unpinned probes. Scores drift with the test set and trend lines become meaningless.
- Publishing before notification. A public jailbreak write-up against a deployed model, sent before the vendor heard of it, causes harm and burns trust with the maintainers you need.
- Payload leakage. Working attack strings, customer text or internal URLs in a public issue tracker. Once indexed, they cannot be withdrawn.
- Benchmark worship. A strong public safety score says nothing about your tools, retrieval sources or users. Community benchmarks complement product-specific evaluation; they do not replace it.
- Taxonomy drift. Ids copied from an old release no longer match the current entries.
- Licence surprises. A non-commercial dataset baked into a commercial product's CI.
Trade-offs
| Choice | Gains | Costs |
|---|---|---|
| Pin and review every artefact | Reproducible scores, controlled risk | Slower adoption of new probes |
| Track community latest | Newest attack coverage | Unexplained score changes |
| Contribute probes upstream | Shared maintenance, wider protection | Review cycles, licence and policy work |
| Keep findings private | No disclosure risk | Others rediscover the issue; no upstream fix |
| Join working groups | Influence on standards and benchmarks | Ongoing time commitment |
What to do next
- Inventory every community artefact your team uses (taxonomies, scanners, probe sets, benchmarks) and write a pinned manifest with versions, hashes, licences and owners.
- Add the manifest check to CI and move harmful-content corpora into access-controlled storage.
- Label your open findings with current OWASP GenAI Top 10 and MITRE ATLAS ids.
- Write a one-page contribution policy: approvers, licences, the notify-before-publish rule and what never leaves the company.
- Pick one upstream finding or scanner gap and take it through the full pipeline, from notification to a merged probe or published report.
- Assign a quarterly reviewer for artefact upgrades and report the programme metrics alongside your red-team results.