The G7 Hiroshima AI Process is the set of agreements the G7 started under Japan's presidency in May 2023 to give governments and AI developers a shared baseline for advanced AI systems, chiefly foundation models and generative AI. It is voluntary. No regulator fines anyone for ignoring it. Yet it matters to engineering teams for a practical reason: its Code of Conduct is the template behind a public reporting framework run by the OECD, and its eleven actions overlap heavily with what the EU AI Act, the NIST AI Risk Management Framework and ISO/IEC 42001 ask you to show. If you build the evidence once, you can answer all of them.
This article treats the Process as an engineering input rather than a diplomatic event. It explains what was actually agreed and when, reads each of the eleven actions as a control with an owner and a piece of evidence, shows how the OECD reporting framework turns that evidence into a public report, and gives a small controls register with a gap checker you can adapt. Country-level policy in Japan, which led the Process, is covered separately in Japan AI policy in depth.
What the Process actually produced
Three documents and one mechanism make up the Process. Keeping them apart avoids the most common confusion, which is treating the guiding principles, the code and the report as one thing.
| When | What | Who it addresses | What it asks for |
|---|---|---|---|
| May 2023 | Process launched at the G7 Hiroshima summit | G7 governments | Work towards common ground on generative AI governance |
| 30 Oct 2023 | International Guiding Principles for Organizations Developing Advanced AI Systems | Developers first, other AI actors as appropriate | Eleven high-level principles |
| 30 Oct 2023 | International Code of Conduct for Organizations Developing Advanced AI Systems | Organisations developing advanced AI | The same eleven items, written as actions with detail |
| Dec 2023 | Comprehensive Policy Framework, endorsed by G7 digital and tech ministers | Governments and all AI actors | Bundles an OECD report on generative AI, guiding principles extended to all AI actors, the Code, and project-based cooperation |
| Feb 2025 | OECD Hiroshima AI Process Reporting Framework launched | Developers, deployers and providers that volunteer | A public questionnaire answered against the Code's actions |
The first reporting round produced 19 published reports in spring 2025 from organisations including Anthropic, Fujitsu, Google, KDDI, Microsoft, NEC, NTT, OpenAI, Preferred Networks, Rakuten, Salesforce and SoftBank, alongside smaller firms and research bodies. The OECD accepts submissions on a rolling basis and encourages annual updates, and later OECD analysis covered roughly two dozen reports as more arrived. Treat any exact count as dated: the portal is the source of truth.
Two properties shape how you should use it. First, the Code is principles-based and risk-based: it says what outcome to achieve, not which technique to use, so the interpretation work is yours. Second, the report is public. Anything you write is read by customers, journalists, competitors and attackers, which changes how much security detail you can include.
The eleven actions as controls and evidence
The eleven actions below are paraphrased from the Code; read the official text before quoting it. The third column is the useful part for engineers: the artefact that proves the action is happening rather than intended.
| # | Action (paraphrased) | Evidence that shows it is real |
|---|---|---|
| 1 | Identify, evaluate and mitigate risks across the AI lifecycle, before and during deployment | Pre-release evaluation plan, red-team results, release gate record |
| 2 | Identify and mitigate vulnerabilities, incidents and patterns of misuse after deployment | Abuse monitoring dashboards, vulnerability intake, incident tickets |
| 3 | Publicly report capabilities, limitations and appropriate and inappropriate uses | Model cards or system cards per release, acceptable use policy |
| 4 | Share information responsibly and report incidents with industry, government, civil society and academia | Disclosure policy, sharing agreements, incident notifications sent |
| 5 | Develop, implement and disclose risk-based governance and risk management policies, including privacy | AI policy, risk register, governance committee minutes |
| 6 | Invest in robust security controls, including physical, cyber and insider-threat safeguards | Weights access controls, audit logs, insider-risk programme, pen tests |
| 7 | Deploy reliable content authentication and provenance where technically feasible, such as watermarking | Provenance metadata spec, watermark evaluation, user-facing labels |
| 8 | Prioritise research to mitigate societal, safety and security risks and invest in mitigations | Research budget lines, publications, funded mitigation work |
| 9 | Prioritise developing advanced AI to address global challenges such as climate, health and education | Programmes and deployments with measured outcomes |
| 10 | Advance the development and, where appropriate, adoption of international technical standards | Standards participation, adopted standards list |
| 11 | Implement appropriate data input measures and protections for personal data and intellectual property | Data provenance records, opt-out handling, PII filtering, licence review |
Grouped by who does the work, the actions fall into four clusters. Actions 1, 2 and 6 belong to engineering and security: evaluations, monitoring and the protection of weights and infrastructure. Actions 3, 4 and 7 are about transparency to outsiders: documentation, incident sharing and labelling generated content. Actions 5, 10 and 11 are governance and data management. Actions 8 and 9 are strategic investment choices, and they are where reports most often slide into marketing. A team that only deploys a third-party model can answer most of 2, 3, 5, 6 and 11 from its own operations and should say plainly that 1, 7 and 8 depend partly on the upstream provider.
Architecture: one evidence store, several frameworks
The structure that makes this tractable is the same one used for any compliance regime: a controls register keyed by requirement, an evidence store holding dated artefacts, and a crosswalk that lets one artefact answer several frameworks. The Hiroshima Code becomes one more column in that crosswalk rather than a separate project.
The crosswalk below is our own mapping, not an official one. It is approximate because the frameworks are written at different levels: the EU AI Act imposes legal obligations on providers of general-purpose AI models, NIST's framework organises risk activities into GOVERN, MAP, MEASURE and MANAGE functions, and ISO/IEC 42001 specifies an auditable management system.
| Code action | EU AI Act | NIST AI RMF | ISO/IEC 42001 |
|---|---|---|---|
| 1 Lifecycle risk | Systemic-risk evaluation and mitigation | MAP, MEASURE | AI risk assessment and treatment |
| 2 Post-deployment misuse | Serious incident tracking for systemic-risk models | MANAGE | Operation and monitoring |
| 3 Public reporting | Technical documentation and downstream information | GOVERN, MAP | Information for interested parties |
| 6 Security | Cybersecurity protection for systemic-risk models | MANAGE | Security controls via the management system |
| 7 Provenance | Generated-content transparency duties on AI system providers | MEASURE | Partly, through system design controls |
| 11 Data and IP | Copyright policy and training-content summary | MAP | Data management for AI systems |
Deeper treatment of each framework lives in the EU AI Act in depth, the NIST AI RMF and ISO/IEC 42001.
A controls register and gap checker in code
A spreadsheet works for ten controls; past that, a small typed register pays for itself because the gap check runs in CI and the report draft is generated from the same data the auditors see. The sketch below is deliberately plain Python with no dependencies.
from dataclasses import dataclass, field
from datetime import date, timedelta
@dataclass
class Evidence:
title: str
url: str # internal link; never paste secrets into a public report
produced: date
public: bool # may this artefact be cited in the HAIP report?
@dataclass
class Control:
action: str # Code action number, "1" .. "11"
name: str
owner: str
max_age_days: int = 365
evidence: list = field(default_factory=list)
def gaps(controls, today=None):
today = today or date.today()
covered = {c.action for c in controls}
out = [("action %s" % a, "no control mapped") for a in map(str, range(1, 12))
if a not in covered]
for c in controls:
if not c.owner:
out.append((c.name, "no owner"))
if not c.evidence:
out.append((c.name, "no evidence"))
continue
newest = max(e.produced for e in c.evidence)
if today - newest > timedelta(days=c.max_age_days):
out.append((c.name, "stale: newest evidence %s" % newest))
if not any(e.public for e in c.evidence):
out.append((c.name, "nothing citable publicly"))
return out
def draft_report(controls, action_text):
"""One section per action: what we do, citable evidence, known gaps."""
lines = []
for a in map(str, range(1, 12)):
lines.append("## Action %s: %s" % (a, action_text[a]))
mine = [c for c in controls if c.action == a]
if not mine:
lines.append("Not yet addressed. Planned owner and date: TODO.")
for c in mine:
cites = [e.title for e in c.evidence if e.public]
lines.append("- %s (owner: %s). Evidence: %s"
% (c.name, c.owner, "; ".join(cites) or "internal only"))
return "\n".join(lines)Three design choices matter. Evidence carries a public flag, because the same red-team result that proves action 1 internally may be unsafe to describe in a public report; the gap checker flags controls with nothing citable so you decide early what to summarise. Every control has a maximum evidence age, which catches the classic failure of reporting last year's evaluation for this year's model. And the draft writes Not yet addressed for unmapped actions instead of skipping them, because an honest gap reads better than an omission that a reviewer will find.
Worked example: a first report from a mid-size company
Consider a 300-person software company that fine-tunes an open-weight model and ships a writing assistant inside its product in Europe and Japan. It wants to file a HAIP report because an enterprise customer asked for one during procurement.
Scoping. The company is a developer of a fine-tuned system and a deployer of it; it does not train frontier models. It states that boundary at the top of the report and attributes upstream properties, such as base-model pre-training data and base-model evaluations, to the model provider's published documentation.
Register. It maps 14 controls to the 11 actions. Actions 1 to 6 and 11 have real controls: a pre-release evaluation suite run on every fine-tune, abuse monitoring on the assistant, a published model card, a disclosure policy, an AI policy approved by a governance committee, access controls on adapter weights, and a training-data review that strips personal data from customer documents used for fine-tuning. Action 7 is partial: the product labels generated text in the interface but does not watermark it. Actions 8, 9 and 10 have one control each, honestly small.
Gap run. The checker reports three findings: the evaluation evidence for the current fine-tune is from the previous release (stale), the insider-risk control has no owner after a reorganisation, and the abuse-monitoring control has only internal dashboards, nothing citable. The team re-runs the evaluation, assigns an owner, and writes a public summary of abuse categories and response times without detection thresholds.
Report. The generated draft is edited by the governance lead, checked against the model card so the two documents make the same claims, reviewed by security for over-disclosure, and submitted. The same register later answers the EU customer's AI Act questionnaire with no new evidence gathering. Total effort was about three engineer-weeks, most of it closing the gaps the checker found.
Failure modes
Aspirational reporting. Statements such as "we are committed to safety" with no artefact behind them. Reviewers and journalists compare reports side by side, and vague answers stand out. Every sentence should point to something that exists.
Stale evidence. Citing an evaluation of a model version that is no longer served. Tie evidence to model versions and enforce a maximum age.
Inconsistent documents. The HAIP report says one thing about intended use, the model card another, the terms of service a third. Generate shared facts from one source and diff the documents before publication.
Over-disclosure. Publishing filter thresholds, monitoring blind spots or the architecture of weight storage. Describe that controls exist and how they are tested; keep parameters internal.
Overclaiming provenance. Implying that watermarking makes generated content reliably detectable. Text watermarks degrade under paraphrase and translation; see watermarking LLM outputs for the limits, and say what your mechanism does and does not survive.
Treating voluntary as optional forever. Customers increasingly ask for these reports in procurement, and binding regimes cover the same ground, so skipping the evidence work defers it rather than avoiding it.
Trade-offs
Transparency against security. Action 3 asks for public detail; action 6 asks you to protect the system. Resolve it per artefact with the public flag and a security review, not with a blanket rule.
Voluntary against binding. A voluntary code lets you report gaps honestly without legal exposure, which is its main strength. It also means no one verifies the claims, so its credibility depends on how specific you are.
Breadth against depth. Eleven actions invite thin coverage of all of them. It is better to show strong evidence for the actions that match your actual risk, typically 1, 2, 3, 6 and 11, and state plainly where you are small.
Cost. The first report costs weeks; later ones cost days if the register is maintained. The cost is mostly evidence you should be producing anyway.
What to do next
- Download the official Code of Conduct text and the OECD reporting questionnaire, and record which version you are working against.
- Write a one-paragraph scope statement: developer, deployer or both, which systems, and what you inherit from upstream providers.
- Map at least one control with a named owner to each of the 11 actions, marking honest gaps rather than inventing coverage.
- Attach dated evidence to every control, flag what can be cited publicly, and set a maximum age per control.
- Run the gap checker in CI and fix stale, ownerless and uncitable controls before drafting anything.
- Generate the draft, reconcile it with your model cards (see model cards) and terms, and have security review it for over-disclosure.
- Reuse the same register for EU AI Act, NIST and ISO requests, and schedule an annual refresh tied to your release calendar.