The G7 Hiroshima AI Process is the set of agreements the G7 started under Japan's presidency in May 2023 to give governments and AI developers a shared baseline for advanced AI systems, chiefly foundation models and generative AI. It is voluntary. No regulator fines anyone for ignoring it. Yet it matters to engineering teams for a practical reason: its Code of Conduct is the template behind a public reporting framework run by the OECD, and its eleven actions overlap heavily with what the EU AI Act, the NIST AI Risk Management Framework and ISO/IEC 42001 ask you to show. If you build the evidence once, you can answer all of them.

This article treats the Process as an engineering input rather than a diplomatic event. It explains what was actually agreed and when, reads each of the eleven actions as a control with an owner and a piece of evidence, shows how the OECD reporting framework turns that evidence into a public report, and gives a small controls register with a gap checker you can adapt. Country-level policy in Japan, which led the Process, is covered separately in Japan AI policy in depth.

What the Process actually produced

Three documents and one mechanism make up the Process. Keeping them apart avoids the most common confusion, which is treating the guiding principles, the code and the report as one thing.

WhenWhatWho it addressesWhat it asks for
May 2023Process launched at the G7 Hiroshima summitG7 governmentsWork towards common ground on generative AI governance
30 Oct 2023International Guiding Principles for Organizations Developing Advanced AI SystemsDevelopers first, other AI actors as appropriateEleven high-level principles
30 Oct 2023International Code of Conduct for Organizations Developing Advanced AI SystemsOrganisations developing advanced AIThe same eleven items, written as actions with detail
Dec 2023Comprehensive Policy Framework, endorsed by G7 digital and tech ministersGovernments and all AI actorsBundles an OECD report on generative AI, guiding principles extended to all AI actors, the Code, and project-based cooperation
Feb 2025OECD Hiroshima AI Process Reporting Framework launchedDevelopers, deployers and providers that volunteerA public questionnaire answered against the Code's actions

The first reporting round produced 19 published reports in spring 2025 from organisations including Anthropic, Fujitsu, Google, KDDI, Microsoft, NEC, NTT, OpenAI, Preferred Networks, Rakuten, Salesforce and SoftBank, alongside smaller firms and research bodies. The OECD accepts submissions on a rolling basis and encourages annual updates, and later OECD analysis covered roughly two dozen reports as more arrived. Treat any exact count as dated: the portal is the source of truth.

Two properties shape how you should use it. First, the Code is principles-based and risk-based: it says what outcome to achieve, not which technique to use, so the interpretation work is yours. Second, the report is public. Anything you write is read by customers, journalists, competitors and attackers, which changes how much security detail you can include.

The eleven actions as controls and evidence

The eleven actions below are paraphrased from the Code; read the official text before quoting it. The third column is the useful part for engineers: the artefact that proves the action is happening rather than intended.

#Action (paraphrased)Evidence that shows it is real
1Identify, evaluate and mitigate risks across the AI lifecycle, before and during deploymentPre-release evaluation plan, red-team results, release gate record
2Identify and mitigate vulnerabilities, incidents and patterns of misuse after deploymentAbuse monitoring dashboards, vulnerability intake, incident tickets
3Publicly report capabilities, limitations and appropriate and inappropriate usesModel cards or system cards per release, acceptable use policy
4Share information responsibly and report incidents with industry, government, civil society and academiaDisclosure policy, sharing agreements, incident notifications sent
5Develop, implement and disclose risk-based governance and risk management policies, including privacyAI policy, risk register, governance committee minutes
6Invest in robust security controls, including physical, cyber and insider-threat safeguardsWeights access controls, audit logs, insider-risk programme, pen tests
7Deploy reliable content authentication and provenance where technically feasible, such as watermarkingProvenance metadata spec, watermark evaluation, user-facing labels
8Prioritise research to mitigate societal, safety and security risks and invest in mitigationsResearch budget lines, publications, funded mitigation work
9Prioritise developing advanced AI to address global challenges such as climate, health and educationProgrammes and deployments with measured outcomes
10Advance the development and, where appropriate, adoption of international technical standardsStandards participation, adopted standards list
11Implement appropriate data input measures and protections for personal data and intellectual propertyData provenance records, opt-out handling, PII filtering, licence review

Grouped by who does the work, the actions fall into four clusters. Actions 1, 2 and 6 belong to engineering and security: evaluations, monitoring and the protection of weights and infrastructure. Actions 3, 4 and 7 are about transparency to outsiders: documentation, incident sharing and labelling generated content. Actions 5, 10 and 11 are governance and data management. Actions 8 and 9 are strategic investment choices, and they are where reports most often slide into marketing. A team that only deploys a third-party model can answer most of 2, 3, 5, 6 and 11 from its own operations and should say plainly that 1, 7 and 8 depend partly on the upstream provider.

Architecture: one evidence store, several frameworks

The structure that makes this tractable is the same one used for any compliance regime: a controls register keyed by requirement, an evidence store holding dated artefacts, and a crosswalk that lets one artefact answer several frameworks. The Hiroshima Code becomes one more column in that crosswalk rather than a separate project.

From a voluntary code to evidence you can reuseCode of Conduct11 actions, Oct 2023Controls registerowner + control per actionEvidence storedated artefacts, linksinterpretproducequeryGap checkermissing / stale / thinRemediation backlogfix before you reportcloses gapsHAIP reportOECD portal, yearlyEU AI ActGPAI obligationsNIST AI RMFprofile, MAP / MEASUREISO/IEC 42001AI management systemOne evidence store, several outputs: the crosswalk decides which artefact answers which question.
Each Code action maps to controls with owners; controls produce dated evidence; a gap checker flags missing or stale evidence before any report is drafted.

The crosswalk below is our own mapping, not an official one. It is approximate because the frameworks are written at different levels: the EU AI Act imposes legal obligations on providers of general-purpose AI models, NIST's framework organises risk activities into GOVERN, MAP, MEASURE and MANAGE functions, and ISO/IEC 42001 specifies an auditable management system.

Code actionEU AI ActNIST AI RMFISO/IEC 42001
1 Lifecycle riskSystemic-risk evaluation and mitigationMAP, MEASUREAI risk assessment and treatment
2 Post-deployment misuseSerious incident tracking for systemic-risk modelsMANAGEOperation and monitoring
3 Public reportingTechnical documentation and downstream informationGOVERN, MAPInformation for interested parties
6 SecurityCybersecurity protection for systemic-risk modelsMANAGESecurity controls via the management system
7 ProvenanceGenerated-content transparency duties on AI system providersMEASUREPartly, through system design controls
11 Data and IPCopyright policy and training-content summaryMAPData management for AI systems

Deeper treatment of each framework lives in the EU AI Act in depth, the NIST AI RMF and ISO/IEC 42001.

A controls register and gap checker in code

A spreadsheet works for ten controls; past that, a small typed register pays for itself because the gap check runs in CI and the report draft is generated from the same data the auditors see. The sketch below is deliberately plain Python with no dependencies.

from dataclasses import dataclass, field
from datetime import date, timedelta

@dataclass
class Evidence:
    title: str
    url: str            # internal link; never paste secrets into a public report
    produced: date
    public: bool        # may this artefact be cited in the HAIP report?

@dataclass
class Control:
    action: str         # Code action number, "1" .. "11"
    name: str
    owner: str
    max_age_days: int = 365
    evidence: list = field(default_factory=list)

def gaps(controls, today=None):
    today = today or date.today()
    covered = {c.action for c in controls}
    out = [("action %s" % a, "no control mapped") for a in map(str, range(1, 12))
           if a not in covered]
    for c in controls:
        if not c.owner:
            out.append((c.name, "no owner"))
        if not c.evidence:
            out.append((c.name, "no evidence"))
            continue
        newest = max(e.produced for e in c.evidence)
        if today - newest > timedelta(days=c.max_age_days):
            out.append((c.name, "stale: newest evidence %s" % newest))
        if not any(e.public for e in c.evidence):
            out.append((c.name, "nothing citable publicly"))
    return out

def draft_report(controls, action_text):
    """One section per action: what we do, citable evidence, known gaps."""
    lines = []
    for a in map(str, range(1, 12)):
        lines.append("## Action %s: %s" % (a, action_text[a]))
        mine = [c for c in controls if c.action == a]
        if not mine:
            lines.append("Not yet addressed. Planned owner and date: TODO.")
        for c in mine:
            cites = [e.title for e in c.evidence if e.public]
            lines.append("- %s (owner: %s). Evidence: %s"
                         % (c.name, c.owner, "; ".join(cites) or "internal only"))
    return "\n".join(lines)

Three design choices matter. Evidence carries a public flag, because the same red-team result that proves action 1 internally may be unsafe to describe in a public report; the gap checker flags controls with nothing citable so you decide early what to summarise. Every control has a maximum evidence age, which catches the classic failure of reporting last year's evaluation for this year's model. And the draft writes Not yet addressed for unmapped actions instead of skipping them, because an honest gap reads better than an omission that a reviewer will find.

Worked example: a first report from a mid-size company

Consider a 300-person software company that fine-tunes an open-weight model and ships a writing assistant inside its product in Europe and Japan. It wants to file a HAIP report because an enterprise customer asked for one during procurement.

Scoping. The company is a developer of a fine-tuned system and a deployer of it; it does not train frontier models. It states that boundary at the top of the report and attributes upstream properties, such as base-model pre-training data and base-model evaluations, to the model provider's published documentation.

Register. It maps 14 controls to the 11 actions. Actions 1 to 6 and 11 have real controls: a pre-release evaluation suite run on every fine-tune, abuse monitoring on the assistant, a published model card, a disclosure policy, an AI policy approved by a governance committee, access controls on adapter weights, and a training-data review that strips personal data from customer documents used for fine-tuning. Action 7 is partial: the product labels generated text in the interface but does not watermark it. Actions 8, 9 and 10 have one control each, honestly small.

Gap run. The checker reports three findings: the evaluation evidence for the current fine-tune is from the previous release (stale), the insider-risk control has no owner after a reorganisation, and the abuse-monitoring control has only internal dashboards, nothing citable. The team re-runs the evaluation, assigns an owner, and writes a public summary of abuse categories and response times without detection thresholds.

Report. The generated draft is edited by the governance lead, checked against the model card so the two documents make the same claims, reviewed by security for over-disclosure, and submitted. The same register later answers the EU customer's AI Act questionnaire with no new evidence gathering. Total effort was about three engineer-weeks, most of it closing the gaps the checker found.

Failure modes

Aspirational reporting. Statements such as "we are committed to safety" with no artefact behind them. Reviewers and journalists compare reports side by side, and vague answers stand out. Every sentence should point to something that exists.

Stale evidence. Citing an evaluation of a model version that is no longer served. Tie evidence to model versions and enforce a maximum age.

Inconsistent documents. The HAIP report says one thing about intended use, the model card another, the terms of service a third. Generate shared facts from one source and diff the documents before publication.

Over-disclosure. Publishing filter thresholds, monitoring blind spots or the architecture of weight storage. Describe that controls exist and how they are tested; keep parameters internal.

Overclaiming provenance. Implying that watermarking makes generated content reliably detectable. Text watermarks degrade under paraphrase and translation; see watermarking LLM outputs for the limits, and say what your mechanism does and does not survive.

Treating voluntary as optional forever. Customers increasingly ask for these reports in procurement, and binding regimes cover the same ground, so skipping the evidence work defers it rather than avoiding it.

Trade-offs

Transparency against security. Action 3 asks for public detail; action 6 asks you to protect the system. Resolve it per artefact with the public flag and a security review, not with a blanket rule.

Voluntary against binding. A voluntary code lets you report gaps honestly without legal exposure, which is its main strength. It also means no one verifies the claims, so its credibility depends on how specific you are.

Breadth against depth. Eleven actions invite thin coverage of all of them. It is better to show strong evidence for the actions that match your actual risk, typically 1, 2, 3, 6 and 11, and state plainly where you are small.

Cost. The first report costs weeks; later ones cost days if the register is maintained. The cost is mostly evidence you should be producing anyway.

What to do next

  1. Download the official Code of Conduct text and the OECD reporting questionnaire, and record which version you are working against.
  2. Write a one-paragraph scope statement: developer, deployer or both, which systems, and what you inherit from upstream providers.
  3. Map at least one control with a named owner to each of the 11 actions, marking honest gaps rather than inventing coverage.
  4. Attach dated evidence to every control, flag what can be cited publicly, and set a maximum age per control.
  5. Run the gap checker in CI and fix stale, ownerless and uncitable controls before drafting anything.
  6. Generate the draft, reconcile it with your model cards (see model cards) and terms, and have security review it for over-disclosure.
  7. Reuse the same register for EU AI Act, NIST and ISO requests, and schedule an annual refresh tied to your release calendar.
Key takeaway: The Hiroshima AI Process is voluntary, but its Code of Conduct and the OECD reporting framework built on it describe the evidence every serious AI governance regime now asks for. Treat the eleven actions as controls with owners and dated artefacts, check for gaps before you write, report specifically and honestly, and reuse the same evidence for binding frameworks.