Singapore is a useful case study for anyone building LLM systems, even if you never ship there. It has not passed a horizontal AI act like the EU's. Instead it pairs binding law (the Personal Data Protection Act and a few harm-specific statutes) with unusually concrete voluntary frameworks and open-source testing tools. For engineers, Singapore governance arrives as questions: who is accountable, what was tested, how is personal data used, what happens when the agent errs. Buyers and auditors expect evidence-backed answers.

This article explains which layers bind, maps each framework to engineering controls, and works through a refund assistant deployed by a Singapore company. Dates come from official IMDA, PDPC and Ministry of Law announcements; the rest is a summary, so read the primary documents before relying on it for compliance.

The governance stack: what binds and what guides

Separate what can get you fined from what shapes expectations. The binding layer is the PDPA (any personal data your system collects, uses or discloses) plus targeted harm laws. On top sits a voluntary layer from the Infocomm Media Development Authority (IMDA) and the AI Verify Foundation it set up. Voluntary is not ignorable: banks, insurers and agencies use these frameworks as procurement checklists.

Singapore's AI governance stack: binding law at the bottom, voluntary frameworks and tools on topTesting and assurance toolsAI Verify testing framework (11 principles), Project Moonshot for LLM red teaming and benchmarksModel AI Gov. Framework2019, 2nd edition 2020MGF for Generative AI2024, nine dimensionsMGF for Agentic AIJan 2026, updated May 2026voluntary: no penalties, but they set what a regulator, auditor or customer will ask forSector rules and principlese.g. MAS FEAT principles for financePDPC advisory guidelinespersonal data in AI recommendation and decision systems, 2024PDPAconsent, purpose, protection, breachElections deepfake rulesamendment passed Oct 2024Online Safety (R&A) Actcommission operating from Jun 2026binding: harm-specific and data-specific statutes, not a horizontal AI act
The layers, simplified. Red boxes are binding law; blue and purple boxes are voluntary guidance and tooling that buyers and auditors use as a yardstick.
InstrumentWhenBinding?
Model AI Governance Framework2019, second edition 2020No
Advisory Guidelines on personal data in AI recommendation and decision systems1 March 2024Guidance on a binding law
Model AI Governance Framework for Generative AI2024No
AI Verify testing framework and Project Moonshot2022 onward; Moonshot 2024No
Elections (Integrity of Online Advertising) AmendmentPassed 15 October 2024Yes
Model AI Governance Framework for Agentic AI22 January 2026, updated 20 May 2026No
Online Safety (Relief and Accountability) Act 2025Commission operating from 29 June 2026Yes

The Model AI Governance Framework

The original framework is the backbone the later documents assume. Its four areas each become a design decision.

  • Internal governance structures and measures. Name an owner for each AI system, with a risk-review step before launch. In code terms: a system card and an approval record attached to every deployment.
  • Determining the level of human involvement. The framework distinguishes human-in-the-loop (a person approves each decision), human-over-the-loop (a person supervises and can intervene) and human-out-of-the-loop (fully automated). You choose by weighing the probability and severity of harm. That is an architecture choice: where the approval queue sits, and which actions can bypass it.
  • Operations management. Data lineage, model selection, robustness testing, monitoring and retraining discipline. The usual MLOps hygiene, written down.
  • Stakeholder interaction and communication. Telling users that AI is involved, explaining decisions where it matters, and giving them a feedback or appeal channel.

Teams often skip the probability-severity choice. A reversible SGD 50 refund can run out of the loop; account closure keeps a human in it. Record why.

Generative AI: nine dimensions as controls

The 2024 framework for generative AI keeps the same spirit and widens the scope from one organisation to the whole supply chain: model developers, application builders and cloud providers. It names nine dimensions. The useful exercise is to turn each into a control you own and a piece of evidence you can produce.

DimensionControl you can buildEvidence
AccountabilityRACI across model vendor, your team and any integrator; contract clauses on incidentsSigned responsibility matrix
DataTraining and retrieval data inventory with legal basis and licence per sourceData register, PDPA basis per field
Trusted development and deploymentModel and system cards, baseline safety evals, staged rolloutCards versioned with each release
Incident reportingSeverity scale, on-call, user report button, post-incident regression testsIncident log with time to fix
Testing and assurancePre-release eval suite, periodic external or third-party testingEval reports tied to model hash
SecurityPrompt-injection tests, tool sandboxing, secrets isolation, supply-chain checksRed-team findings and fixes
Content provenanceLabel AI output; watermark or sign media you generate where tools allowProvenance policy, sample outputs

The remaining two, safety and alignment R&D and AI for public good, mostly concern model developers and governments rather than application teams.

AI Verify, Moonshot and testing as evidence

AI Verify is a testing framework and open-source toolkit. It assesses a system against 11 governance principles: transparency, explainability, repeatability or reproducibility, safety, security, robustness, fairness, data governance, accountability, human agency and oversight, and inclusive growth with societal and environmental well-being. Each principle has desired outcomes, each outcome has process checks validated by documentary evidence, and a subset (fairness, explainability, robustness) also has technical tests. The original toolkit targeted classical supervised models; the framework has since been extended with generative AI considerations.

Project Moonshot, launched in 2024, is the LLM-side tool: it connects to a model endpoint, runs benchmark suites and supports red-teaming sessions. You do not need to adopt either tool to follow the approach, but you do need its shape: a repeatable test run, tied to an exact model and prompt version, with results stored as evidence. A minimal harness looks like this.

import hashlib, json, time

def run_suite(model_call, suite, system_prompt, model_id):
    """Run a fixed test suite and emit an evidence record tied to exact versions."""
    results = []
    for case in suite:                      # case: {"id", "principle", "prompt", "check"}
        out = model_call(system_prompt, case["prompt"])
        results.append({"id": case["id"], "principle": case["principle"],
                        "passed": case["check"](out)})
    by_principle = {}
    for r in results:
        agg = by_principle.setdefault(r["principle"], [0, 0])
        agg[0] += r["passed"]; agg[1] += 1
    return {
        "model_id": model_id,
        "prompt_sha256": hashlib.sha256(system_prompt.encode()).hexdigest(),
        "suite_sha256": hashlib.sha256(json.dumps([c["id"] for c in suite]).encode()).hexdigest(),
        "run_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
        "pass_rate": {k: round(p / n, 3) for k, (p, n) in by_principle.items()},
    }

Store that record with the release: a hash-pinned file, not a slide.

Turning frameworks into evidence: one control register feeds tests, logs and the assurance packFramework clausesMGF, GenAI, AgenticControl registercontrol id, owner, testCI test suitesafety, robustness, biasRuntime telemetryguardrail and tool logsAssurance packper releaseCustomer, auditor, regulatorasks: show me the evidenceIncident processtriage, notify, fix, add regression testThe register is the single source: every clause maps to a control, every control to a test or a log query.
A control register maps every framework clause to a test or log query, and each release produces an assurance pack from those sources.

Personal data: the PDPA and PDPC guidance

The PDPA is binding, and it applies whenever personal data flows through your system: in training data, in retrieval corpora, in prompts and in logs. The PDPC's March 2024 advisory guidelines explain how the commission applies the Act to AI recommendation and decision systems. Note the scope: they address systems that recommend or decide, and they explicitly do not cover generative AI. For an LLM product you still apply the PDPA itself, and the guidelines are the best available signal of how the PDPC reasons.

The guidelines centre on three questions. Consent or exception: model development may rely on the business improvement or research exceptions instead of fresh consent, within their conditions. Protection during development: minimise, pseudonymise or anonymise where the use case allows. Transparency: when consent is the basis, tell people enough for it to mean something. Encode the decision per dataset.

DATASET_REGISTER = {
    "support_tickets_2025": {
        "contains_personal_data": True,
        "pdpa_basis": "business_improvement_exception",   # or "consent", "research_exception"
        "basis_review": "DPO-2026-014",                   # who decided, where it is written down
        "minimisation": ["names->tokens", "NRIC regex removed", "phone masked"],
        "allowed_uses": ["fine_tune_router", "eval_set"],
        "retention_days": 365,
    },
}

def check_use(dataset, use):
    entry = DATASET_REGISTER[dataset]
    if use not in entry["allowed_uses"]:
        raise PermissionError(f"{dataset} is not cleared for {use}; see {entry['basis_review']}")

This stops debug logs quietly becoming fine-tuning data: unlisted uses are refused.

Agentic AI: bounding what agents can do

The agentic framework, launched at Davos on 22 January 2026 and updated on 20 May 2026, addresses systems that plan and act on a user's behalf. It is organised around four dimensions: assessing and bounding the risks, ensuring meaningful human accountability, implementing technical controls and processes, and enabling end-user responsibility. The May update added case studies and practices covering multi-agent systems, third-party agents and automation bias.

"Bounding the risk" is the dimension that turns most directly into code. An agent's risk is set by what its tools can do, so the bound is a policy check between the model's proposed action and its execution.

POLICY = {
    "lookup_order":   {"auto": True},
    "issue_refund":   {"auto_max_sgd": 50, "approver": "support_lead", "daily_cap_sgd": 2000},
    "close_account":  {"auto": False, "approver": "ops_manager"},
}

def gate(action, args, ctx):
    rule = POLICY.get(action)
    if rule is None:
        return "deny", "tool not in allow-list"
    if rule.get("auto"):
        return "allow", None
    if action == "issue_refund":
        if ctx.refunded_today_sgd + args["amount"] > rule["daily_cap_sgd"]:
            return "deny", "daily cap reached"
        if args["amount"] <= rule["auto_max_sgd"]:
            return "allow", None
    return "escalate", rule["approver"]   # a named human role owns the decision

Human accountability means escalations land on a named role with context, not a queue nobody owns. Automation bias is approvers rubber-stamping: measure approval rate and time-to-approve. End-user responsibility means telling users what the agent can do and how to stop or reverse it.

Targeted harm laws that reach AI output

Two binding laws reach generative output directly. The Elections (Integrity of Online Advertising) Amendment, passed on 15 October 2024, prohibits online election advertising that realistically depicts a candidate saying or doing something they did not, whether made with AI or ordinary editing. If your product generates images, video or voice of real people, an election-period policy belongs in your content filters.

The Online Safety (Relief and Accountability) Act 2025 created an Online Safety Commission, operating since 29 June 2026, whose first phase covers intimate image abuse, image-based child abuse, doxxing, online harassment and online stalking. For a product hosting user or model content, those map to classifier labels and a takedown workflow.

Worked example: a refund assistant

A Singapore retailer launches an LLM assistant that answers order questions and can issue refunds. It uses a third-party model through an API, retrieves from a help-centre corpus, and logs conversations. Here is the governance work, in the order a sensible team does it.

  1. Risk tiering (MGF). Order lookups: low severity, out of the loop. Refunds up to SGD 50: reversible, out of the loop with a daily cap. Larger refunds: human over the loop through the approval queue. Account closure: human in the loop.
  2. Data (PDPA). The retrieval corpus has no personal data. Conversation logs do. The DPO records the business improvement exception as the basis for using masked logs to build an eval set, and bans their use for vendor fine-tuning. Retention: one year.
  3. Testing (GenAI framework, AI Verify approach). 400 test cases across safety, robustness, prompt injection through order notes, and refund-policy adherence. Release gate: zero unauthorised refunds in 200 adversarial attempts and at least 95 percent correct answers on the factual set.
  4. Agent bounding (Agentic framework). The policy gate above, plus idempotency keys on the refund API so a retried tool call cannot pay twice.
  5. Incidents. A severity scale where any unauthorised refund is Sev-2, a user-facing report button, and a rule that every incident adds a regression case.
  6. Communication. The UI says it is an AI assistant, lists what it can do, and offers a human handoff.

Three months in, the queue shows 1,240 escalations, 1,236 approved at a median of 9 seconds: automation bias. The team raises the auto limit to SGD 80, where disputes are rare, and requires a written reason above SGD 300.

Failure modes

Failure modeWhat it looks likePrevention
Treating voluntary as optionalNo evidence when an enterprise buyer sends an AI Verify-style questionnaireKeep a control register from day one
Evidence not tied to versionsEval report cannot be matched to the deployed model or promptHash model id, prompt and suite in every record
Decorative human reviewNear-100 percent approvals in secondsMonitor approval rate and latency; redesign thresholds
Ignoring targeted harm lawsImage tool generates realistic candidate depictions during an electionElection-period filters, takedown runbook

Trade-offs and related reading

The voluntary approach is fast to update, as the four months between the agentic framework and its revision show. The cost is ambiguity: no conformity assessment says you are done, so the bar is your most demanding customer. Evidence in the AI Verify shape travels well; one control register can map to the EU AI Act, NIST AI RMF and ISO/IEC 42001. Heavier human review cuts harm but adds latency and, past a point, only paperwork. External testing builds trust but exposes prompts and tool designs.

For related reading, compare the binding approach in the EU AI Act, the risk vocabulary in the NIST AI RMF, the management-system view in ISO/IEC 42001, how national testing bodies work in AI safety institutes, and a neighbouring soft-law regime in Japan's AI governance.

What to do next

  1. List every AI system you run in or for Singapore, with an owner and a probability-severity tier that sets its level of human involvement.
  2. Build a dataset register with a PDPA basis, minimisation steps and allowed uses per dataset, and make pipelines enforce it.
  3. Map the nine GenAI dimensions to controls and evidence in one control register.
  4. Stand up a hash-pinned eval suite and run it on every model, prompt or tool change.
  5. For agents, put a policy gate between proposed actions and execution, with caps, idempotency and named approvers; watch approval rates for automation bias.
  6. Add election-period and online-harm categories to content filters and takedown runbooks.
  7. Read the current text of each framework before relying on this summary; they are revised often.
Key takeaway: Singapore governs AI through binding data and harm laws plus detailed voluntary frameworks that buyers and auditors treat as the bar. Engineer for evidence: a control register, version-pinned tests, a dataset register with a PDPA basis, policy gates around agent actions and measured human oversight.