Most engineers meet GPU regulation sideways: a procurement order stalls, a customer in another country cannot be onboarded to a cluster, or legal asks how many floating-point operations went into last quarter's model. The rules behind those moments are written in hardware and compute terms, which means engineers are the people best placed to answer them, if they understand what is being measured.

This article covers how United States export controls classify accelerators, where the China policy stood at the end of September 2026, how the EU AI Act and California's SB 53 use training compute as a trigger, and what systems a GPU team should build to answer those questions with evidence. It explains rules that change often and is not legal advice; confirm dates and thresholds with counsel before acting.

Advertisement

Three layers: the chip, access to the chip, the model

Regulation touches an AI stack in three places. The first is the physical chip: export controls decide who may receive an accelerator, or a server containing one, based on its measured capability and the destination. The second is access: even when the chip never moves, rules and licence conditions increasingly care about who uses it remotely, who owns the company that bought it, and what the compute is used for. The third is the model: AI laws such as the EU AI Act and California's Transparency in Frontier Artificial Intelligence Act (SB 53) use the amount of compute spent training a model as a trigger for obligations.

Three places regulation touches an AI compute stack1. The chipexport classification: TPP, density2. Access to the chipwho may use it, from where3. The modeltraining compute thresholdsECCN 3A090 / 4A090licences by destination and end userKYC, end-use, affiliates,remote and cloud access conditionsEU AI Act GPAI: 1e23, 1e25 FLOPCalifornia SB 53: 1e26 operationsYour evidence layerhardware inventory with ECCN per SKU, customer screening records, compute ledger per training runAnswers you must producecan we ship it, can they use it, must we reportRules that movedates, thresholds, suspensions, guidanceBuild the evidence layer once; re-map it to each rule as the rules change.
The three layers and the shared evidence layer beneath them. The rules on the right move; the evidence you collect should not have to.

How export controls measure an accelerator

Since the October 2023 update, the United States Export Administration Regulations classify advanced integrated circuits under Export Control Classification Number 3A090, and computers containing them under 4A090. Two metrics decide the classification.

Total processing performance (TPP) is defined as 2 times the dense MacTOPS times the bit length of the operation, aggregated over all processing units on the chip, and taking the maximum across the bit lengths the chip supports. A multiply-accumulate counts as two operations, so in practice TPP is the dense TOPS figure at a precision times its bit width. Dense matters: many headline figures assume structured sparsity and are double the dense rate.

Performance density (PD) is TPP divided by the applicable die area in square millimetres.

The thresholds from the October 2023 rule are:

  • 3A090.a: TPP of 4800 or more; or TPP of 1600 or more with performance density of 5.92 or more.
  • 3A090.b: TPP of 2400 or more but below 4800 with performance density of 1.6 or more but below 5.92; or TPP of 1600 or more with performance density of 3.2 or more but below 5.92.

Classification is only the first step: licensing then depends on destination, end user, end use and policies that change far more often than the thresholds.

Classifying a data-centre accelerator under 3A090 (October 2023 thresholds)spec sheetdense TOPS per precision, die areaTPPmax over bit lengths of TOPS x bitsperformance densityTPP / applicable die area3A090.aTPP >= 4800, or TPP >= 1600 and PD >= 5.923A090.bTPP 2400-4800 and PD 1.6-5.92, or TPP >= 1600 and PD 3.2-5.92neithernot 3A090 on these tests; check other entriesThen: destination, end user, end use and any case-by-case policy (such as TPP below 21,000 for China, Jan 2026).
From spec sheet to classification. The arithmetic is simple; the hard part is getting dense per-precision rates and the right die area.
from dataclasses import dataclass

@dataclass
class Accelerator:
    name: str
    dense_tops: dict        # {bit_length: dense tera-operations per second, MAC = 2 ops}
    die_area_mm2: float     # applicable die area, per the regulation's definition

def tpp(acc: Accelerator) -> float:
    # 2 x MacTOPS x bit length == dense TOPS x bit length; take the max over precisions
    return max(tops * bits for bits, tops in acc.dense_tops.items())

def classify_3a090(acc: Accelerator) -> str:
    t = tpp(acc)
    pd = t / acc.die_area_mm2
    if t >= 4800 or (t >= 1600 and pd >= 5.92):
        return f"3A090.a (TPP {t:.0f}, PD {pd:.2f})"
    if (2400 <= t < 4800 and 1.6 <= pd < 5.92) or (t >= 1600 and 3.2 <= pd < 5.92):
        return f"3A090.b (TPP {t:.0f}, PD {pd:.2f})"
    return f"not 3A090 on TPP/PD tests (TPP {t:.0f}, PD {pd:.2f}); confirm with counsel"
Advertisement

Worked example: two hypothetical accelerators

Take a hypothetical training accelerator, call it Alpha, whose datasheet lists 1,000 dense TFLOPS at FP8 and 500 dense TFLOPS at BF16, on a die whose applicable area is 800 square millimetres. At FP8, TPP is 1,000 times 8, or 8,000. At BF16 it is 500 times 16, again 8,000. The maximum is 8,000, above 4,800, so Alpha is 3A090.a. Its performance density is 8,000 divided by 800, or 10. Using its sparse FP8 figure would double TPP to 16,000 and could move it across a policy line it does not cross.

Now an inference card, Beta: 125 dense TOPS at INT8 and 62.5 dense TFLOPS at FP16, on 800 square millimetres. TPP is 125 times 8, or 1,000 (FP16 gives the same). That is below 1,600, so neither 3A090 test applies. Density only matters once TPP reaches 1,600, so even a much smaller die would not change the answer.

For real parts, compute TPP from dense rates at every supported precision, and keep the arithmetic and datasheet version in your hardware inventory. For the architecture behind these numbers on real Hopper and Blackwell parts, see the H100 deep dive and the B200 deep dive.

Where the chip rules stood at the end of September 2026

Controls on advanced chips to China began in October 2022 and were tightened in October 2023. The AI Diffusion rule of January 2025, which would have tiered the world into access categories, was rescinded in May 2025. In April 2025 the China-specific H20 became subject to licensing; in December 2025 the administration said the H200 could be sold to approved Chinese customers, and a BIS rule effective 15 January 2026 implemented that narrowly.

ItemStatus at end of September 2026Engineering consequence
January 2026 case-by-case policyChips with TPP below 21,000 and total DRAM bandwidth below 6,500 GB/s (the rule was framed around NVIDIA H200 and AMD MI325X) move from presumption of denial to case-by-case review for China and Macau. Parts at or above either line stay under presumption of denial.Memory bandwidth is now a policy input alongside TPP; record it per SKU.
Conditions on those licencesUS availability and supply certification; China and Macau shipments at most 50% of US volume; testing by a US-headquartered lab on US soil; no military or WMD end use or listed parties; consignees must prevent unauthorised remote and IaaS access.Cloud operators must be able to show who accessed which hardware.
TariffA 25% tariff under Section 232 was announced on 14 January 2026 alongside the policy.Cost models for affected exports change.
Affiliates abroadBIS stated on 1 June 2026 that licence requirements apply to companies headquartered in, or with a parent in, China regardless of where the buyer is located.Screen on ownership, not just shipping address.
50% Affiliates RuleSuspended from 10 November 2025 until 9 November 2026 under the US-China trade arrangement; it reinstates automatically unless BIS acts.Ownership screening may need to tighten within weeks; plan for both outcomes.
Remote accessAugust 2026 reporting described access to advanced GPUs through Southeast Asian data centres; the regime focuses on physical chips, and proposals to cover remote access were under discussion.Expect access-based obligations; log them now.

None of this is stable enough to hard-code; express each rule as data, with thresholds, destinations, ownership tests and effective dates, that a screening service evaluates.

How AI laws measure a model: training compute

Model-side rules use cumulative training compute as a proxy for capability. Under the EU AI Act, the Commission's guidelines treat training compute above 10^23 floating-point operations, for a model that can generate language, images or video, as an indicative criterion for being a general-purpose AI model at all. A general-purpose model trained with 10^25 FLOP or more is presumed to have high-impact capabilities, making it a model with systemic risk; the provider must notify the Commission within two weeks of reasonably foreseeing that it will cross that line. Obligations for general-purpose model providers have applied since 2 August 2025, and the Commission's enforcement powers started on 2 August 2026, with fines of up to 3% of worldwide annual turnover or 15 million euros, whichever is higher. See the AI safety frameworks deep dive for how these obligations map to NIST AI RMF and ISO/IEC 42001.

California's SB 53, operative since 1 January 2026, defines a frontier model as a foundation model trained with more than 10^26 integer or floating-point operations, counting the initial run plus later fine-tuning or material modification. A large frontier developer is one whose revenue with affiliates exceeded 500 million US dollars in the preceding year, and it carries heavier duties such as publishing a safety framework.

You can estimate training compute two ways and should do both. From first principles, a dense transformer costs about 6 times parameters times training tokens. From the hardware, compute equals GPU-hours times 3,600 times the per-GPU peak rate times the achieved model FLOPs utilisation.

def flops_from_model(params: float, tokens: float) -> float:
    return 6.0 * params * tokens            # dense transformer rule of thumb

def flops_from_cluster(gpu_hours: float, peak_flops: float, mfu: float) -> float:
    return gpu_hours * 3600 * peak_flops * mfu

THRESHOLDS = {"EU GPAI indicative": 1e23, "EU systemic-risk presumption": 1e25,
              "California SB 53 frontier": 1e26}

def report(run_id: str, cumulative_flops: float) -> None:
    for name, limit in THRESHOLDS.items():
        share = cumulative_flops / limit
        flag = "CROSSED" if share >= 1 else ("WATCH" if share >= 0.5 else "ok")
        print(f"{run_id}: {name}: {share:.2%} of threshold [{flag}]")

# 70B dense model on 15T tokens
report("base-70b", flops_from_model(70e9, 15e12))        # 6.3e24 FLOP
# fine-tune: 1,000 GPUs x 30 days, 1e15 FLOP/s peak, 40% MFU
report("ft-run-7", flops_from_cluster(1000 * 24 * 30, 1e15, 0.40))  # ~1.0e24 FLOP

A 70-billion-parameter model on 15 trillion tokens is about 6.3 times 10^24 FLOP: above the EU indicative criterion, below the systemic-risk presumption. The fine-tune adds about 1.0 times 10^24, which SB 53 counts cumulatively. A 400-billion-parameter model on the same data would be near 3.6 times 10^25, inside the EU presumption; learn that at planning, not after the run.

The compliance engineering: an evidence layer

  • Hardware inventory with classification. Per SKU: dense rates, die area, TPP, PD, memory bandwidth, confirmed ECCN and datasheet version. Per unit: serial, site, country, owning entity. GPU infrastructure planning covers the capacity side of the same inventory.
  • Customer and tenant screening. Record legal entity, ultimate parent, headquarters country and dated restricted-party screening results; re-screen when lists or rules change.
  • Access logs tied to hardware. Map every job to tenant, user, node serials and region. Specialist providers such as those described in the CoreWeave deep dive already expose the per-tenant accounting this needs.
  • A compute ledger. Every training job appends its lineage, GPU-hours, hardware, utilisation and both FLOP estimates.
  • Rules as configuration. A versioned, counsel-reviewed file that the services read.
-- compute ledger: one row per job, lineage links fine-tunes to their base
CREATE TABLE compute_ledger (
  job_id          TEXT PRIMARY KEY,
  model_lineage   TEXT NOT NULL,       -- e.g. 'base-70b'
  parent_job_id   TEXT,                -- checkpoint this job started from
  gpu_type        TEXT NOT NULL,
  gpu_hours       DOUBLE PRECISION NOT NULL,
  measured_mfu    DOUBLE PRECISION,
  flops_model_est DOUBLE PRECISION,    -- 6 * N * D
  flops_hw_est    DOUBLE PRECISION,    -- hours * 3600 * peak * mfu
  region          TEXT NOT NULL,
  tenant_entity   TEXT NOT NULL,
  started_at      TIMESTAMPTZ NOT NULL,
  ended_at        TIMESTAMPTZ
);

SELECT model_lineage,
       SUM(GREATEST(flops_model_est, flops_hw_est)) AS cumulative_flops
FROM compute_ledger GROUP BY model_lineage ORDER BY cumulative_flops DESC;

Failure modes

FailureHow it happensPrevention
Sparse rates in TPPMarketing TFLOPS with 2:4 sparsity doubles the figureStore dense rates only; cite datasheet page
Wrong precisionTPP computed at BF16 when FP8 or INT8 gives a higher valueCompute every supported bit length and take the max
Screening by addressA buyer abroad is owned by a restricted parentRecord ultimate parent and headquarters; re-screen on rule changes
Uncounted fine-tuningLineage lost when a checkpoint is copied between teamsLedger rows carry parent_job_id; block untracked checkpoints
Hard-coded rulesThresholds buried in code go stale after a rule changeRules as reviewed configuration with effective dates
Stale snapshotA policy summary from last quarter is treated as currentDate-stamp every summary; review on each BIS or AI Office publication

What to do next

  1. Build a hardware inventory with dense rates per precision, die area source, computed TPP and PD, memory bandwidth and confirmed ECCN for every SKU you own or plan to buy.
  2. Run the classification code on your fleet and have counsel confirm the results for anything near a threshold.
  3. Record ultimate parent and headquarters country for every customer or tenant, and schedule re-screening before 9 November 2026.
  4. Tie every job to tenant, user, node serials and region, and set a retention period with counsel.
  5. Create a compute ledger with model lineage, and back-fill it for every model you still ship.
  6. Add a planning check that estimates 6ND for each proposed run against 10^23, 10^25 and 10^26 before it is scheduled.
  7. Move thresholds and dates into a versioned, counsel-reviewed rules file consumed by screening and ledger services.
  8. Review BIS rules and EU AI Office guidance on a calendar, and date-stamp every internal summary.
Key takeaway: Export controls measure an accelerator by total processing performance and performance density, and since January 2026 also by memory bandwidth for China case-by-case review; AI laws measure a model by cumulative training compute, at 10^23 and 10^25 FLOP in the EU and 10^26 operations in California. Both are engineering quantities. Compute them correctly from dense rates and real job logs, keep them in an inventory and a compute ledger alongside customer screening and access records, and express the fast-moving rules as reviewed configuration, so each change in policy becomes a lookup rather than an investigation.