A cloud security baseline is the short list of controls that every account, subscription or project in your organisation must satisfy, no matter who owns it or what runs in it. It is not a compliance framework and not a hardening guide for one workload. It is the floor: the settings whose absence has caused the most common cloud breaches, such as a public storage bucket, an unmonitored root credential, a deleted audit log or an instance metadata endpoint that hands credentials to anyone who can make the server fetch a URL.
Most organisations have a baseline document. Far fewer have a baseline that is enforced by the platform, checked continuously and measured. This article shows how to build the second kind: how to choose controls, which to prevent and which to detect, how to express them as code on AWS and Google Cloud, how exceptions work, and what breaks. It assumes you already have an account structure; the landing zone article covers how to build one.
What belongs in a baseline, and what does not
A control earns a place in the baseline when three things are true: it applies to every account regardless of workload, violating it creates a serious and well-understood risk, and it can be checked automatically. Encryption at rest by default qualifies. A rule that every API must use mutual TLS does not, because it depends on the workload and belongs in a service-level standard.
Keeping the list short is a feature. A baseline of fifteen controls that is 99 percent compliant protects more than a two-hundred-item checklist that nobody can measure. Published benchmarks such as the CIS Foundations Benchmarks for each major provider are a good source of candidates; treat them as a menu, pick the controls that meet the three tests, and write down why each one is in.
The control catalogue
Group controls by the failure they prevent. Each needs an identifier, a preventive mechanism where one exists, a detective check, and an owner. The table is a realistic starting set.
| Area | Control | Prevent with | Detect with |
|---|---|---|---|
| Identity | No long-lived root or owner credentials; MFA on privileged humans | Org policy; disable key creation | Credential reports, account summary |
| Identity | Humans federate through the identity provider; no local users | Deny user creation outside a pipeline role | Inventory of IAM users |
| Logging | Org-wide audit trail to a separate log account, tamper-resistant | Deny stopping or deleting the trail | Trail status check |
| Data | Storage cannot be made public by default | Account-level public access block; org constraint | Public resource findings |
| Data | Disks and databases encrypted at rest by default | Default encryption settings | Config rule per region |
| Compute | Instance metadata requires session tokens | Deny launches without IMDSv2 | Instance inventory |
| Network | No administrative ports open to the internet | Firewall policy at org or folder level | Security group analysis |
| Scope | Resources only in approved regions | Region deny policy | Resource inventory by region |
| Detection | Threat detection enabled in every region and account | Delegated admin auto-enable | Service status per account |
Identity design itself, roles, boundaries and federation, is covered in Cloud IAM architecture; keys and envelope encryption in Cloud KMS architecture. The baseline only asserts the minimum both must satisfy.
Preventive, detective and responsive: why you need all three
Preventive controls stop a violation from being created: an organisation policy that refuses the API call. They are the strongest, because there is nothing to clean up, but they are blunt, can break legitimate work, and can only express what the provider's policy language can express. Detective controls evaluate existing state and raise findings. They catch everything preventive controls cannot express, plus drift in accounts created before the control existed. Responsive controls act on findings, either automatically, such as re-enabling a public access block, or by opening a ticket with an owner and a deadline.
The rule of thumb: prevent what is unambiguous and cheap to get right, detect everything, and auto-remediate only when the fix cannot break a running workload. Turning on default disk encryption is safe to auto-apply to new volumes. Deleting a security group rule that opens port 22 might cut off a team's only access path, so it should page a person.
Preventive guardrails as code: AWS
On AWS, service control policies (SCPs) attached to the organisation root or to organisational units set the maximum permissions for every principal in the member accounts below them. The example below protects the audit trail, requires IMDSv2 on new instances and restricts activity to approved regions.
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "ProtectAuditTrail",
"Effect": "Deny",
"Action": ["cloudtrail:StopLogging", "cloudtrail:DeleteTrail", "cloudtrail:UpdateTrail"],
"Resource": "*"
},
{
"Sid": "RequireIMDSv2",
"Effect": "Deny",
"Action": "ec2:RunInstances",
"Resource": "arn:aws:ec2:*:*:instance/*",
"Condition": {"StringNotEquals": {"ec2:MetadataHttpTokens": "required"}}
},
{
"Sid": "ApprovedRegionsOnly",
"Effect": "Deny",
"NotAction": ["iam:*", "sts:*", "organizations:*", "support:*", "cloudfront:*", "route53:*"],
"Resource": "*",
"Condition": {"StringNotEquals": {"aws:RequestedRegion": ["eu-west-1", "eu-central-1"]}}
}
]
}Three details matter. The region statement uses NotAction to exempt global services whose API calls are made against a single region; without that exemption, the policy breaks identity, DNS and CDN management. The list shown is illustrative, not complete; start from AWS's own published region-restriction example SCP and extend its exemptions for the global services you use. SCPs only deny: they never grant, so a principal still needs an IAM policy that allows the action. And SCPs do not apply to the organisation's management account, which is one reason nothing should run in it. Test every change in a sandbox organisational unit first, because an SCP mistake takes effect across every account at once.
Preventive guardrails as code: Google Cloud
Google Cloud expresses the same idea with organisation policy constraints set on the organisation, a folder or a project and inherited downward. Two high-value constraints block service account key creation, which removes the most common source of leaked long-lived credentials, and enforce public access prevention on Cloud Storage.
# Organization policies (GCP): one policy per file, each applied with
# gcloud org-policies set-policy FILE.yaml
# Enforced at the organization node, inherited by every folder and project.
# disable_sa_keys.yaml
name: organizations/ORG_ID/policies/iam.disableServiceAccountKeyCreation
spec:
rules:
- enforce: true
# public_access_prevention.yaml
name: organizations/ORG_ID/policies/storage.publicAccessPrevention
spec:
rules:
- enforce: trueAzure's equivalent is Azure Policy assigned at a management group, with deny effects for prevention and audit effects for detection. Whatever the provider, keep these definitions in the same repository as the rest of your infrastructure code, review them like code, and deploy them through a pipeline rather than a console.
Detective checks as code
Providers ship managed detection for many baseline controls, such as AWS Config rules and Security Hub, Google Security Command Center, and Microsoft Defender for Cloud. Use them. But it is worth being able to run the baseline yourself, because it forces the control definitions to be precise and gives you a check that works identically in a pipeline, a new account and an incident. A minimal checker for a few AWS controls looks like this:
import boto3
def check_account(session, account_id, trail_name):
# Return a list of (control_id, passed, detail) for one AWS account.
results = []
summary = session.client("iam").get_account_summary()["SummaryMap"]
results.append(("IAM-1 root has no access keys", summary["AccountAccessKeysPresent"] == 0, summary["AccountAccessKeysPresent"]))
results.append(("IAM-2 root has MFA", summary["AccountMFAEnabled"] == 1, summary["AccountMFAEnabled"]))
for region in ("eu-west-1", "eu-central-1"):
ebs = session.client("ec2", region_name=region).get_ebs_encryption_by_default()
results.append((f"DATA-1 EBS default encryption {region}", ebs["EbsEncryptionByDefault"], region))
try:
pab = session.client("s3control").get_public_access_block(AccountId=account_id)
cfg = pab["PublicAccessBlockConfiguration"]
results.append(("DATA-2 S3 account public access block", all(cfg.values()), cfg))
except session.client("s3control").exceptions.NoSuchPublicAccessBlockConfiguration:
results.append(("DATA-2 S3 account public access block", False, "not configured"))
status = session.client("cloudtrail").get_trail_status(Name=trail_name)
results.append(("LOG-1 org trail logging", status["IsLogging"], trail_name))
return resultsRun it across accounts by assuming a read-only audit role in each, and write the results to a table keyed by account, control and date. That table is the baseline's source of truth: it drives the dashboard, the tickets and the metrics. Treat a check that cannot run, for example because the audit role is missing, as a failure rather than a skip, or blind spots will look like compliance. Config drift reconciliation covers how to feed findings back into desired state.
The audit trail is the control that protects the others
Every other control assumes you can reconstruct what happened. The logging design therefore has a stricter bar: an organisation-wide trail or sink, written to storage in a dedicated log archive account, with write-once or retention locks so that even an administrator of a compromised workload account cannot delete history. Only the security team reads it, and access to it is itself logged.
Keep management events on everywhere. Decide deliberately about data events such as object reads, because they are high volume and cost real money; enable them for buckets that hold sensitive data rather than globally. Set retention from your incident-response and legal needs, not from the default.
Worked example: onboarding a new account
A team requests an account for a new service. The account factory creates it inside the workloads organisational unit, so the SCPs apply from its first second: it cannot leave the approved regions, cannot stop the trail, and cannot launch instances without IMDSv2. A pipeline role then applies account-level settings that policies cannot express: default EBS encryption in each approved region, the S3 account-level public access block, and enrolment in threat detection through the delegated administrator.
Within an hour the checker runs. Suppose it reports DATA-1 failing in one region because the bootstrap job timed out there. That is a finding with the platform team as owner; the fix is safe, so remediation re-runs the bootstrap automatically and the next scan passes. The account reaches 100 percent baseline compliance before the team deploys anything, and the evidence is a row in the results table rather than a screenshot.
Exceptions without erosion
Some workloads genuinely need to violate a control: a public website bucket, a legacy appliance without IMDSv2 support. Handle this with an exception register, not by weakening the control. Each entry names the account, the resource, the control, the business owner, the compensating control and an expiry date. The checker reads the register and reports excepted findings separately, and an expired exception turns back into a failure automatically. Exceptions without expiry dates are how baselines quietly decay.
Failure modes
- Detective-only baselines. Findings accumulate faster than teams fix them and the dashboard goes permanently red. Prevent what you can.
- Preventive policy outages. A region deny without global-service exemptions, or an untested SCP applied at the root, breaks every account at once. Stage changes through a sandbox unit.
- Assuming SCPs cover everything. The management account is exempt, and service-linked roles are not restricted by SCPs.
- Silent scan gaps. A missing audit role or a new region makes checks skip rather than fail.
- Break-glass that is never tested. Emergency access that bypasses guardrails must exist, be alarmed on use and be rehearsed.
- Alert fatigue. Mixing baseline findings with thousands of low-severity advisories buries the ones that matter.
Measuring the baseline
Three numbers tell you whether the baseline is working. Coverage: the percentage of accounts passing each control, reported per control rather than as one average that hides a failing control behind passing ones. Time to remediate: how long findings stay open, by severity. Exception health: open exceptions and how many are past expiry. Review them monthly with account owners; a control nobody can pass is a signal to fix the control or the platform, not to accept the red. Zero trust architecture is the natural next layer once the floor holds.
What to do next
- Write down ten to fifteen candidate controls and keep only those that apply everywhere, address a serious risk and can be checked automatically.
- Give each control an identifier, an owner, a preventive mechanism if one exists, and a detective check.
- Move audit logging to a dedicated log archive account with retention locks, and deny stopping or deleting it.
- Deploy preventive policies from a repository through CI, staged via a sandbox organisational unit.
- Run detective checks across every account daily, treating any check that cannot run as a failure.
- Create an exception register with owners and expiry dates, and report coverage, time to remediate and exception health monthly.