Ransomware is usually discussed as malware, as if the problem were a file that encrypts other files. In practice the encryption is the last step of an intrusion that took hours to weeks, and by the time it runs the attacker typically holds administrative credentials, knows where the backups are, and has often copied data out. What decides whether an organisation has a bad week or a catastrophic quarter is architecture: how far one stolen credential reaches, whether the backup system shares fate with production, whether anyone notices the preparation, and whether a restore has ever been rehearsed at full scale.
This article builds that architecture from first principles. It walks the attack stages and the control that breaks each, explains the rule that matters most (backups must survive an attacker who owns production), gives detection logic, and works a recovery-time estimate so you can tell whether your recovery objective is real.
How an intrusion becomes an encryption event
Most human-operated incidents follow a similar sequence, and each stage is a place to stop it.
- Initial access. A phished credential, an unpatched internet-facing system (VPN appliances and file-transfer products are frequent targets), exposed remote desktop, or access bought from someone who already has it.
- Credential theft and escalation. Harvesting passwords and tokens from memory, browsers and scripts, until the attacker holds a domain administrator or cloud administrator identity.
- Discovery and lateral movement. Mapping file servers, databases, hypervisors and, above all, the backup system, then moving to them over remote management protocols.
- Sabotage of recovery. Deleting backup jobs and catalogs, shortening retention, wiping snapshots and volume shadow copies, disabling endpoint protection.
- Exfiltration. Copying data out so it can be used for extortion even if the victim restores.
- Encryption. Usually pushed everywhere at once, often at night or on a weekend, frequently by encrypting virtual machine storage on the hypervisor so thousands of systems go down together.
The design consequence is that the decisive moments are stages two to four. If a stolen credential cannot become an all-powerful one, if the backup system cannot be reached with production credentials, and if sabotage of recovery triggers an alarm, stage six becomes an outage you restore from rather than an existential event.
The architecture in four layers
Each layer is designed on the assumption that the previous one has already failed. Prevention assumes attackers will get in anyway, so containment limits how far one foothold reaches. Containment assumes some administrator account will be compromised, so detection watches for what attackers must do before encrypting. Detection assumes it will sometimes be too late, so recovery is built to work even when every production administrator credential is in hostile hands.
Identity is the blast radius
Encryption at scale requires administrative reach, so the strongest single control is making administrative credentials hard to steal and narrow in scope.
- Phishing-resistant MFA for every privileged and remote-access login. FIDO2 security keys or platform passkeys resist the proxy-phishing kits that defeat one-time codes and push approvals. Start with administrators, VPN and email.
- Tiered administration. Accounts that administer identity (domain controllers, the identity provider) are never used to log on to workstations or ordinary servers. A workstation compromise then exposes only workstation-tier credentials. Enforce it with logon restrictions, not policy documents.
- Separate admin accounts and admin workstations. Administrators browse and read mail as a normal user and administer from a hardened device that does nothing else.
- Just-in-time elevation. Standing domain admin membership should be close to empty; elevation is requested, approved, time-boxed and recorded. See privileged access management for the machinery.
- Unique local administrator passwords. A shared local admin password turns one compromised laptop into all of them. Microsoft's LAPS (now built into Windows) rotates a unique one per machine.
- Service accounts with least privilege. Backup and monitoring service accounts often hold sweeping rights and never-expiring passwords; inventory them, narrow them, and alert on interactive logons.
In cloud environments the same principle applies to the root or organisation-level identities, break-glass accounts and CI/CD roles that can modify storage policies. The policy-evaluation side of this is covered in IAM architecture.
Segmentation and the admin path
A flat network lets an attacker reach every server from any desk. Segmentation does not need to be perfect to help; it needs to make the high-value paths explicit and narrow.
- Block SMB, RDP, WinRM and SSH between workstations. Workstations rarely have a legitimate reason to talk to each other, and these protocols are the lateral-movement highways.
- Allow administrative protocols to servers only from jump hosts or admin workstations, so an attacker must compromise that small, monitored set first.
- Put hypervisor management interfaces and storage arrays on a management network reachable only from the admin tier, and do not join hypervisor hosts to the same directory domain whose admins you are trying to contain.
The broader model of authenticating every connection rather than trusting the network is in zero trust architecture; for ransomware the specific goal is that compromising one segment does not hand over the management plane of all the others.
Backups that survive an attacker who owns production
The classic 3-2-1 rule (three copies, two media types, one off-site) was designed for disk failures and fires. Ransomware adds an adversary who will log in to the backup console with stolen credentials and delete everything. The modern form is often written 3-2-1-1-0: add one copy that is immutable or offline, and zero errors in restore verification. The requirements that make it real:
- A separate control plane. The backup servers, console and storage are administered with identities that do not exist in the production directory, protected by their own phishing-resistant MFA. A production domain admin should have no route to log in.
- Pull, not push, where possible. The backup system reaches into production to read data. Production holds no credentials that can write to, let alone delete from, the backup store.
- Immutability enforced by the storage, not the software. A retention setting in the backup application can be changed by anyone who owns the application. Object storage with a write-once lock, hardened immutable repositories, or physically offline media enforce it below that layer.
- Retention longer than the time to notice. If attackers were inside for three weeks, a 14-day immutable window may hold only copies taken after they arrived. Keep at least some immutable points older than any intrusion you would plausibly miss.
- Back up the things you need to restore everything else. Directory services, DNS, the identity provider configuration, certificate authorities, infrastructure-as-code repositories, secrets vaults and the backup catalog itself.
On AWS, S3 Object Lock provides storage-enforced immutability. It requires versioning, and offers two retention modes: Governance, which users with a specific permission can override, and Compliance, which nobody, including the account root user, can shorten or remove until the retention date passes. A legal hold is a separate flag with no expiry that blocks deletion until removed. A dedicated backup account with a Compliance-mode default looks like this:
# Run in a separate AWS account used only for backups
aws s3api create-bucket --bucket corp-backup-immutable \
--region us-east-1 --object-lock-enabled-for-bucket
aws s3api put-object-lock-configuration --bucket corp-backup-immutable \
--object-lock-configuration '{"ObjectLockEnabled":"Enabled",
"Rule":{"DefaultRetention":{"Mode":"COMPLIANCE","Days":35}}}'Compliance mode is unforgiving: a mistaken 10-year retention on a petabyte cannot be undone and must be paid for. Test the retention period in Governance mode first, then switch. The details of the feature are in S3 Object Lock.
Detection: the signals before and during encryption
Attackers must do certain things before encrypting, and those are better detection points than the encryption itself. High-value signals, roughly in the order they appear:
- New members of privileged groups, especially outside change windows.
- Interactive logons by service accounts, or admin logons from non-admin workstations (a direct violation of tiering).
- Endpoint protection being disabled or uninstalled on several hosts in a short period.
- Deletion of volume shadow copies, snapshots, backup jobs or reduction of backup retention.
- Unusual outbound volume from a file server or database host.
- Canary files touched: decoy documents placed in shares that no legitimate process opens.
- Mass rename or rewrite: one host modifying hundreds of files per minute with high-entropy content.
Backup-tampering alerts deserve special care: they must be raised by the backup control plane itself and delivered to a channel the attacker does not control, because the attacker may already own the SIEM's admin accounts. A sketch of a file-activity detector for shares, using a sliding window per writer:
from collections import defaultdict, deque
import math, time
WINDOW_S, MAX_WRITES, ENTROPY_MIN = 60, 300, 7.5
CANARIES = {r"\\fs01\finance\~budget_2026.xlsx", r"\\fs01\hr\~salaries.docx"}
recent = defaultdict(deque) # writer (user, host) -> event times
def entropy(sample: bytes) -> float:
if not sample:
return 0.0
counts = [sample.count(b) for b in set(sample)]
return -sum(c / len(sample) * math.log2(c / len(sample)) for c in counts)
def on_file_event(user, host, path, op, first_4k: bytes):
if path in CANARIES:
return isolate(host, user, reason="canary touched")
if op not in ("write", "rename"):
return
q, now = recent[(user, host)], time.time()
q.append(now)
while q and now - q[0] > WINDOW_S:
q.popleft()
if len(q) > MAX_WRITES and entropy(first_4k) > ENTROPY_MIN:
isolate(host, user, reason=f"{len(q)} writes/min, high entropy")Here isolate would call your EDR's network-containment API and disable the user's sessions. Compressed formats are naturally high-entropy, so tune per share and treat the canary as the high-confidence signal. Correlating these events across sources is the job of the SIEM pipeline.
Exfiltration changes what backups can fix
Backups answer the encryption, not the leak. If attackers copied customer data, restoring does not undo the breach, and the incident becomes a data-protection event with notification obligations. The architectural responses are egress filtering on servers (most databases have no business talking to arbitrary internet hosts), alerts on large outbound transfers and on new cloud-storage tools appearing on servers, encryption of sensitive data at rest with keys outside the attacker's reach, and data minimisation: data you deleted on schedule cannot be stolen.
Recovery: the clean-room restore and the arithmetic
Restoring onto infrastructure the attacker still controls is the most common way to be hit twice. A sound plan restores into a clean environment: rebuilt or isolated network, fresh credentials, identity restored first from a known-good point, then core services, then applications by business priority. Every restored system is scanned and its credentials rotated before reconnecting.
Then check whether the recovery time objective is physically possible. Worked example: an organisation with 300 virtual machines totalling 60 TB, backed up to immutable object storage. A restore test measures 800 MB/s aggregate throughput from the store into the recovery cluster.
| Tier | Contents | Data | Transfer at 800 MB/s | Plus rebuild and validation |
|---|---|---|---|---|
| 0 | Directory, DNS, PKI, secrets vault | 2 TB | about 42 minutes | 4-8 hours (forest recovery, credential reset) |
| 1 | Payments, order database, core APIs | 10 TB | about 3.5 hours | 6 hours |
| 2 | Remaining business applications | 48 TB | about 17 hours | 1-2 days |
Transfer alone for 60 TB is 60,000,000 MB / 800 MB/s = 75,000 seconds, nearly 21 hours, and that is the best case: it assumes the store, the network and the target storage sustain that rate concurrently, and that nobody is waiting on a forensic team to declare a segment clean. If the business has written down a four-hour RTO for everything, this table shows it is fiction. The fixes are to tier the restore, keep a warm copy of tier 0 and 1 in a separate environment, raise throughput (parallel streams, larger recovery cluster), or renegotiate the objective. The general method of setting and testing objectives is in disaster recovery.
Failure modes
- Backup server joined to the production domain. The attacker logs in with the domain admin account they already have and deletes everything. Separate the identity domain.
- Immutability that the application can switch off. Retention settings in software are not immutability. Enforce it in storage.
- Immutable window shorter than dwell time. Every immutable copy contains the attacker's persistence. Keep older points and scan before reconnecting.
- Never-tested restores. Missing catalog, missing encryption keys, licence servers that are themselves encrypted. Restore something real every quarter and the whole tier 0 at least yearly.
- Hypervisor management on the user network. One credential encrypts every VM at once. Isolate management.
- Alerts delivered only inside the compromised environment. The attacker reads or silences them. Send backup-tamper alerts out of band.
Trade-offs
Tiered administration and separate admin workstations cost administrator time and are resented at first; they are also the controls that most directly shrink the blast radius. Compliance-mode immutability removes your own ability to delete data early, which conflicts with deletion requests under data-protection law; keep the immutable window as short as the dwell-time argument allows and document the conflict. Automated isolation stops encryption quickly but can cause outages from false positives; begin in alert-only mode, measure, then enable containment for the canary and backup-tamper signals first. A warm standby environment shortens recovery dramatically but roughly doubles the cost of the protected tiers, which is why it is usually reserved for tier 0 and tier 1.