Sending one email is a few lines of code. Sending millions a day for many customers, without getting your IP addresses blocked, losing messages, or letting one customer's bad mailing list damage everyone else's delivery, is a distributed-systems problem. The interesting parts are not SMTP itself but the machinery around it: where messages wait, who decides when each one is sent, how the system learns from receivers' replies, and how it keeps tenants from hurting each other.

The email delivery and spam architecture article covers the protocol layer: the SMTP transaction, SPF, DKIM and DMARC, transport security, the retry schedule from RFC 5321, bounce classes and the bulk-sender rules of the large mailbox providers. This article assumes that and looks inside the sending platform: the spool and message state machine, scheduling per receiving provider, IP pools and warm-up, DKIM key management, the event pipeline that turns replies into suppression and reputation signals, and the capacity math to size it.

Advertisement

The shape of the system

A sending platform has two halves joined by durable storage. The front half is a synchronous API: authenticate the tenant, validate the message, check the suppression list, store the message, and return an ID. The back half is asynchronous: a scheduler decides when each stored message may be attempted, MTA workers sign it and hold SMTP sessions with receiving servers, and every reply becomes an event. Events flow back to update suppression lists, notify tenants and adjust the scheduler.

Inside a sending platform: a durable spool, a scheduler that is polite per receiving provider, and an event loop that feeds backSend APIauth, idempotency keyAccept pathvalidate, suppressSpooldurable message storeSchedulerqueue per provider x IP poolMTA workersDKIM sign, SMTP sessionstoken grantedReceiving MXGmail, Microsoft, othersEvent logdelivered, deferred, bouncedSMTP repliesSuppression listhard bounces, complaintsWebhooks to tenantssigned, retriedReputation metricsper tenant x providerpause / slow downchecked on acceptKey serviceDKIM selectors
The accept path writes to a durable spool and returns. The scheduler releases messages per receiving provider and IP pool under token buckets, MTA workers sign and deliver, and every SMTP reply lands in an event log whose consumers update suppression, tenant webhooks and reputation metrics, which in turn slow or pause sending.

The key property is that nothing after the accept path depends on the caller. A tenant's request returns in milliseconds even when Gmail is deferring everything, and a crash of an MTA worker loses no message because the spool, not the worker, is the source of truth. The same accept-then-deliver split appears in notification systems, and the accept call should carry an idempotency key so a client retry does not send twice.

The spool and the message state machine

The spool stores each message body once and tracks a delivery record per recipient, since one message to five recipients at three providers is really several independent deliveries. Each delivery record moves through an explicit state machine, and every transition is written before the action it permits, so a crash at any point is recoverable by replaying from the last durable state.

StateEntered whenLeaves to
acceptedAPI stored the messagequeued, or suppressed if the address is on the list
queuedReady, waiting for a tokenin_flight
in_flightA worker leased it with a timeoutdelivered, deferred, bounced; back to queued if the lease expires
deferred4xx reply or connection failurequeued at next_attempt_at, or expired after the give-up time
delivered250 after DATAterminal (a later asynchronous bounce or complaint is a new event)
bounced / expired / suppressed5xx, give-up, or suppression hitterminal

Two details prevent the classic duplicate-send bug. The worker leases a delivery with an expiry rather than locking it forever, so a dead worker's messages return to the queue. And the transition to delivered happens after the receiver's 250 reply but before anything else; if the worker dies between those two steps the message will be sent again on lease expiry, so at-least-once delivery is the honest guarantee: generate the Message-ID once at accept time and reuse it on every retry. Messages that keep failing for non-SMTP reasons, such as a rendering error, belong in a dead-letter queue rather than an infinite retry loop.

Advertisement

Group by receiving provider, not by domain

Receivers rate-limit by their own infrastructure, not by the recipient domain. Thousands of company domains are hosted by Google Workspace or Microsoft 365, and they share the same receiving limits and reputation view as gmail.com or outlook.com. If you queue by domain, you might open a polite five connections to each of two hundred Workspace domains and hit Google with a thousand at once. So resolve each domain's MX record and map the MX host to a provider bucket.

import dns.resolver   # dnspython

PROVIDERS = {          # MX host suffix -> provider bucket (extend from your own logs)
    "google.com": "google",
    "googlemail.com": "google",
    "protection.outlook.com": "microsoft",
    "outlook.com": "microsoft",
    "yahoodns.net": "yahoo",
}

def provider_for(domain, cache):
    if domain in cache:
        return cache[domain]
    try:
        answers = sorted(dns.resolver.resolve(domain, "MX"), key=lambda r: r.preference)
        host = str(answers[0].exchange).rstrip(".").lower()
    except dns.resolver.NXDOMAIN:
        cache[domain] = ("invalid", None)            # domain does not exist
        return cache[domain]
    except dns.resolver.NoAnswer:
        host = domain                                # RFC 5321 implicit MX: try its A/AAAA
    bucket = next((p for suffix, p in PROVIDERS.items() if host.endswith(suffix)), host)
    cache[domain] = (bucket, host)                   # respect the DNS TTL in a real cache
    return cache[domain]

Cache results for the DNS TTL, since lookups at this scale are a load problem. Domains with no MX and no A record are undeliverable and should be rejected at accept time, which saves a five-day retry cycle. Domains whose MX hosts you do not recognise get their own bucket keyed by host, with conservative defaults.

The scheduler: token buckets with feedback

Each (provider, IP pool) pair gets a token bucket for message rate and a cap on open connections. The scheduler releases messages only when a token is available, and the bucket adapts to receivers' replies: a 4.7.x deferral means the receiver wants you to slow down, so halve the rate and drop a connection; a run of successful deliveries slowly raises the rate towards a ceiling. This additive-increase, multiplicative-decrease shape is the same one TCP uses, for the same reason: it backs off fast when the other side is overloaded and probes gently for more capacity.

import time

class Bucket:
    def __init__(self, rate_per_s, burst, max_conns):
        self.rate, self.burst, self.tokens = rate_per_s, burst, burst
        self.max_conns, self.open_conns, self.stamp = max_conns, 0, time.monotonic()

    def take(self):
        now = time.monotonic()
        self.tokens = min(self.burst, self.tokens + (now - self.stamp) * self.rate)
        self.stamp = now
        if self.tokens >= 1 and self.open_conns < self.max_conns:
            self.tokens -= 1
            return True
        return False

    def on_reply(self, code, enhanced):
        if 400 <= code < 500 and enhanced.startswith("4.7"):   # receiver says slow down
            self.rate = max(0.1, self.rate * 0.5)                # multiplicative decrease
            self.max_conns = max(1, self.max_conns - 1)
        elif 200 <= code < 300:
            self.rate = min(self.rate + 0.05, self.ceiling())   # additive increase

    def ceiling(self):
        return 50.0   # per (provider, IP) cap from warm-up stage and observed limits

def dispatch(queues, buckets, workers):
    # queues[(provider, pool)] holds ready message ids, oldest first
    for key, q in queues.items():
        b = buckets[key]
        while q and b.take():
            workers.submit(key, q.popleft())

In a multi-node deployment the buckets must be shared, otherwise each node is individually polite and collectively rude. Either give each node a fixed share of every bucket or keep buckets in a shared store; the trade-offs are those of any distributed rate limiter. Within a bucket, order by priority first and age second, so password resets and receipts never wait behind a newsletter. Keeping transactional and bulk traffic on separate IP pools makes this ordering much easier, and protects the transactional reputation from marketing complaints.

Reuse connections. An SMTP session can carry many messages, one after another, to the same MX, and opening a fresh TLS connection per message multiplies handshake cost and looks abusive to receivers. Workers should keep a small pool of warm sessions per (provider, IP) and close them on idle timeout.

IP pools and warm-up

Mailbox providers judge a sending IP by its history. A brand-new IP that suddenly sends a hundred thousand messages looks like a compromised machine, so new IPs are warmed up: daily volume starts small and grows only while deferral and complaint rates stay low. The ramp below is an illustrative geometric schedule, not a published standard; real schedules are tuned against each provider's responses, and some providers apply limits per IP and per sending domain.

def warmup_cap(day, target_per_day, start=500, growth=1.5):
    # Daily volume cap for a new IP: grow geometrically, never past the target.
    # Advance a day only if yesterday's deferral and complaint rates stayed under your thresholds.
    return min(target_per_day, int(start * growth ** day))

# day 0: 500, day 5: 3,796, day 10: 28,832, day 15: 218,946 ...

Organise IPs into pools with different trust levels: a transactional pool, a bulk pool for established senders, a probation pool for new tenants or newly imported lists, and dedicated IPs for very large tenants who want their own reputation. Reputation is increasingly attached to the authenticated sending domain as well as the IP, which is why a tenant with poor practices cannot escape it just by moving IPs, and why custom DKIM signing domains for each tenant matter.

DKIM signing and key rotation

Every message is DKIM-signed by the worker just before transmission, after any last header changes, because altering a signed header afterwards breaks the signature. The private keys live in a key service or KMS, and workers either fetch them into memory with short leases or call a signing endpoint. A 2048-bit RSA signature costs on the order of a millisecond of CPU, small per message but not negligible at a thousand messages per second; measure it.

Rotation uses selectors, the label that tells receivers which DNS record holds the public key. To rotate, generate a new key pair, publish its public key under a new selector, wait for DNS propagation and the old record's TTL, switch signing to the new selector, and keep the old public key published for several days so messages still in transit or being re-verified (forwarded mail, delayed retries) still validate. Then revoke the old key by publishing an empty key value. Automate it; a rotation that is done by hand once a year is one that fails when the person who did it last has left.

The event pipeline

Every attempt produces an event: delivered, deferred with the reply text, bounced in-session, or later an asynchronous bounce, a complaint from a feedback loop, an unsubscribe, an open or a click. Write all of them to an append-only log keyed by message ID, and let independent consumers build what they need. The suppression consumer adds hard bounces, complaints and unsubscribes to the tenant's suppression list. The webhook consumer delivers signed events to tenants with retries, following the usual webhook delivery practice. The metrics consumer aggregates rates per tenant, provider and IP over sliding windows.

Those metrics close the loop. If a tenant's hard-bounce rate jumps, they probably imported an old list; if their complaint rate approaches the threshold the big providers publish (Gmail and Yahoo ask bulk senders to keep reported spam below 0.3 percent), pause them automatically before the shared pool's reputation suffers. Treat these circuit breakers as product features with clear notifications, because to the tenant a pause looks like an outage.

Tenant isolation on shared infrastructure

In a shared pool one tenant's behaviour affects every other tenant's delivery, so isolation is a reputation problem as much as a resource one. Combine four controls: per-tenant quotas at the accept path so one tenant cannot fill the spool; per-tenant weighting within each provider bucket so a large campaign cannot starve a small sender's receipts; automatic pause rules driven by bounce and complaint rates; and pool placement by track record, with new tenants starting in probation and graduating on evidence. Per-tenant DKIM domains make reputation attributable, which is what makes graduation and demotion fair.

Worked example: sizing for 20 million messages a day

Assume 20 million messages per day, an average size of 50 KB, and a peak hour carrying five times the average rate because campaigns cluster. All of these are assumptions to replace with your own figures. The average rate is 20,000,000 / 86,400, about 231 messages per second, so the peak is about 1,160 per second.

Throughput per SMTP connection depends on the receiver's latency; suppose you measure four messages per second on a reused connection. Peak then needs about 290 concurrent connections, split across providers by your recipient mix, and each provider's share must fit within the connection caps your buckets have learned. Outbound bandwidth at peak is 1,160 times 50 KB, about 58 MB/s or roughly 465 Mbit/s. DKIM signing at about a millisecond per signature needs a little over one core at peak, so plan several for headroom. The spool absorbs about 1 TB of writes per day, but most messages are deleted soon after delivery; provision instead for the deferred backlog of a large provider deferring everything for several hours.

Failure modes

  • Queueing by domain. Hosted domains hide a shared receiver, so per-domain politeness adds up to a flood.
  • Unshared buckets. Ten scheduler nodes each respecting a limit means ten times the limit.
  • Retry storms. A provider outage ends and every deferred message is released at once; release deferred mail through the same buckets as new mail.
  • Head-of-line blocking. Bulk campaigns delay password resets because they share a queue or an IP pool.
  • Lost events. Webhook consumers fall behind and suppression lags, so the platform keeps mailing addresses that already hard-bounced.
  • Broken DKIM after rotation. The old key is revoked before delayed and forwarded mail is verified.
  • Silent reputation decay. Nobody watches per-provider deferral and complaint rates until a block arrives.

What to do next

  1. Separate the synchronous accept path from asynchronous delivery, with a durable spool, idempotency keys and leased in-flight records.
  2. Write the message state machine down explicitly and persist each transition before acting on it.
  3. Group recipients by resolved MX provider and give each (provider, IP pool) a shared, adaptive token bucket with connection caps.
  4. Split transactional and bulk traffic onto different pools, and warm new IPs on a schedule gated by deferral and complaint rates.
  5. Automate DKIM rotation with overlapping selectors and per-tenant signing domains.
  6. Build the event log first and drive suppression, webhooks and reputation metrics from it, with automatic pause rules.
  7. Size the spool for a multi-hour deferral at peak, not for steady-state throughput.
Key takeaway: An email platform is a queueing system whose most important job is to be polite to each receiver while keeping tenants from harming each other. Accept fast into a durable spool, schedule per receiving provider with adaptive token buckets, protect reputation with separate pools, warm-up and per-tenant signing domains, and turn every SMTP reply into an event that drives suppression, tenant notifications and automatic slow-downs.