Uploading a 20 GB video over a home connection is a different problem from uploading a profile photo. The transfer takes most of an hour, the laptop may sleep, the train enters a tunnel, the browser tab is closed and reopened. A design that sends the file as one HTTP request fails almost every time at that size, and the failures waste everything sent so far. Large upload design is the discipline of making progress durable: split the file, send pieces in parallel, record which pieces arrived, and resume from there.
This page covers the client and transport side: why single requests fail, the two protocol families (offset-based tus and part-based multipart), how to size parts, how many to send at once, how to resume, and how to issue credentials and clean up. What happens to the bytes after they land, including storage layout, deduplication and processing pipelines, is covered in designing a file upload service.
The upload path in one picture
Why one big request fails
Consider a 20 GB file over a 50 Mbit/s uplink. That is 160 gigabits, so 3,200 seconds, about 53 minutes, at full speed. Suppose that in any given minute there is a 1% chance of something that kills a long-lived connection: a Wi-Fi roam, a NAT timeout, a load balancer draining a node. The chance of 53 consecutive clean minutes is 0.99 to the power 53, about 59%. Four uploads in ten fail, and each failure restarts from zero, so expected wasted transfer is enormous and some users never finish.
Split the same file into 64 MiB parts and each takes about 11 seconds. A failure now costs one part, at most 11 seconds of work, and the chance that any given part is interrupted is under 0.2%. The whole design follows from that arithmetic: the unit of retry must be small relative to the mean time between interruptions, and progress must be recorded somewhere that survives the interruption.
Two protocol families: offsets and parts
There are two ways to make an upload resumable. Offset-based protocols treat the upload as one growing resource. In tus, the open protocol most libraries implement, the client creates an upload with POST and gets a URL back, then sends bytes with PATCH and an Upload-Offset header. After any failure it asks the server for the current offset with HEAD and continues from there. If the client's offset disagrees with the server's, the server answers 409 Conflict and changes nothing; with the checksum extension, a chunk whose Upload-Checksum does not match is discarded with status 460. Values in Upload-Metadata are base64-encoded.
POST /files # creation extension
Tus-Resumable: 1.0.0
Upload-Length: 21474836480
Upload-Metadata: filename dmFjYXRpb24ubXA0,filetype dmlkZW8vbXA0
-> 201 Created
Location: https://upload.example.com/files/7f3a
PATCH /files/7f3a
Tus-Resumable: 1.0.0
Content-Type: application/offset+octet-stream
Upload-Offset: 0
<bytes ...connection drops after 1,342,177,280 bytes>
HEAD /files/7f3a # after reconnecting: where did we get to?
Tus-Resumable: 1.0.0
-> 200 OK
Upload-Offset: 1342177280
Upload-Length: 21474836480
Cache-Control: no-store
PATCH /files/7f3a
Upload-Offset: 1342177280 # continue exactly there
...
-> 204 No Content
Upload-Offset: 21474836480Part-based protocols split the file into numbered parts that can arrive in any order, in parallel, and are stitched together by a final complete call. S3's multipart upload is the common example and most object stores offer something similar. Offset-based is simpler for one stream and works through a server you control; part-based allows parallelism and lets the client send bytes straight to object storage, which matters at scale because your API servers never carry file traffic. An IETF HTTP working group draft for resumable uploads, derived from tus, also exists; check its current status before depending on it.
Choosing a part size
Part size is a trade-off between retry cost and overhead. Small parts lose less on failure but add a request, a signature and a round trip each; large parts amortise that overhead but cost more to resend. The object store's limits bound the choice. S3 currently documents a maximum object size of 48.8 TiB, at most 10,000 parts, and parts between 5 MiB and 5 GiB, with no minimum for the last part. The two limits meet exactly: 10,000 parts of 5 GiB is 48.8 TiB, so the biggest objects need the biggest parts.
MiB, GiB = 1 << 20, 1 << 30
MIN_PART, MAX_PART, MAX_PARTS = 5 * MiB, 5 * GiB, 10_000 # S3 limits; last part may be smaller
def plan_parts(size: int, target: int = 64 * MiB) -> list[tuple[int, int, int]]:
"""Return (part_number, offset, length) triples."""
part = max(target, MIN_PART, -(-size // MAX_PARTS)) # ceil(size / 10,000)
part = -(-part // MiB) * MiB # round up to a whole MiB
if part > MAX_PART:
raise ValueError(f"{size} bytes cannot fit in {MAX_PARTS} parts of at most 5 GiB")
return [(i + 1, off, min(part, size - off))
for i, off in enumerate(range(0, size, part))]
parts = plan_parts(20 * GiB) # 320 parts of 64 MiB
parts = plan_parts(1024 * GiB) # 10,000 cap forces 105 MiB parts: 9,987 partsIn practice, pick a target between 8 and 128 MiB, raise it when the file would exceed the part count, and fix it for the life of the upload: the part size is part of the resume state. On mobile, lean towards the small end because connections break more often; on servers moving data between data centres, lean large.
Parallelism, throughput and backoff
Why send parts in parallel at all? A single TCP connection's throughput is bounded by its window divided by the round-trip time. With a 4 MiB effective window and 100 ms RTT, one stream tops out near 335 Mbit/s however fat the pipe; with packet loss its congestion window shrinks further. Several concurrent parts fill the bandwidth-delay product and also isolate failures, because one stalled part does not stop the others.
More is not always better. On a 20 Mbit/s mobile uplink, eight parallel 64 MiB parts each run at 2.5 Mbit/s and take over three minutes, which raises the chance each one is interrupted, and they compete with everything else the device is doing. Start with three or four, measure aggregate throughput, add a slot while throughput rises and drop one when it stops rising or errors appear. Retries need exponential backoff with jitter so thousands of clients recovering from one outage do not return in lockstep, and the upload API needs per-user quotas, as in any rate limiter design.
async def upload_all(session, parts, store, slots=4, max_slots=8):
done = await store.completed_parts(session.upload_id) # survives restarts
pending = [p for p in parts if p[0] not in done]
sem = asyncio.Semaphore(slots)
async def one(part_no, offset, length):
async with sem:
for attempt in range(8):
try:
url = await session.signed_url(part_no) # short-lived, one part
body = await read_slice(session.file, offset, length)
etag = await http_put(url, body, checksum=sha256_b64(body))
await store.mark_done(session.upload_id, part_no, etag)
return
except RetryableError:
await asyncio.sleep(min(30, 0.5 * 2 ** attempt) * random.random())
raise UploadFailed(part_no)
await asyncio.gather(*(one(*p) for p in pending))
etags = await store.completed_parts(session.upload_id)
await session.complete(sorted(etags.items())) # ordered by part number
Resuming after a crash or restart
Resume state has two copies and the server's wins. The client persists the upload id, file identity (name, size, modification time, and ideally a hash of the first and last megabyte), part size and the parts it believes are done, in IndexedDB in browsers or a local database on mobile. On restart it checks the file is unchanged, then asks the server which parts actually exist; with S3 that is ListParts, which returns at most 1,000 parts per response, so a 9,987-part upload needs ten paginated calls. Parts the client marked done but the server lacks are resent; the reverse case, a part the server has but the client never recorded, is simply adopted.
Retrying a part is naturally idempotent: uploading part 17 twice replaces it. The complete call is not, so give it an idempotency key and make the server treat a second complete for the same upload as a lookup, the pattern described in idempotency design.
Credentials, verification and cleanup
Clients should upload straight to storage without holding long-lived credentials. The upload API authenticates the user, checks quota and declared size, creates the multipart upload and returns presigned URLs: each signs one part number for one upload id with a short expiry, typically minutes. Issue them in batches as the scheduler needs them rather than all 10,000 up front, so a leaked list is worth little and expired URLs are refreshed naturally. Send a per-part checksum header that the store verifies on receipt, so corrupted parts fail immediately rather than at the end.
On complete, the API checks that the part list is contiguous, the total size matches what was declared, and the user still owns the session, then asks storage to assemble the object. Treat the result as untrusted input until scanned and validated; that pipeline is the subject of the server-side article, and the storage internals behind multipart assembly are in object storage internals.
Finally, clean up. Parts of uploads that are never completed or aborted are stored and billed indefinitely. Set a bucket lifecycle rule that aborts incomplete multipart uploads after a few days, and have the API abort sessions the user cancels explicitly.
Browser and mobile clients
The engine runs in hostile environments, and each platform adds constraints. In browsers, read parts with File.slice(), which returns a lazy reference rather than loading the file into memory, and compute checksums in a Web Worker so hashing a multi-gigabyte file does not freeze the page. A closed tab kills the upload, so the engine must resume from persisted state the next time the page opens, ideally prompting the user to reselect the same file because browsers do not grant persistent access to it.
On mobile, the operating system suspends apps that leave the foreground. iOS and Android both provide system-managed background transfer facilities that continue uploads from a file on disk after the app is suspended, so write each part to a temporary file and hand it to the system rather than streaming from memory. Respect metered connections: offer a Wi-Fi-only setting for very large files and pause rather than fail when the network disappears. Report progress from acknowledged parts, not bytes written to the socket, or the bar jumps backwards after every retry.
Failure modes
- Part size changed between sessions. A client update changes the default size and resumed uploads stitch misaligned bytes. Persist the size with the upload.
- File changed under the upload. The user edits the file and resume mixes old and new parts. Check size, modification time and sampled hashes before resuming.
- ListParts not paginated. Only the first 1,000 parts are seen, the rest are resent or the complete call omits them.
- Orphaned parts. Abandoned uploads quietly accumulate storage cost. Use the lifecycle abort rule and report its volume.
- Retry storms. Clients retry immediately after a storage blip and prolong it. Use capped exponential backoff with full jitter.
- Stale signed URLs. URLs signed ahead in a batch expire while their parts wait in the queue, or a retry reuses an old URL; storage checks expiry when a request starts, so these fail with 403. Sign just before each attempt and re-sign on 403.
Trade-offs
Offset-based uploads through your own servers give you one simple protocol, inline inspection and easy throttling, at the cost of carrying every byte through your fleet. Direct-to-storage multipart scales without that bandwidth bill and gives parallelism, but moves complexity into the client and makes inline validation impossible, so all checks happen after completion. Larger parts reduce request costs and signing load; smaller parts make mobile uploads more robust. Many products use both: tus-style resumable uploads for small and medium files and browser clients, and direct multipart for very large files and server-to-server transfers.
What to do next
- Measure upload success rate by file-size bucket; the large-file bucket is where single-request designs fail.
- Pick a protocol: tus through your edge for simplicity, or direct multipart with presigned URLs at scale.
- Implement part sizing that respects the 10,000-part limit and persist the chosen size with the session.
- Persist resume state on the client and reconcile with the server's part list, paginating it, on every restart.
- Start at three or four parallel parts, adapt to measured throughput, and back off with jitter.
- Issue short-lived per-part URLs in batches and send per-part checksums.
- Make complete idempotent and add a lifecycle rule that aborts incomplete uploads after a few days.