Building a file upload feature: size limits, validation and safe storage
"Users should be able to upload files" takes one line in a spec and hides four separate decisions: which path the bytes travel, what you agree to accept, where and under what name you store the result, and how you hand it back to a browser. The short version of the answer: keep the bytes off your application server when you can, allowlist the extensions you accept, never store a file under the name the user gave it, and never serve uploaded files from your main domain. Getting these right on day one costs a few days. Retrofitting them means migrating every file already in production.
There is a measurable reason to take this seriously. In the 2025 CWE Top 25, published by MITRE and CISA on December 11, 2025, unrestricted upload of file with dangerous type (CWE-434) sits at number 12. Number 25 on the same list is allocation of resources without limits or throttling (CWE-770), which in practice usually means somebody forgot a size cap. An upload form is one of the few places where an unauthenticated visitor gets to write a file onto your infrastructure. Treat it accordingly.
Which path do the bytes take?
This is the decision everything else hangs off. The traditional route sends the file to your application server, which validates it and writes it to storage. It is easy to write, you see the whole flow in one place, and for small files it works fine. The catch is that every layer between the browser and your code enforces a limit of its own. nginx defaults client_max_body_size to 1 MB and returns a 413 past it. A stock PHP install ships with upload_max_filesize at 2M and post_max_size at 8M. On Vercel a function request body cannot exceed 4.5 MB before you get FUNCTION_PAYLOAD_TOO_LARGE. Behind Cloudflare, request bodies are capped at 100 MB on Free and Pro, 200 MB on Business and 500 MB by default on Enterprise, and since September 4, 2026 Enterprise customers can raise that themselves up to 5 GB from the dashboard.
The second route has the browser write straight to object storage. The flow: the browser asks your app for permission, the app checks the user's authorization and mints a presigned URL or POST policy with the constraints baked in, the browser sends the bytes directly to S3, R2 or Blob storage, and your app only records that a given user uploaded a given object. Your application server never touches the payload. If you expect anything larger than a few megabytes, this should be your default.
The trap in direct uploads
If your server never sees the bytes, it cannot check their size or type either. The common mistake is minting a presigned URL and treating the rest as somebody else's problem. Whoever gets hold of that URL can write 5 GB to a slot you sized for 5 MB. The fix is to put the constraints inside the signature. The content-length-range condition in an S3 POST policy exists exactly for this: you give a minimum and maximum in bytes, S3 compares the incoming Content-Length, and rejects anything outside the range with a 400 before the upload completes. The same policy can pin the object key, the allowed content type and the encryption setting.
Second rule: an object is not real until your application says it is. Do not trust the browser's "upload finished" callback, because it may never arrive, or it may arrive for bytes that were never sent. Listen for the storage event notification, or verify existence and size server side with a HEAD request. The lifetime of the signed URL is a decision too. SigV4 allows up to 7 days (links generated in the S3 console top out at 12 hours), but no upload link needs to outlive 15 minutes. Signatures created through an assumed role die when the role session ends, which is where almost every "I asked for 7 days and it broke after an hour" report comes from.
A size limit never lives in one place
The limit lives in three places and needs to exist in all of them. In the browser, so a user does not spend ten minutes uploading 400 MB only to be rejected at the end. At the edge or reverse proxy, so an abusive request dies early. In the storage policy, because that is the real gate. When those three numbers disagree, the user gets an error nobody can explain.
For genuinely large files, the numbers are worth knowing. A single PUT to S3 tops out at 5 GiB; past that you need multipart, and AWS suggests switching to multipart from around 100 MB anyway. Parts run from 5 MiB to 5 GiB, with a ceiling of 10,000 parts per upload. Cloudflare R2 caps single-operation uploads at 5 GiB as well, reaching about 5 TiB per object through multipart, again with a 10,000 part limit. For uploads that survive a dropped connection you either use the tus protocol or drive the storage provider's multipart API from the browser. The standard version of this is still in flight: the IETF draft draft-ietf-httpbis-resumable-upload reached revision 12 on July 6, 2026 and is not an RFC yet. A field team pushing 500 MB videos over mobile needs resumable uploads. Accounting uploading a 2 MB PDF from the office does not.
Who decides what you accept?
Validation has one governing rule: nothing the client sends you tells you what the file actually is. The OWASP File Upload Cheat Sheet spells this out, and the practical summary looks like this:
- Allowlist extensions, never blocklist them. Blocklists turn into an endless game against
.jpg.php,.phtml,.pHp, double extensions and null bytes. Name the extensions your business actually needs and reject everything else. - Do not trust the
Content-Typeheader. It comes from the client and spoofing it is one line of code. It helps catch a user who picked the wrong file, not an attacker. - Check the file signature. Reading the magic bytes is a necessary step, but on its own it is beatable: a polyglot file that opens with a valid JPEG header and carries executable content is not hard to build.
- Re-encode images. What OWASP calls image rewriting means decoding the upload and writing it out again with your own library. Embedded payloads, stray metadata and malformed structures do not survive the round trip. If you already generate thumbnails, this costs you nothing extra.
- Do not treat SVG as an image. An SVG is a document and script inside it runs. Either sanitize it with a library built for the job or only ever serve it as a download.
- Archives and office files need one more layer. For archives, cap the decompressed size (decompression bombs) and sanitize the entry paths (zip slip). Office documents still carry macro risk.
Storing it: the name is data, not a path
The filename a user gives you is a string to display, not a location on disk. Store the object under a random key such as a UUID, keep the original name in a separate column, and return it in a header at download time. That single habit removes path traversal (../../), accidental overwrites and the whole family of platform-specific filename oddities in one move.
OWASP ranks storage locations clearly: best is a different host entirely, then outside the webroot, and only as a last resort inside the webroot with tight permissions. Keep the bucket private by default. Storage buckets left open to the internet remain one of the most common cloud accidents, which we covered in cloud misconfiguration and shared responsibility. Keeping the storage access keys out of your source code belongs to the same story, and that one is in secrets management.
Serving it: why a separate domain?
GitHub serves user content from raw.githubusercontent.com and Google from googleusercontent.com. This is not branding. An HTML or SVG file uploaded by a user and served from your main domain runs in the same origin as your session cookies, which turns it into stored XSS. A separate domain, or at minimum a separate subdomain, cuts that link.
Three headers do most of the remaining work: write Content-Type from your own record rather than from what the user claimed, send X-Content-Type-Options: nosniff so the browser stops guessing, and use Content-Disposition: attachment for anything meant to be downloaded rather than rendered. Then check authorization on every single download. A system where incrementing the number in /files/1234 hands you somebody else's invoice is wide open no matter how carefully the upload side was built; we went through where and how to run those checks in authorization models. If private files sit behind a CDN, read the cache rules twice. One careless Cache-Control and one customer's document gets served to another, a failure mode we illustrated in caching strategy.
Scanning and quarantine
If your users send files to each other, malware scanning stops being optional. The pattern that works is plain: write the upload to a quarantine bucket, move it to the real bucket if it scans clean, delete it and tell the user if it does not. Teams running their own stack use ClamAV; on the managed side, GuardDuty Malware Protection for S3, generally available since June 11, 2024, scans newly uploaded objects in selected buckets automatically. Either way the scan is asynchronous, so the record sits in a "scanning" state for a while and the file is served to nobody until it clears. Getting that kind of work off the request path is what background jobs and queues are for.
The part nobody plans for: orphaned files
Six months after the feature ships, three kinds of junk have accumulated in your bucket: abandoned multipart uploads, objects left behind by deleted records, and eight copies of the same document. All three cost money and all three leak into your backups. The fixes are cheap on day one: a lifecycle rule that aborts incomplete multipart uploads, a job that deletes the object when the record goes, and a per-account quota. Storage itself is usually cheap; egress is what gets expensive, and we measured where that cost accumulates in cloud cost optimization. Since uploaded files often contain personal data, building retention periods and a deletion path into the product from the start also works out cheaper than bolting them on later.
Eight checks before you ship
- Do the bytes pass through your application server? If so, write down every limit on the path (proxy, runtime, CDN) and make sure they agree.
- Does the signed URL carry size and type constraints? If not, you do not have a limit.
- How long does the signed URL live? Fifteen minutes for uploads and a few minutes for downloads covers most cases.
- Do you have both an extension allowlist and a signature check? Are you re-encoding images?
- Is the object stored under a random key? Is the user's filename used for display only?
- Is the bucket private and are files served from a separate domain? Does the authorization check run on every download request?
- Is there malware scanning, and is the file unreachable until the scan clears?
- Who cleans up orphaned objects? Are the lifecycle rule and per-account quota actually defined?
You can walk an existing system through this list in about an hour, and most teams stall on the first three. At Wedevit we map the upload path end to end for products that move files, put the limits and validation back where they belong, and build the storage and delivery layers together with authorization. Where it matters, we test the upload endpoints alongside the checks covered in API security. All of it is delivered remotely.
Need help with this topic?