It works in staging but breaks in production: where the gap comes from
When the same code passes in staging and fails in production, the code is rarely the problem. The difference almost always comes from one of five places: configuration values, the shape and volume of the data, how many copies of the app are running, the network layer sitting in front of it (proxy, CDN, firewall), and the slow drift that pulls two environments apart over months. Hunting the bug by rereading the code usually ends with nothing found, because there is nothing in the code to find.
Write the differences down instead of guessing them
The Twelve-Factor App has described this since 2011 and names three gaps: the time gap (a developer writes code that reaches production days or weeks later), the personnel gap (one person writes it, another deploys it), and the tools gap (SQLite locally, MySQL in production). Containers have mostly closed the tools gap over the last fifteen years. The other two are still with us.
The practical first step is a small table. Rows: runtime version (Node, Python, JVM), database version and extensions, operating system and locale, the list of environment variable keys, instance count and memory limit, proxy and CDN settings, which mode each third-party service is in, data volume. Columns: local, staging, production. Filling this in once removes most of the guesswork from the next "but it worked in staging" ticket. In most teams the exercise takes about two hours and turns up at least one surprise.
Build once, deploy everywhere
If you build a separate artifact for staging and another for production, the thing you tested is not the thing you shipped.
Three habits break this rule. The first is mutable tags: a Docker tag like :latest or :staging can point at a different image next week, so two environments can run different code under the same label. Pin by digest (image@sha256:...) when you need certainty. The second is skipping the lockfile: npm ci installs exactly what the lockfile says, while npm install may update it. A pipeline that calls npm install will eventually produce different dependency versions in different environments.
The third is quieter: values baked in at build time. The Next.js documentation is explicit about this. Variables prefixed with NEXT_PUBLIC_ are inlined into the browser bundle during next build, and if you promote a single image from one environment to another, those values stay frozen at whatever they were when the image was built. What you see in production is a server that talks to the right API and a browser that keeps calling the staging one. Anything that must change per environment should be served at runtime, not compiled into the client bundle.
Config belongs to the environment, but only if you validate it
Reading values from environment variables is not enough on its own. The real failure is a missing or misspelled variable that nobody notices until the code path runs for the first time. You do not want to learn that the payment provider key is absent during a customer's first checkout.
The fix is a startup check: validate the configuration against a schema when the process boots, and refuse to start if something required is missing. Keep the required keys, their types and their allowed values in one place. That single check catches most of the "we forgot to set it" class of release bugs at the health-check stage rather than in front of a user.
Defaults deserve their own rule. A fallback like DB_HOST ?? "localhost" means that forgetting the variable in production produces no error at all, just a silent connection to the wrong place. Quiet wrong behavior costs more than a loud crash every time.
Then there are credentials. Staging should never run with the production database user, the production payment key or the production mail-sending key. Give each environment its own identity and keep the staging one narrow. The mechanics are in the secrets management post. The worst version we have seen is a staging app wired to the live database, where one bulk-update experiment wiped real customer records and nobody filed it as an environment problem.
The data gap is the expensive one
Staging holds 5,000 rows, production holds 5 million. A cost-based query planner looks at the small table, decides a sequential scan is the cheapest path, and returns in milliseconds. The same query in production picks a different plan, spills to disk because an index is missing, and times out. Same code, same SQL, different outcome. The planner side of this is covered in the slow queries post.
Volume is not the only difference. Real data is messy: line breaks in address fields, apostrophes in names, spaces in phone numbers, emoji in product titles. If your MySQL columns are utf8 rather than utf8mb4, an emoji throws an error in production and never appears in your tidy seed data. Postgres has a less familiar trap: text indexes are built using the operating system's collation rules, and the rule changes that shipped with glibc 2.28 made two databases installed on different distribution releases sort the same strings differently. Postgres 13 and later warn about the version mismatch, and the warning means a REINDEX is due. Running your environments on different OS releases is how you end up there.
The answer is not copying production data down. Build a masked dataset that keeps production's row counts and distribution while replacing personal fields. Why that matters and how the masking is done is in the test data post.
It runs on one instance and falls apart on two
Staging usually runs a single instance. Production runs two or more. That one difference hides an entire class of bugs.
Keep sessions or caches in process memory and users appear logged out the moment they land on the second instance. Run scheduled jobs inside the app and three instances will run the same job three times, sending the same email three times. Write uploads to the local disk and the file exists on one instance only, until the next deploy removes it. Race conditions between two instances updating the same row also live here. Add a read replica and you get one more: a user saves a record, the next read hits the replica, and they see the old value.
The only real antidote is running at least two instances in staging. Two small instances teach you more than one large one. For the instance count and runtime model decision itself, see do you need Kubernetes.
The layer in front of the app does not exist on your laptop
On a developer machine the request reaches the application directly. In production it passes a CDN, a load balancer, a reverse proxy and possibly a web application firewall first. Each of those rewrites something.
The common ones: the proxy terminates TLS, so the app sees plain HTTP, refuses to send a Secure cookie, and the login silently fails. The fix is trusting and reading X-Forwarded-Proto (trust proxy in Express). If a payment return or an embedded iframe needs SameSite=None, browsers only accept it together with Secure, which you will never notice while developing over HTTP. File uploads return 413 in production because nginx defaults client_max_body_size to 1 MB. The CDN adds a caching layer that your dev setup does not have, so customers keep seeing last week's price. The firewall blocks a legitimate form submission because the text inside it looks like SQL.
What these share is that none of them show up in application logs. Put proxy and CDN configuration in the difference table, and when debugging, start by checking what the request actually looked like by the time it reached your code.
A third party's test mode is not the third party
A payment sandbox approves everything and will not reproduce a real bank decline. Shipping providers often have no test endpoint at all. An SMS provider in test mode sends nothing. Some details are environment-specific by design: webhook signing secrets are usually per endpoint, so the staging secret cannot verify a production signature. Rate limits and IP allowlists are typically active only in production.
Two things help. Record the provider's real response shape once and replay it in your own tests, so you find out when their contract changes. And always keep a path that works when the service is unavailable, which is the subject of the third-party outage post.
Two environments drift apart on their own
Drift is never one decision. Someone SSHes into production to fix an urgent problem, changes a setting, and forgets to mirror it. An index gets added by hand. An environment variable is edited in a dashboard. Six months later there are dozens of differences and no map.
Three countermeasures. Make version-controlled files the only route for infrastructure and configuration changes, treat manual edits as exceptions, and write them back the same day. Put a schema comparison in the pipeline so the build complains when the two databases diverge. Compare environment variable keys (keys, not values) on a schedule: a key that exists in production and not in staging is the address of your next surprise. For applying schema changes without downtime, see zero-downtime deployments.
Long-lived staging or a throwaway environment per branch?
A single shared staging environment becomes a lock as the team grows. Three people's work queues up behind each other, nobody can tell whose change broke it, and half-finished data accumulates with no owner. The pattern that has taken over is a temporary environment created automatically for each pull request and deleted when it closes. Vercel and Netlify give you this for the frontend out of the box; backends and databases take setup, but the logic is identical.
A durable environment still earns its place: long-running integration work with external systems, load testing, and the acceptance test the customer signs off on. The two are not exclusive. Per-branch environments speed up feedback during development, and the durable one carries the final pre-release check. Structuring that sign-off is covered in the acceptance testing post, and the load side in load testing and capacity planning.
Some gaps cannot be closed, so verify in production
No staging environment reproduces production's traffic, data volume and real user behavior. Accepting that and moving the risk into how you release is the more honest position.
Three tools carry most of the weight. Feature flags: the code ships to production switched off, then opens to the internal team, then to a small percentage of users (feature flags and progressive rollout). Synthetic checks: a robot running your critical flows (login, add to cart, start checkout) against the live system every minute. And watching the first fifteen minutes after every deploy: error rate, latency, queue depth. With those in place, "it doesn't work in production" stops being a customer complaint and becomes an alert. For thresholds and targets, see SLOs and error budgets.
Three steps that fit into this week
One: export the environment variable keys from both environments and diff them. Keys only, not values. Where they differ, and they will, write down each one and find its owner.
Two: load staging with a masked dataset at production scale, then compare the query plans of your five busiest queries across both environments. Where the plans diverge, you have found a missing index.
Three: run two instances in staging and leave it that way for a week. In-memory sessions, duplicated scheduled jobs and files written to local disk will all surface within days. Each of the three is about a day of work, and each one stops the same class of failure from happening again.
Need help with this topic?