İçeriğe geç
wedevit

September 23, 2026 · 9 min read · software

İlhan Buğra Aslan

The shipping API is down and we cannot take orders: designing for third-party outages


The shipping integration stops responding, the checkout spinner keeps turning, and support starts taking calls about a site that will not accept orders. Nothing is wrong with your servers. The outage still belongs to you, because a third-party call sitting in the middle of a user flow will hold your entire application hostage for as long as your code lets it. You cannot stop your vendors from failing. What you can decide is whether their failure stops your product, and that comes down to three questions: does every outbound call have a timeout, do you stop sending requests once errors pile up, and what does the user see while that service is gone.

In most projects those three questions get answered after the first bad outage rather than while the integration is being written. Two incidents from last year show what the delay costs.

Outages are not rare, and 2025 made the point twice

At 11:48 PM PDT on October 19, 2025, the DNS records for DynamoDB's public endpoint in the AWS us-east-1 region went empty. The cause was not an attack or a hardware failure. A latent race condition between two components that automate DNS updates let one of them apply a stale plan while the other deleted it, and every IP address for the endpoint disappeared. According to Amazon's own post-event summary the DNS problem was fixed at 2:40 AM on October 20, but the story did not end there. New EC2 instance launches kept failing until 1:50 PM, Lambda invocations stayed impaired until 2:15 PM, and full recovery across affected systems landed at 2:20 PM. Roughly three hours of primary failure, fifteen hours of tail.

Cloudflare's turn came on November 18, 2025. A database permissions change deployed at 11:05 UTC made the query behind the Bot Management feature file return duplicate rows. The file more than doubled in size, the module that preallocates memory for 200 features could not process it, and the proxy crashed. Sites behind Cloudflare started seeing HTTP 5xx errors from 11:20, core traffic was largely restored by 14:30, and everything was normal at 17:06. The most instructive detail is a smaller one: Cloudflare's status page, hosted entirely outside Cloudflare's own infrastructure, went down at the same moment, which briefly led the team to believe they were under attack.

The shared lesson is not "pick a bigger vendor." These were the biggest vendors. The lesson is that if your product's ability to function depends on a system you do not control, you either describe in advance how that link breaks or you find out on the day it does.

Start with an inventory, not with code

Resilience work starts on one page, not in a repository. List every external service your product talks to: the payment processor and acquiring bank, shipping carriers, the e-invoicing provider, your SMS and OTP vendor, the email service, address validation, the identity provider behind SSO, the ERP, the CDN, DNS, object storage, search, analytics, and the chat widget on your site.

For each line, answer three questions. Which user flow stops if this service does not answer right now? What does an hour of that flow being down cost the business? Is there an acceptable half-working mode while the service is gone?

Answers tend to sort into three buckets. Things that stop money or login (payments, authentication, OTP) are tier one. Things that slow a flow down without killing it (shipping rates, address lookup, invoicing) are tier two. Things nobody would notice (analytics, recommendations, chat) are tier three, except that a tier three service called synchronously behaves exactly like a tier one service, which is where the most frustrating outages come from.

Without this table, the conversation during an incident is an argument about whether to turn something off, conducted under pressure. With it, the decision was already made on a calm day.

No timeout means their outage is your outage

Leaving an outbound call without a timeout is as common as it is quiet, and library defaults are not on your side. Python's requests applies no timeout unless you set one, so a silent server means waiting forever. In Go, the zero value of http.Client means no timeout at all, and that is what http.Get uses. Node's fetch, through undici underneath it, ships with 300 second limits for headers and body, and five minutes on a web request is indistinguishable from infinity.

What follows is predictable. The other side slows down, your workers pile up on those calls, connection pools and threads fill, and pages that have nothing to do with that vendor stop loading too. The user says the site is down, and from where they sit that is accurate.

Pick the number from measured latency rather than from habit. The approach Amazon describes in its engineering library is to choose an acceptable rate of false timeouts first, say 0.1%, then take the matching latency percentile of the downstream service, p99.9 in that example, as your timeout. Calls that cross the public internet need extra headroom for network latency. Setting the value too low hurts as well: a 200 ms limit on a service whose p99 has drifted to 400 ms cancels work that was about to succeed and burns capacity on requests nobody is waiting for anymore.

Set connect and read timeouts separately, and keep a budget for the request as a whole. If a page has to answer in three seconds, no single dependency can be allowed ten.

A slow dependency is worse than a dead one

A service that is fully down is the easy case: the connection is refused, the error returns immediately, and your fallback path runs. The hard case is the one that half works. Five percent of requests error, the rest take nine seconds instead of 200 milliseconds, and the health check endpoint still returns 200 OK. Dashboards green, product broken.

The standard remedy is compartmentalisation. The bulkhead pattern in Microsoft's architecture guidance takes its name from the watertight compartments in a ship's hull: give each dependency its own connection pool and its own concurrency limit, so slow calls to the shipping API cannot drain the pool that serves your login screen. When the 20 connections allocated to one dependency are all busy, new calls fail fast instead of queuing, and the flow moves onto the fallback you designed.

The second remedy is getting work off the request path entirely. Any call the user does not need to wait for (issuing an invoice, sending a notification, pushing an order into the ERP) belongs in the background. The mechanics are in background jobs and queues.

Retries help, until they become the attack

Retrying transient failures is correct. Retrying without a budget turns your clients into a load generator aimed at a service that is already struggling. Two limits from Google's SRE book are a good starting point: at most three attempts per request, and a per-client budget where retries stay under 10% of total requests. The same book shows how layered retries multiply. If every layer retries three times, giving four attempts each, three stacked layers turn a single user action into 64 requests against the database.

The practical rules are short. Put exponential backoff with random jitter between attempts, otherwise every client returns at the same instant and the second wave is worse than the first. Do not retry 4xx responses, because the request is wrong, not the service. Honour Retry-After when the vendor sends it. Every retried write needs to be idempotent or carry an idempotency key, or "no response, let's try again" turns into a customer charged twice and an invoice issued twice. The payment-specific version of this is in payment webhooks and reconciliation.

Circuit breakers: code that knows when to stop asking

A circuit breaker is a small state machine that stops sending requests to a dependency once the error rate crosses a threshold. Microsoft's architecture guidance describes the three states: closed, where calls pass normally; open, entered when failures exceed the threshold within a window, where calls fail instantly without ever reaching the vendor; and half-open, reached after a timer expires, where a limited number of trial requests decide whether the circuit closes again or reopens.

It earns its keep twice. It protects you, because the user drops into a fallback in 50 milliseconds instead of waiting 30 seconds for an error. It protects the vendor, since the last thing a recovering service needs is the backlog of every client hitting it the second it comes back.

Decide what the user sees when it breaks

This is the real design work. For every dependency, write down what "absent" looks like.

  • On reads, serve the last good answer. Shipping rates, exchange rates, product listings and similar data can be cached and served from the last successful response. In HTTP terms that is the stale-if-error and stale-while-revalidate directives, covered in caching strategy.
  • On writes, accept and queue. Take the order, issue the invoice later, push to the ERP when the service returns. "Order received, your invoice will arrive by email" beats "something went wrong" by a wide margin.
  • Define a safe fallback value. If shipping cost cannot be calculated, continuing with a flat rate that protects your margin may be cheaper than losing the cart. That is a product decision, not a developer decision.
  • Keep a second route into the account. If nobody can log in when your SMS provider fails, the outage is total. A second SMS vendor or an app-based code are the usual answers, compared in multi-factor authentication methods.
  • Be able to switch the feature off. A kill switch per dependency lets you disable a feature during an incident without waiting for a deploy. The infrastructure for that is in feature flags and progressive rollout.

When none of that is possible, at least be honest. A screen that says what happened and when to try again produces fewer support tickets and fewer abandoned carts than a spinner that never resolves.

Third-party scripts in the browser are dependencies too

Resilience discussions tend to stay server-side, but the page your user sees also depends on files fetched from other people's domains: tag managers, chat widgets, analytics, maps, fonts, A/B testing tools. When one of them slows down, your page slows down with it, and a synchronously loaded script will block rendering outright.

Load anything non-critical with async or defer, serve fonts from your own domain, and test that your interface still works when a third-party script fails to load at all. The measured cost of these scripts and what they do to your speed metrics is covered in Core Web Vitals and site speed.

A second vendor is not always worth it

"Let's add a backup provider" sounds simple and bills as two integrations, two reconciliation processes, two contracts and two sets of tests. Worse, a fallback path that never runs will not run on the day you need it. Code nobody executes is code nobody knows works.

The test is straightforward: if an hour of downtime on that flow costs more than the second integration costs per year, the backup is justified. For SMS, email and card payments that maths often works out. For maps, address validation or recommendations it usually does not, and graceful degradation is enough. If you do build the second path, route a small percentage of real traffic through it regularly. A path exercised monthly is a path that works during an incident.

If you are not measuring it yourself, the customer tells you first

Track three numbers per dependency: error rate, latency percentiles and the state of its circuit breaker. They belong on a dashboard with alerts, not only in application logs. Do not make the vendor's status page your primary signal; it updates late, and as Cloudflare's own incident showed, it can be unreachable at exactly the wrong moment.

Add a synthetic check that runs from outside your network. A script that walks the critical flow (log in, add to cart, reach the payment step) every minute catches both your failures and your vendors' before a customer does. How to set the thresholds and alerts behind that is in observability, SLOs and error budgets, and who does what once the alert fires belongs to the first 24 hours of incident response.

SLA credits will not cover your losses

A 99.9% uptime commitment allows roughly 43 minutes of downtime in a 30 day month. That sounds tolerable until you remember dependencies chain. If your checkout touches four external services and all four honour their commitments exactly, your expected combined downtime is close to three hours a month, with nobody in breach of anything.

The credit you get when a vendor does breach is usually a percentage of what you paid that month for that service, not a percentage of the orders you lost. The clause worth negotiating is therefore not the credit rate but the communication: how quickly incidents get reported, when the root cause analysis arrives, how much notice planned maintenance gets. Which terms deserve that kind of attention in a supplier agreement is collected in software maintenance and support contracts.

An untested plan is not a plan

Everything above can exist in your code and configuration and still fail on the day. The only way to know is to try it.

Block the dependency's domain in your test environment and watch what the application does. Then build the harder scenario: instead of taking the service down, add five seconds of latency to its responses and return 500s for 20% of requests. Add a degraded-vendor scenario to your load tests, using the method in load testing and capacity planning. Flip your kill switches for real once a quarter, because an untested switch is exactly as trustworthy as an untested fallback.

What you can do this week

  1. Write the dependency list, with the flow each one blocks and the hourly cost of that flow being down.
  2. Hunt for calls with no timeout. Grepping the codebase for outbound calls and finding the ones left on defaults usually takes half a day.
  3. Pick the most critical tier one dependency and define its circuit breaker and its degraded behaviour, agreed with the product side rather than decided in code.
  4. Turn that dependency off in the test environment and watch. Whatever surprises you there is your actual backlog.

At Wedevit we run this as one engagement: build the dependency inventory, put timeout and retry policies, circuit breakers and degradation paths into the product, then verify them with fault injection in a test environment. In the same pass we review your integration endpoints against the checks in API security. All of it is delivered remotely.


Need help with this topic?

get in touch →← all posts