One client is slowing the whole API down: designing rate limits and quotas
When a single client drags your API down, the cause is usually not an attack. It is an honest integration stuck in a loop. Rate limiting does not fix that with one setting, because it is really four decisions: what you count against (IP, API key, account, endpoint), how you count it (fixed window, sliding window, token bucket), what happens when the limit is hit (reject, slow down, queue) and how you tell the client (429, Retry-After, remaining-quota headers). Two more limits get mixed into the same conversation and shouldn't be: a concurrency limit and a commercial quota. Design those six separately, or you end up with a rule that either blocks paying customers or stops nothing at all.
Three different reasons to put a limit in place
The first is abuse: scraping, credential stuffing, bulk account creation. The second is accident: a misconfigured cron job, a client that never exits its pagination loop, a pile of retries that all come back in the same second after an outage. The third is money. If every request carries a metered cost behind it (a model call, an SMS, a carrier lookup, egress traffic) then an unlimited endpoint is a billing exposure.
The first reason keeps growing. Imperva's 2026 bot report puts automated traffic at 53% of the web in 2025, with human traffic down to 47%, and says 27% of bot attacks went straight at APIs without touching the user interface. The second and third reasons, though, are what most teams actually hit. In the incidents we see, the client filling up the limit is usually code written by a customer who has a contract with you.
What you count against is half the design
IP address is the easiest key and the most misleading one. An entire office leaves through one address behind NAT, mobile carriers put thousands of subscribers behind the same address, so a tight per-IP limit punishes a legitimate customer's whole staff together. The reverse problem exists too: an IPv6 subscriber typically gets a /64 or larger block, so counting single addresses counts nothing. You have to count per prefix.
Then there is the layer in front of you. If your application sits behind a CDN or load balancer, the IP it sees belongs to that layer and the real client is in the X-Forwarded-For header. You can only read that header if you trust your own proxy and nothing else. Trust it from anywhere, and an attacker forges the header and walks around the limit entirely. Why proxy headers are invisible in local development is covered in the environment parity post.
For authenticated traffic the right unit is the account or API key, not the address. In a multi-tenant setup, a per-tenant limit is the cheapest tool you have against one customer slowing down everyone else (multi-tenant SaaS architecture).
The login endpoint is its own case and needs both keys at once. Count only per account and an attacker deliberately fills the limit to lock the real user out, turning your protection into a denial-of-service tool. Count only per IP and credential stuffing spread across thousands of addresses stays invisible. Count both, and set the lockout policy using the framing in the authentication build-versus-buy post.
Rate, concurrency and quota are not the same limit
Stripe's published limits are a clean illustration. In live mode the global limit is 100 requests per second per account, individual endpoints default to 25 requests per second, and a sandbox gets 25 globally. Alongside that sits a separate concurrency limit that counts requests in flight at a given moment rather than requests per second. When a 429 comes back, the Stripe-Rate-Limited-Reason header names which limit fired: global-rate, endpoint-rate, global-concurrency, endpoint-concurrency or resource-specific.
A third type is per object: at most 1,000 updates per PaymentIntent per hour, 10 new invoices per subscription per minute. A fourth is purely commercial. Read requests get an allocation averaging 500 per transaction over a rolling 30 days, with a floor of 10,000 reads per month for every account regardless of volume.
Four limit types inside one product, and none substitutes for another. Your design needs the same separation, because "how many requests per second" covers neither long-running heavy queries nor what a customer is entitled to by the end of the month.
The fixed window will fool you
The most common implementation is a counter that resets every minute. The trouble shows up at the boundary. With a limit of 100 per minute, a client can send 100 requests at 00:59 and another 100 at 01:00: two hundred requests in about two seconds. As far as the database you are protecting is concerned, there was no limit. The second problem is herding, since every client learns to fire at the top of the minute.
A sliding window that genuinely counts the last 60 seconds fixes both, at the cost of keeping more state. The common middle ground is to weight the counters of two adjacent windows and approximate the sliding one.
The third option, and usually the best, is a bucket. A token bucket allows a burst up to the bucket's capacity and settles the long-run average at the refill rate. AWS API Gateway throttles with this algorithm and exposes it as two knobs: rate (tokens added per second) and burst (bucket capacity). The leaky bucket is the same family. Nginx's limit_req module implements it:
limit_req_zone $binary_remote_addr zone=api:10m rate=10r/s;
limit_req zone=api burst=20 nodelay;
Shopify uses a leaky bucket across its APIs too, with bucket size and restore rate varying by API and by the merchant's plan. The practical value of the model is that real clients do not send traffic evenly, they send it in bursts. Size the burst allowance to a few seconds of the rate and you protect the average without blocking honest callers.
Not every request costs the same
Counting requests treats a plain GET and a report query that scans ten tables as equal. GitHub moved to points for exactly that reason. On the REST API, GET, HEAD and OPTIONS cost 1 point while POST, PATCH, PUT and DELETE cost 5, and no endpoint may exceed 900 points per minute. The primary limit is 5,000 requests per hour for an authenticated user and 60 per hour per IP for unauthenticated requests.
The secondary limits are the more instructive list: no more than 100 concurrent requests, no more than 90 seconds of CPU time per 60 seconds of real time, and no more than 80 content-generating requests per minute or 500 per hour. On the GraphQL side the accounting is entirely cost-based, with 5,000 points per hour and a ceiling of 500,000 nodes in a single call, the cost derived from the page sizes in the query.
The rule that follows is simple. When your endpoints differ in cost by two orders of magnitude, request count is the wrong unit. Either attach weights, or give the expensive endpoints (search, reports, exports, bulk import) their own much lower limits. To work out which query is expensive and why, see the posts on database bottlenecks and site search.
Rejecting should be the last resort
Returning 429 is the easiest response, not the only one. Three alternatives usually land better.
Shaping. In nginx, if you leave out nodelay, requests inside the burst allowance are not rejected but queued and released at the configured rate. The client waits a moment and still gets its work done, which beats a 429 for an integration that runs once a minute.
Queueing. If the work behind the request is genuinely heavy, stop doing it synchronously. Accept the job, return 202 and deliver the result through a separate endpoint or a webhook. The queue side is covered in the background jobs post.
Shedding by priority. Two of the four limiters Stripe describes on its engineering blog do precisely this. One watches fleet usage and permanently reserves a slice of capacity for critical requests; the other watches worker utilization and sheds low-priority and test-mode traffic first. Your version of that: payments and logins always get through, reports and exports are the first thing to be squeezed under load. The same thinking applies on the load testing and capacity planning side.
What you tell the client
The right status code is 429 Too Many Requests, defined back in 2012 by RFC 6585. Send Retry-After with it. Per RFC 9110 section 10.2.3 that header carries either a number of seconds or an HTTP date, and both are valid.
A small detail that gets missed: when limit_req rejects a request, nginx returns 503 by default. Without limit_req_status 429; client libraries read it as a generic server fault rather than a rate limit, and cannot pick the right backoff behaviour.
For remaining quota, the de facto standard is the shape GitHub uses: x-ratelimit-limit, x-ratelimit-remaining, x-ratelimit-used, x-ratelimit-reset and x-ratelimit-resource. The IETF effort to standardise this is still a draft. draft-ietf-httpapi-ratelimit-headers reached its eleventh revision in May 2026 and is not an RFC yet. It defines two fields, one for the policy and one for the current state:
RateLimit-Policy: "burst";q=100;w=60,"daily";q=1000;w=86400
RateLimit: "default";r=50;t=30
The draft's own warning belongs in your design as well: a client must not treat the advertised quota as a service level agreement, and a saturated server may report lower numbers than its nominal policy.
Finally, make the error body say which limit was hit. Stripe's Stripe-Rate-Limited-Reason is a good model, because knowing whether it was global, per endpoint or concurrency is what lets an integrator make the correct fix. Skip it and your support queue fills with undiagnosable "we're getting 429s" tickets. Publish the limits in your documentation, and remember that tightening one is a breaking change for everyone integrated against it (versioning and backward compatibility).
When you are the one receiving the 429
Carriers, invoicing services, payment providers and model APIs all have limits, so the other half of this topic is your own client code.
Honour Retry-After when it is there. When it is not, back off exponentially and add randomness. Stripe recommends the same thing in its docs and states the reason plainly: without jitter every client returns at the same instant and you get a thundering herd. Retrying three times at a fixed one-second interval does not solve the problem, it sends the same pile three times.
Do not blindly retry a non-idempotent write. In payment flows, use an idempotency key, or your retry logic will manufacture duplicate records (payment integration and reconciliation).
Cap your worker pool concurrency against the other side's limit, too. Fifty workers running in parallel will pile 429s onto a service that accepts 10 requests per second within minutes, and once retries stack on top, the queue grows instead of draining. This is where separating transient from permanent errors and adding a circuit breaker earns its keep (surviving third-party outages).
The most effective measure is never reaching the limit: run a token bucket on your own side and throttle outbound traffic before it leaves. That is what Stripe suggests as well.
Three servers, three separate counters
If each instance keeps the counter in its own memory, your effective limit is multiplied by the instance count. Three replicas and a limit of 100 per second means 300 per second. With autoscaling it gets worse: more traffic brings more instances, more instances loosen the limit, and the limit disappears exactly when you need it.
The usual answer is a shared counter in Redis. The simple version is an atomic increment with an expiry; the better one is a token bucket implemented in a single script. Stripe's limiters keep their state in Redis as well.
That path costs you two things. Latency, because the counter is now a network call on every request. And a failure decision: what happens when the counter is unreachable? Stripe's rule is to fail open, so a bug or outage in the limiter does not start rejecting traffic. Pair that with a local ceiling, otherwise a Redis outage removes all of your protection at once.
Volume traffic is best kept away from the application entirely. Cloudflare's rate limiting rules differ sharply by plan: the free plan gives one rule, counting by IP only, with a 10-second period; Business adds custom counting expressions and periods up to 10 minutes; Enterprise adds counting on response status or headers and much longer windows. The edge cuts volume, the application protects business logic, and neither replaces the other (DDoS protection).
Measure the threshold, don't guess it
Picking the number in a meeting produces either a limit that never triggers or one that stops a customer on day one. The reliable route is a week of data: plot the distribution of request rate per key, find the peak of your busiest legitimate client, set the threshold at a few times that, and give expensive endpoints their own figure.
Roll it out in observation mode first. Nginx ships a switch for this: limit_req_dry_run on; applies no limit but keeps counting excess requests in the shared zone. Writing the rule, watching for a week to see who would have been blocked, talking to those callers and only then enforcing removes most of the risk of a surprise outage.
After that, put the measurement on a dashboard: 429s per key, per endpoint, per customer. A key that never hit the limit and suddenly starts hitting it is either abuse or growth, and in both cases you want to hear it before the customer tells you. Thresholds and alerting are covered in the SLO and error budget post.
Quota is a product decision, not a technical one
A per-second rate limit protects the system. A quota defines what you are selling. Monthly allowance per plan, the difference between a soft limit (warn but allow) and a hard one (stop), whether overage is billed: these are pricing decisions, and engineering should not be making them alone.
AWS API Gateway's usage plans show the shape well. Each API key gets a rate and burst setting, a quota per day, week or month sits on top, and the limits apply in a defined order: per client and per method first, then account level, then the regional ceiling. The same structure transfers to your own product.
Two practical notes. Let users see their own consumption in the dashboard, because otherwise the first place they learn about the quota is your support queue. And put a per-user ceiling on every feature with a metered cost behind it (model calls, SMS, egress), since that limit is billing protection rather than technical protection (cloud cost, adding an AI feature to a product).
Three steps that fit in this week
One: list your five most expensive endpoints (search, reports, export, bulk import, login) and measure current traffic per key. In most teams, writing that list also reveals which customer is closest to the edge.
Two: put a high threshold on those five in observation mode, wait a week, see who hits it, then bring the threshold down to the real number.
Three: standardise the 429 response. Retry-After, remaining-quota headers and an error body that names the limit that fired. On the same day, fix backoff and jitter in your own outbound clients, because you are as often the side receiving 429s as the side sending them.
At Wedevit we map the cost profile of your APIs, split rate limiting and quota enforcement between the edge and the application, document the 429 contract for your integrators and repair retry behaviour in your outbound integrations. A concrete first step: if a customer's code went into a loop today, is there a layer that would stop it, or would you find out from your database load?
Need help with this topic?