İçeriğe geç
wedevit

September 25, 2026 · 10 min read · software

İlhan Buğra Aslan

Nothing changes until the user hits refresh: choosing between WebSocket, SSE and polling


One question settles most of this: does the client have to send data at the same rate it receives it? If yes, as in chat, collaborative editing, shared cursors or a live trading screen, you want WebSocket. If only the server has news to deliver, as in order status, a notification counter, an operations dashboard or report progress, server-sent events cover it with less code and less maintenance. And if nobody minds the screen being ten seconds out of date, plain polling is enough and it is the cheapest thing you can build. Picking the protocol is the easy decision. What costs you weeks is what happens when the connection drops, who is allowed to subscribe to which channel, how an event produced on one server instance reaches a client connected to another, and how many seconds the proxy in front of you waits before killing an idle connection.

Write down a staleness budget per screen

"Make it live" is not a requirement. The requirement is a number: how many seconds out of date is this screen allowed to be? Write that number next to each screen and most of the argument disappears. An order status page is fine at 10 to 30 seconds. A courier tracking map wants something near 5. A chat window with a typing indicator needs roughly 300 milliseconds. If you want to see someone else's cursor move, you are talking about 50. An internal operations dashboard at 60 seconds bothers nobody.

Two more questions belong on the same page. How many concurrent connections: a warehouse team of 30, or 50,000 visitors sitting on the same campaign page? And can you afford to lose an event? A notification badge that misses one update is harmless, since the next page load corrects it. A payment result that goes missing is not harmless, because someone is watching that screen and will try paying again.

Polling is still the right answer for most screens

Teams skip polling because it feels crude. Do the arithmetic first. Five thousand open tabs asking every 5 seconds is 1,000 requests per second, which sounds bad until you notice that nearly all of those requests answer "nothing changed", and making that answer cheap is standard HTTP work.

Start with conditional requests. Put an ETag on the response, let the client ask with If-None-Match, return 304 when nothing moved. The body is never rendered and never transferred. Next, give the client a cursor: it sends the last version it saw as ?since=... and the server returns only what came after. Then the step almost everyone forgets, which is to stop polling while the tab is in the background. The Page Visibility API makes that a three-line change, and it usually removes half the traffic, because people open a tab and walk away from it. The header mechanics are in the caching article.

Also check where your interval came from. "Every 3 seconds" is a default somebody typed, not a decision. If the staleness budget says 30 seconds, ask every 30 seconds. Polling also happens to be the option that scales with no thought at all, because you are never holding connections open.

Long polling is a middle option that costs what SSE costs

With long polling the server holds the request instead of answering it, and replies when something happens. The Socket.IO client still starts this way by default before trying to upgrade to WebSocket, which is why Vercel's own documentation reminds you to configure the client with transports: ['websocket'] when you want the upgrade skipped.

The economics are the point. Long polling keeps a connection open too, so the server-side cost matches SSE, and in exchange you write your own reconnection, sequencing and timeout handling. Unless you are supporting browsers old enough to lack EventSource, there is not much reason left to choose it.

SSE: the least work for the widest set of cases

Server-sent events are an ordinary HTTP response. Content type text/event-stream, the connection stays open, the server writes events line by line. On the browser side EventSource hands you two things for free: it reconnects by itself when the stream drops, and it tells the server the id of the last event it saw through the Last-Event-ID header. Keep a short buffer of recent events on the server and you can replay whatever was missed during the gap. The retry field lets the server dictate the reconnection delay.

Know two limits before you commit. Over HTTP/1.1 the browser will not open more than 6 connections to one domain, and as MDN spells out, that ceiling is per browser rather than per tab, so the seventh open tab simply hangs with no error. Over HTTP/2 the number of concurrent streams is negotiated with the server and defaults to 100, which makes the problem go away. If you already serve over HTTP/2, skip this paragraph.

The second limit is authentication. The EventSource constructor accepts only withCredentials, so there is no way to attach a custom header. No Authorization: Bearer .... That leaves cookies or a token in the URL, and a token in the URL ends up in access logs, proxy logs and sometimes a Referer header. If you must go that way, issue a single-use ticket that expires in minutes. The one-way nature of the stream is rarely a problem: when the client needs to send something, it sends a normal POST.

When WebSocket is genuinely the right call

It is bidirectional and the per-message overhead is tiny, a small frame header instead of the dozens of lines of HTTP headers that every poll drags along. Chat, collaborative document editing, presence indicators, games and any screen where messages flow both ways several times a second are what the protocol is for.

The bill arrives as code you have to write: reconnection with exponential backoff, resubscription, state resync after a gap, heartbeat messages, and backpressure handling for when a consumer cannot keep up. Vercel's example client walks a delay from 1 second up to 30 for exactly this reason. If nobody on the team plans to write that loop, you are not ready to run the protocol.

The infrastructure layer cuts these connections quietly

Persistent connections are usually ended by something other than your application. An AWS Application Load Balancer idles out at 60 seconds by default, adjustable anywhere from 1 to 4,000 seconds. Set your heartbeat to 60 as well and you have built a race condition: the connection drops now and then, and nobody can reproduce it.

On nginx, WebSocket needs proxy_http_version 1.1 plus the Upgrade and Connection headers passed through by hand, or the upgrade never happens. proxy_read_timeout also defaults to 60 seconds. SSE additionally needs proxy_buffering off, or an X-Accel-Buffering: no header on the response, otherwise events sit in a buffer and the screen that worked on the developer's laptop shows nothing in production. That second one is a classic, because there is no bug in the code to find.

Managed platforms write their limits down. API Gateway's WebSocket APIs in AWS idle out after 10 minutes and cap a single connection at 2 hours, and neither quota can be raised. Vercel moved WebSocket support on Functions into public beta on 22 June 2026; a connection closes when the function hits its maximum duration, and a reconnecting client is not guaranteed to land on the same function instance. All of it says the same thing. The connection is not permanent, so treat disconnection as part of the normal flow rather than an error case.

It works on one server and breaks when you add the second

This is the architecture mistake we run into most. The connection lives on one instance while the code that produces the event runs on another. The subscriber list you kept in memory does not exist over there, so some users get the update and some do not. Staging runs a single instance, so nothing surfaces until production.

What you need is a publish/subscribe layer: Redis, NATS, or Postgres LISTEN/NOTIFY. If you pick Postgres, read three warnings in its documentation first. The payload has to be shorter than 8,000 bytes in the default configuration. Notifications are delivered only once the transaction commits. And nothing is durable: with no session listening the notification is gone, and if one listener falls behind far enough to fill the queue (8 GB in a standard install), transactions that call NOTIFY start failing at commit time. So do not ship data through the channel. Send "record 12345 changed" and let the client read the details from your normal API.

Sticky sessions deserve a line too. WebSocket needs the client pinned to an instance; SSE does not, because it rides the same HTTP scaling you already have. For how many instances and which runtime model, see do you need Kubernetes.

The bug users actually report: they came out of a tunnel and the screen is wrong

Live features fail during the gap. The connection drops for 40 seconds, three events are produced, the connection comes back, and the screen carries on displaying stale data with no idea it missed anything. Nobody sees an error, because technically nothing errored.

The pattern that fixes it is a monotonically increasing sequence number on every event. The client remembers the last one it saw and sends it on reconnect. The server either resumes from there or answers "my buffer does not reach back that far, refetch everything". You have to implement that second answer as well, otherwise long disconnections turn into permanent silent drift. SSE gives you the transport half of this through Last-Event-ID; the replay buffer is still yours to keep, sized by something like the last 200 events or the last 60 seconds.

The acceptance test fits in one sentence. Cut the network for 60 seconds, restore it, and wait without touching the page. If the screen repairs itself, the feature is live. If it needs a refresh, what you have is the appearance of a live feature.

A live channel is the easiest place to skip authorisation

The WebSocket handshake is not covered by the same-origin policy, and the browser attaches cookies to it anyway. If the server does not check the Origin header against an explicit allowlist, another site can open a connection with your user's session and read whatever flows through it. The attack is called cross-site WebSocket hijacking and OWASP keeps a cheat sheet for it. CORS instincts mislead people here, because the browser completes the connection even without the server granting cross-origin access and leaves the decision to the server.

The second mistake is more common: authorisation is checked when the connection opens and never again. A "subscribe to order:12345" message arrives and goes straight through. A channel name being hard to guess is not authorisation. Every subscription request has to be validated server-side against the identity on the session, and in a multi-tenant product the tenant id comes from the session, never from the message the client sent. The model itself is covered in the authorisation article, tenant isolation in multi-tenant SaaS architecture, and the general API side in API security.

Third point: what happens to an open connection when permissions change? Strip a user's role and a connection that stays open for two hours can keep streaming data to them. Either keep connection lifetimes short with periodic reauthorisation, or close the affected channels when a permission changes.

Buy the realtime layer or run it yourself

Ably, Pusher, Supabase Realtime and Azure SignalR price the work by connections and messages. A few connections carrying heavy message volume produces completely different arithmetic from many connections carrying occasional messages, so put your own two numbers (peak concurrent connections and daily messages) into the pricing page before deciding anything.

The hidden cost of running it yourself is not server rent. It is the on-call rotation, connection-level load testing, handing over open connections during a release, and absorbing the reconnection storm when a whole region of clients drops at once. A usable rule: if realtime is not the core loop of your product, meaning no customer is paying you for it, buying almost always works out cheaper. If it is the core loop, run it yourself and accept that it is a team's worth of work. The identity-side version of the same decision is in build or buy your login.

An unmeasured live system degrades silently

The metric list is short: concurrent connections, reconnection rate, event lag from produced to rendered at p95, missed event count, and memory per connection.

Alert on the reconnection rate rather than the connection count. A new proxy in the path, a lowered timeout, or a release that broke the heartbeat shows up there first, while the user complaint arrives days later phrased as "it sometimes doesn't update". For the mechanics of attaching targets to those numbers, see the SLO and error budget article. Load testing needs a distinction too: testing concurrent connections is not the same exercise as testing requests per second. What you are measuring is connection count, memory and file descriptor limits. The load testing article is a reasonable place to start building that scenario.

Three things that fit in this week

First, write a staleness budget in seconds next to every screen someone wants to be live. "Order list: 30 seconds", "chat: 0.3 seconds". Protocol arguments held without that table do not resolve.

Second, measure the polling you already have. How many requests, how many return 304, how many come back empty? If conditional requests plus pausing on hidden tabs get you inside the budget, you are done, and you never have to operate a persistent connection.

Third, if the budget is still missed, add SSE on one screen and instrument two things: reconnection rate and event lag. Then cut the network for a minute and watch whether the screen recovers on its own. If what you actually want is to report progress on long-running work, that work has to leave the request path first, which is the subject of the background jobs and queues article.


Need help with this topic?

get in touch →← all posts