Client-Server Communication Protocols: REST, GraphQL, gRPC, WebSockets, SSE, and WebRTC
Introduction
"How does the client talk to the server?" sounds small in a system design interview, but the answer shapes load balancing, caching, authentication, reconnection, and capacity. Candidates who answer "WebSockets, because it is real time" or "gRPC, because it is fast" without naming the trade-off lose points on the first follow-up.
This post gives senior backend engineers the depth to choose and defend a protocol. We cover six options: REST, GraphQL, gRPC, WebSockets, Server-Sent Events (SSE), and WebRTC. For each: when to use it, how it works on the wire, an example, how it scales, and how it fails. The running example is a live interview platform with sessions, questions, a code editor, phase changes, and a voice conversation, which we assemble into one worked design at the end.
These are not six competitors for one slot. A real product uses several at once, each on the path it suits.

Request/Response vs Streaming
Before comparing protocols, separate two shapes of communication.
Request/response means the client asks, the server answers, and the exchange is over. REST, GraphQL queries and mutations, and unary gRPC calls all fit here. The server holds no per-client state between requests, so any healthy instance can serve the next one, and a failed request is simply retried if it is safe to retry.
Streaming means a connection stays open and messages flow over time. WebSockets, SSE, gRPC streaming, and WebRTC all fit here. The server holds state per connection, so you scale on concurrent connections rather than requests per second, deploys must drain connections, and every client needs a reconnection strategy.
Between the two sit short polling (ask on a timer, wasting requests and adding up to one interval of latency) and long polling (hold the request open until there is news, then ask again). Long polling works through nearly any proxy and remains a reasonable fallback, but costs a full HTTP request per message.
Two further distinctions matter:
- Direction. SSE is server to client only. WebSockets and gRPC bidirectional streams carry both directions.
- Transport. REST, GraphQL, SSE, and WebSockets run over HTTP on TCP (HTTP/1.1 or HTTP/2), and REST, GraphQL, and SSE can also run over HTTP/3 on QUIC. gRPC requires HTTP/2. WebRTC media runs over UDP when it can, because for live audio a late packet is as useless as a lost one, and TCP's in-order retransmission turns one lost packet into a stall for every packet behind it.
Most protocol mistakes in interviews come from forcing a streaming need into request/response (polling every 500 ms for a timer) or forcing a request/response need into streaming (sending form saves over a WebSocket and then reinventing acknowledgements, retries, and status codes).
The diagram below previews how the three request/response options look in the running example: REST and GraphQL at the edge, gRPC between services.

REST
When to Use It
REST is the default for public APIs, CRUD over business entities, and anything that benefits from HTTP's machinery: caching, status codes, idempotent methods, and a vast ecosystem of gateways and CDNs. If you cannot name a reason to use something else, use REST.
How It Works
REST (from Roy Fielding's 2000 dissertation) models the system as resources identified by URLs and manipulated with a small set of HTTP methods. In practice it usually means pragmatic JSON over HTTP rather than full hypermedia, which is fine as long as you use the semantics correctly:
GETis safe and idempotent;PUTandDELETEare idempotent;POSTis neither;PATCHis not guaranteed to be idempotent.- Status codes carry meaning that intermediaries understand:
304 Not Modified,409 Conflict,412 Precondition Failed,429 Too Many RequestswithRetry-After. - Conditional requests use
ETagwithIf-None-Match(for cache revalidation) andIf-Match(for optimistic concurrency). - Cache-Control headers let browsers and CDNs cache responses without application code.
Idempotency is the property that makes retries safe. A client that times out on a PUT can resend it. A client that times out on a POST cannot know whether the first attempt landed, which is why payment-style APIs accept an Idempotency-Key header: the server stores the key with the result and returns the stored result on a retry. This is a widely used convention and is being standardized as an IETF draft.
Example: Interview Resources
Fetching an interview and saving the candidate's code with optimistic concurrency:
GET /v1/interviews/42 HTTP/1.1 Authorization: Bearer eyJ... If-None-Match: "v17" HTTP/1.1 304 Not Modified ETag: "v17"
PUT /v1/interviews/42/code HTTP/1.1 Content-Type: application/json If-Match: "v17" {"language": "python", "source": "def two_sum(nums, target): ..."} HTTP/1.1 412 Precondition Failed
The 412 tells the client another tab saved a newer version, so it merges instead of silently overwriting. Final submission is a POST, so it carries an idempotency key:
POST /v1/interviews/42/submissions HTTP/1.1 Idempotency-Key: 7f3c2a9e-1b4d-4e8a-9c55-0d2f6b1a7e10 Content-Type: application/json {"language": "python", "source": "..."}
Scaling
Stateless REST services scale horizontally behind a load balancer. Reads scale further with HTTP caching: a CDN absorbs public GETs, and ETag revalidation turns many private reads into cheap 304s. Use cursor pagination (?after=<opaque cursor>&limit=50) rather than offsets for large or changing lists, because OFFSET 100000 forces the database to walk and discard every skipped row. Rate limit per client and return 429 with Retry-After so well-behaved clients back off.
Failure Modes and Anti-Patterns
- Chatty clients. A screen needing an interview, its questions, and the candidate becomes several round trips, or N+1 calls over a list. Fix it with
?include=expansion or a backend-for-frontend. - RPC in REST clothing.
POST /doEverythingwith anactionfield throws away caching, idempotency, and meaningful status codes. - Unsafe retries. Retrying non-idempotent
POSTs without an idempotency key creates duplicates. - Breaking changes without versioning. Removing or renaming fields breaks old mobile clients that you cannot force to update. Add fields freely; remove them only behind a version.
GraphQL: Over-Fetching, N+1, and Caching
When to Use It
GraphQL fits when many different clients (web, iOS, Android, partner integrations) need different slices of the same connected data, and the cost of building a REST endpoint per screen has become the bottleneck. It was developed at Facebook for their mobile apps and open-sourced in 2015. It is a weaker fit for simple CRUD or cache-heavy public APIs.
How It Works
A GraphQL server publishes a typed schema. Clients send a query document, usually as a POST to a single endpoint such as /graphql, describing exactly the fields they want. The server parses and validates the query against the schema, then executes it by calling a resolver function for each field. Mutations change data; subscriptions stream updates and are usually delivered over a WebSocket (the graphql-ws protocol) or SSE.
This solves over-fetching and under-fetching for the client by moving complexity to the server, which creates three problems you should raise yourself.
N+1 queries. Resolvers run per field, per object. Fetching 50 interviews with candidate names naively runs 1 query plus 50. The fix is the DataLoader pattern: collect keys requested in the same tick, issue one batched WHERE id IN (...) query, and cache per request.
Caching. GraphQL POST requests to one URL defeat CDNs and browser caches. The mitigations are persisted queries (the client sends a hash of a query registered at build time, often with GET, which makes it cacheable again), normalized client caches (Apollo Client and Relay store objects by type and ID so a mutation's response updates every view of that object), and server-side caching inside resolvers.
Cost control. A client can send a deeply nested or very wide query that fans out into thousands of database calls. Production GraphQL servers enforce depth limits, complexity or cost scoring, timeouts, and, for first-party apps, an allowlist of persisted queries so arbitrary queries are not accepted at all.
Example: One Query for the Session Page
type Interview {
id: ID!
title: String!
status: InterviewStatus!
candidate: Candidate!
questions: [Question!]!
}
query SessionPage($id: ID!) {
interview(id: $id) {
title
status
candidate { name }
questions { id title difficulty }
}
}
The candidate resolver, batched with DataLoader:
// Created once per request, so the cache never leaks across users.
const candidateLoader = new DataLoader(async (ids) => {
const rows = await db.query(
"SELECT id, name FROM candidates WHERE id = ANY($1)", [ids]);
const byId = new Map(rows.map((r) => [r.id, r]));
return ids.map((id) => byId.get(id) ?? null); // same order as ids
});
const resolvers = {
Interview: {
candidate: (interview, _args, ctx) =>
ctx.loaders.candidate.load(interview.candidateId),
},
};
Scaling
The GraphQL layer is stateless; the real work is in resolvers: batching, per-request caching, and bounded fan-out. Larger organizations split the graph across teams with a federated gateway, which adds a query-planning hop. Measure per-resolver timing, because one slow field delays the whole response.
Failure Modes and Anti-Patterns
- Errors with 200 OK. Servers often return 200 with an
errorsarray, so 5xx-only monitoring misses failures. - Exposing the database schema. One type per table makes every schema change a breaking API change.
- Unbounded public queries. Without depth and cost limits, a public GraphQL endpoint is an easy denial-of-service target.
- GraphQL for internal service calls. Between backend services you control both sides; gRPC or REST is usually simpler.
gRPC: HTTP/2, Protobuf, and Streaming
When to Use It
gRPC is the strong default for internal service-to-service communication: typed contracts, generated clients in many languages, compact binary encoding, deadlines, and streaming in both directions. It is a poor fit for public browser-facing APIs, because browsers cannot speak native gRPC.
How It Works
You define services and messages in a .proto file. The compiler generates client stubs and server interfaces. On the wire:
- HTTP/2 carries each call as a stream on a long-lived connection. Calls are multiplexed, with no per-call handshake, and headers are HPACK-compressed.
- Protocol Buffers encode messages in a compact binary format where each field is identified by its field number, not its name. Renaming a field is safe for the binary wire format (though it breaks JSON mapping and field masks), but you must never reuse or change a number; mark removed ones
reserved. Unknown fields are ignored, so versions coexist during rolling deploys. - Status is sent in HTTP/2 trailers (
grpc-status,grpc-message) after the response body. Browserfetchdoes not expose trailers, which is the concrete reason browsers need gRPC-Web: a variant that encodes trailers in the body, translated by a proxy such as Envoy.
gRPC has four call types: unary (one request, one response), server streaming, client streaming, and bidirectional streaming. It also has first-class deadlines: the client says "answer within 300 ms", the deadline propagates downstream, and work is cancelled when it expires. Mention deadlines in an interview; missing timeouts are how one slow dependency becomes a cascading outage.
Example: Question and Transcript Services
syntax = "proto3"; package interview.v1; service QuestionService { rpc GetQuestion(GetQuestionRequest) returns (Question); } service TranscriptService { // Server streaming: push transcript segments as they are produced. rpc StreamTranscript(StreamTranscriptRequest) returns (stream Segment); } message GetQuestionRequest { string question_id = 1; } message Question { string id = 1; string title = 2; Difficulty difficulty = 3; reserved 4; // removed field "legacy_hint"; never reuse 4 repeated string tags = 5; } enum Difficulty { DIFFICULTY_UNSPECIFIED = 0; EASY = 1; MEDIUM = 2; HARD = 3; } message StreamTranscriptRequest { string session_id = 1; } message Segment { string speaker = 1; string text = 2; int64 ts_ms = 3; }
A Python client call with a deadline:
try:
q = stub.GetQuestion(GetQuestionRequest(question_id="q-77"), timeout=0.3)
except grpc.RpcError as e:
if e.code() == grpc.StatusCode.DEADLINE_EXCEEDED:
q = cached_question("q-77") # degrade instead of hanging
else:
raise
Scaling
Because HTTP/2 multiplexes many calls on one long-lived connection, gRPC breaks naive layer 4 load balancing. An L4 balancer picks a backend per TCP connection, so a client with one connection sends all its calls to one backend while new backends sit idle. The fixes are an L7 proxy that balances per call (Envoy, a service mesh sidecar, or a load balancer with gRPC support; AWS Application Load Balancer, for example, supports gRPC target groups) or client-side load balancing, where the client resolves all backend addresses and spreads calls across them.
Configure keepalive pings so idle connections are not silently dropped, and set max_connection_age so clients reconnect and rebalance after a scale-out. Many gRPC implementations default to a 4 MB maximum received message size; if you hit it, you probably want streaming rather than a bigger limit.
Failure Modes and Anti-Patterns
- No deadlines. A call without a deadline can wait forever, tying up threads and connections upstream.
- Retry storms. Retrying at every layer multiplies load during an incident (three layers each making three attempts is up to 27 calls). Retry at one layer, with backoff, jitter, and a retry budget.
- Field number reuse. Reusing a number for a different type corrupts data silently between versions.
- gRPC for a public browser API, followed by a gRPC-Web proxy and a JSON transcoder. For browsers and third parties, REST is usually less work.
WebSockets
When to Use It
WebSockets fit bidirectional, low-latency, high-frequency messaging between a browser and a server: chat, multiplayer games, collaborative editing, live cursors, trading screens. If the client mostly listens and only occasionally sends, SSE plus REST is often simpler.
How It Works
A WebSocket (RFC 6455) starts as an HTTP/1.1 request with an Upgrade header. The server replies 101 Switching Protocols, and from then on the TCP connection carries WebSocket frames in both directions instead of HTTP.
GET /ws/sessions/42 HTTP/1.1 Host: app.example.com Upgrade: websocket Connection: Upgrade Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ== Sec-WebSocket-Version: 13 HTTP/1.1 101 Switching Protocols Upgrade: websocket Connection: Upgrade Sec-WebSocket-Accept: s3pPLMBiTxaQ9kYGzzhZRbK+xOo=
The accept value is derived from the client's key, proving the server understood the handshake. Frames carry text or binary with little overhead; client-to-server frames are masked to protect intermediaries from cache poisoning. Control frames (ping, pong, close) handle liveness and shutdown. RFC 8441 defines WebSockets over HTTP/2, but support across servers, proxies, and load balancers varies, so most deployments still use one TCP connection per WebSocket.
WebSockets give you no message IDs, acknowledgements, automatic reconnection, resume, or request/response correlation. Everything above "a pipe of frames" is yours to design.
Example: Editor Presence Channel
A client that reconnects with jittered exponential backoff and resumes from the last sequence number it saw:
let lastSeq = 0, attempt = 0;
function connect() {
const ws = new WebSocket(`wss://app.example.com/ws/sessions/42?since=${lastSeq}`);
ws.onopen = () => { attempt = 0; };
ws.onmessage = (e) => {
const msg = JSON.parse(e.data); // {seq, type, payload}
if (msg.seq <= lastSeq) return; // drop duplicates after resume
lastSeq = msg.seq;
handle(msg);
};
ws.onclose = () => {
const delay = Math.min(30000, 500 * 2 ** attempt++) * Math.random();
setTimeout(connect, delay); // full jitter avoids a thundering herd
};
}
connect();
Scaling
WebSocket servers are stateful. Each open connection holds a socket, buffers, and usually a subscription, so you scale on concurrent connections and memory, not requests per second. Back-of-envelope: if each connection costs around 50 KB of memory in your server (measure your own; it varies widely by runtime and buffer settings), 100,000 connections is about 5 GB, which fits a few large instances with headroom.
The harder problem is fan-out across instances: a room's members are spread over many gateways. The usual design is a pub/sub backbone (Redis Pub/Sub, NATS, Kafka): each instance subscribes to its clients' rooms and publishes what they send. Alternatively, consistent hashing on room ID puts a room on one instance, which removes the hop but turns hot rooms into hot instances.
Load balancers must allow long-lived connections. AWS Application Load Balancer's idle timeout defaults to 60 seconds (configurable up to 4,000), so send application pings more often than that. Managed services add limits: Amazon API Gateway WebSocket APIs document a 10-minute idle timeout and 2-hour maximum connection duration.
Restarting an instance drops all its connections at once. Drain gradually (stop accepting, then close in batches) and rely on client jitter to spread the reconnect wave.
Failure Modes and Anti-Patterns
- No backpressure. A slow client whose send buffer keeps growing will exhaust server memory. Bound per-connection queues and disconnect, or drop non-critical messages, when they fill.
- No resume. After a reconnect the client has missed messages. Use sequence numbers and a replay buffer, or re-fetch state over REST.
- Authentication only at connect. A token checked once at handshake stays "valid" for hours. Re-validate periodically or close connections when sessions are revoked.
- Cross-site WebSocket hijacking. Browsers send cookies on the upgrade request, and WebSockets are not restricted by CORS. Check the
Originheader. - Replacing all REST with WebSockets. You lose caching, status codes, idempotency, and standard tooling, and you rebuild them badly.
Server-Sent Events (SSE)
When to Use It
SSE fits server-to-client streams where the client does not need to send much back: notifications, progress updates, live scores, dashboards, phase changes in a session, and token-by-token output from an LLM. The client sends anything it needs to through ordinary REST calls.
How It Works
SSE is just an HTTP response that never ends. The server responds with Content-Type: text/event-stream and writes UTF-8 text events separated by blank lines:
retry: 5000 id: 1041 event: phase data: {"phase": "CODING", "endsAt": "2026-10-05T14:30:00Z"} : heartbeat id: 1042 event: timer data: {"remainingSec": 120}
The browser's EventSource API parses this format, dispatches events by type, and reconnects automatically when the connection drops. On reconnect it sends a Last-Event-ID header with the last id it received, so the server can replay what was missed. The retry: field tells the browser how long to wait before reconnecting. Lines starting with : are comments, which make cheap heartbeats.
Because it is plain HTTP, SSE reuses your cookies, HTTP/2, proxies, and load balancers.
Example: Session Event Stream
Server side (Node with Express):
app.get("/v1/sessions/:id/events", async (req, res) => {
res.set({
"Content-Type": "text/event-stream",
"Cache-Control": "no-cache",
"X-Accel-Buffering": "no", // tell nginx not to buffer
});
res.flushHeaders();
const since = Number(req.get("Last-Event-ID") || 0);
let lastSent = since, replaying = true;
const pending = [];
const deliver = (e) => { if (e.seq > lastSent) { lastSent = e.seq; send(res, e); } };
// Subscribe before replaying so events published during the replay query are not lost.
const unsubscribe = bus.subscribe(`session:${req.params.id}`,
(e) => (replaying ? pending.push(e) : deliver(e)));
const hb = setInterval(() => res.write(": hb\n\n"), 20000);
req.on("close", () => { clearInterval(hb); unsubscribe(); });
for (const e of await recentEvents(req.params.id, since)) deliver(e);
pending.forEach(deliver);
replaying = false;
});
function send(res, e) {
res.write(`id: ${e.seq}\nevent: ${e.type}\ndata: ${JSON.stringify(e.data)}\n\n`);
}
Client side:
const es = new EventSource("/v1/sessions/42/events", { withCredentials: true });
es.addEventListener("phase", (e) => showPhase(JSON.parse(e.data)));
es.addEventListener("timer", (e) => syncTimer(JSON.parse(e.data)));
Scaling
SSE has the same stateful-connection profile as WebSockets, and the same answer: a pub/sub backbone so any instance can deliver any session's events, plus a small replay store (a Redis stream or a capped list per session) keyed by event ID so reconnects can resume. Heartbeats every 15 to 30 seconds keep load balancers and NAT devices from closing idle connections.

The HTTP version matters. Under HTTP/1.1, browsers limit concurrent connections per origin (six in major browsers), and each open EventSource permanently occupies one, so a user with several tabs open can starve their own REST calls. Under HTTP/2, each stream is multiplexed on one connection and the limit is the negotiated maximum concurrent streams, which the HTTP/2 specification recommends should be no smaller than 100. Serve SSE over HTTP/2 in production.
Failure Modes and Anti-Patterns
- Proxy buffering. nginx and some CDNs buffer responses by default, so events arrive in bursts or not at all. Disable buffering on the SSE route (
proxy_buffering off;orX-Accel-Buffering: no) and test through the full production path. - No custom headers in
EventSource. The standard API cannot set anAuthorizationheader. Use cookie authentication, or a short-lived single-use token in the query string (and keep it out of access logs), or afetch-based SSE client that reads the stream manually. - Binary data. SSE is text only. Base64 works but inflates payloads by about a third; if you need binary, use WebSockets.
- Events without IDs. Without
id:lines, reconnection resumes with no idea what was missed. - Treating SSE as reliable delivery. Replay covers short disconnects. For state that must be correct, the client should re-fetch the authoritative state over REST after reconnecting.
WebRTC: Signaling, STUN/TURN, and SFUs
When to Use It
WebRTC is for real-time audio and video (calls, meetings, screen sharing, voice AI agents) and low-latency peer data where losing a message beats waiting for it. Nothing else here is built for interactive media.
How It Works
Signaling. Endpoints exchange a session description (SDP: codecs, tracks, encryption fingerprints) and ICE candidates (possible addresses) before media flows. WebRTC does not specify how; you build the channel, usually a WebSocket. One side sends an offer, the other an answer, and candidates trickle in as discovered.
NAT traversal with ICE, STUN, and TURN. ICE gathers candidate addresses and tests pairs until one works. A STUN server tells a client the public IP and port its packets appear from, which is enough for many home networks. Behind restrictive NATs or UDP-blocking firewalls, a TURN server relays all media. TURN is expensive, since every byte crosses your bandwidth bill, but without it some users cannot connect. Offer TURN over TLS on port 443 for locked-down corporate networks.
Media transport. Media is encrypted with SRTP, keyed via DTLS; the WebRTC standards make encryption mandatory. Audio commonly uses the Opus codec and video uses VP8, H.264, or newer codecs. Media flows over UDP where possible, with jitter buffers, loss concealment, and congestion control. Data channels carry arbitrary messages over SCTP, reliable or unreliable, ordered or not.
Topologies: Mesh, MCU, SFU
In a mesh, every participant sends its stream to every other participant. That works for two or three people and collapses quickly: with five participants each sending video at 1 Mbps, each participant uploads 4 Mbps, and home upload links are often the tightest constraint.
An MCU decodes, mixes, and re-encodes all streams into one, which is CPU-heavy and adds latency.
An SFU (selective forwarding unit) receives each participant's stream once and forwards it to the others without decoding. Each client uploads one stream and downloads what it needs. With simulcast, senders upload two or three quality layers and the SFU forwards the layer each receiver can handle. SFUs are the standard for group calls and server-side participants such as recorders or AI agents. An SFU decrypts hop by hop, so it can see media unless you add end-to-end encryption with encoded transforms.
Example: Connecting to an SFU
const pc = new RTCPeerConnection({
iceServers: [
{ urls: "stun:stun.example.com:3478" },
{ urls: "turns:turn.example.com:443?transport=tcp",
username: tempUser, credential: tempPass }, // short-lived credentials
],
});
const mic = await navigator.mediaDevices.getUserMedia({ audio: true });
mic.getTracks().forEach((t) => pc.addTrack(t, mic));
pc.onicecandidate = (e) => e.candidate && signaling.send({ type: "ice", candidate: e.candidate });
pc.ontrack = (e) => { remoteAudio.srcObject = e.streams[0]; };
const offer = await pc.createOffer();
await pc.setLocalDescription(offer);
signaling.send({ type: "offer", sdp: offer.sdp });
signaling.on("answer", (a) => pc.setRemoteDescription(a));
signaling.on("ice", (c) => pc.addIceCandidate(c.candidate));
Issue TURN credentials from your backend with a short expiry (a common pattern is an HMAC-based time-limited username and password) so leaked credentials cannot be used as a free relay.

Scaling
Signaling is light; media is bandwidth-bound. Opus voice is commonly configured in the 24 to 64 kbps range; take 50 kbps per stream. A one-to-one call through an SFU has two streams in and two out, roughly 200 kbps of SFU traffic. At 5,000 concurrent calls that is about 1 Gbps through the SFU fleet, before TURN relays. Video is an order of magnitude more per stream. Scale SFUs horizontally by assigning each room to one SFU (or cascading SFUs for very large rooms), place them in regions close to users, and route each user to the nearest region.
Many teams use a hosted WebRTC provider instead of running SFU and TURN fleets, which is usually right when media infrastructure is not the differentiator: you buy NAT traversal and edge presence and pay per participant-minute.
Failure Modes and Anti-Patterns
- No TURN. Everything works in the office and fails for users on strict corporate or mobile networks.
- Ignoring ICE restarts. A user switching from Wi-Fi to cellular changes address; without ICE restart handling the call just freezes.
- WebRTC as a general data protocol. For client-server messages, a WebSocket is simpler to operate and load balance.
- Device issues. Denied microphone permission or the wrong input device are real failures; run a device check before the session.
Head-to-Head Comparison
| Dimension | REST | GraphQL | gRPC | WebSockets | SSE | WebRTC |
|---|---|---|---|---|---|---|
| Communication shape | Request/response | Request/response; subscriptions | Unary and streaming (all four modes) | Bidirectional stream | Server to client stream | Bidirectional media and data |
| Transport | HTTP/1.1, HTTP/2, HTTP/3 over TCP or QUIC | Usually HTTP | HTTP/2 | TCP after HTTP/1.1 upgrade | HTTP/1.1 or HTTP/2 | UDP preferred (SRTP, SCTP over DTLS), TCP/TLS via TURN |
| Payload format | Usually JSON | JSON | Protobuf (binary) | Text or binary frames | UTF-8 text | Encoded audio/video, arbitrary data |
| Contract and typing | OpenAPI (optional) | Schema (required) | .proto (required) | None built in | None built in | SDP negotiation; app data untyped |
| Browser support | Native | Native (via HTTP) | Needs gRPC-Web proxy | Native | Native (EventSource) | Native |
| HTTP caching | Excellent | Poor without persisted queries | None | None | None | None |
| Reconnect and resume | Not applicable | Not applicable | Retries per call; streams restart | Build it yourself | Built in (Last-Event-ID) | ICE restart; app-level session rejoin |
| Server state per client | None | None | Per stream while open | Per connection | Per connection | Per peer connection, media buffers |
| Load balancing | Any L4/L7 | Any L7 | L7 or client-side | Long-lived connections (L4, or L7 with Upgrade support) | Long-lived connections | Room-to-SFU assignment |
| Latency profile | One round trip per call | One round trip, resolver fan-out behind it | Low; multiplexed connections | Very low once open | Low once open | Lowest for media |
| Operational complexity | Low | Medium (cost limits, N+1, caching) | Medium (proto management, L7 LB) | Medium to high (fan-out, drain, backpressure) | Low to medium (buffering, replay) | High (TURN, SFU, NAT, codecs) |
| Typical workloads | Public APIs, CRUD, payments | Multi-client product APIs | Internal microservices, streaming pipelines | Chat, games, collaboration | Notifications, progress, LLM token streams | Calls, meetings, voice agents |
The table hides one important nuance: these are not mutually exclusive layers. GraphQL subscriptions run over WebSockets or SSE. WebRTC needs a signaling channel, which is usually a WebSocket. gRPC streams can feed the service that produces SSE events. The choice is per path, not per product.
When to Pick Which
Start from the shape of the communication on each path, not from the protocol name. The flowchart below captures the order of questions that settles most cases.

Media comes first because nothing else carries it; push comes next because it changes the connection model.
Scenario 1: Ride-Hailing Driver Location on a Rider's Map
Requirements: a rider watches the assigned driver move on a map. The driver's app sends a location every few seconds; the rider's app only displays it.
Choice: the driver app sends locations with plain REST POSTs (or a small batch every few seconds, which tolerates flaky mobile networks well), and the rider app receives them over SSE or a WebSocket. SSE is enough because the rider sends nothing on that path, and it reconnects automatically on a flaky network. If the same connection also carries chat between rider and driver, a WebSocket becomes the simpler single channel.
Scenario 2: Internal Order Pipeline
Requirements: an order service calls inventory, pricing, and fraud services on every checkout. Latency budget for the whole call chain is tight, teams ship independently, and the services are in four languages.
Choice: gRPC with deadlines propagated across every hop, generated clients per language, and an L7 mesh or client-side load balancing. Asynchronous steps (emails, analytics) go to a queue, not a synchronous call. The public checkout API in front of it stays REST, with an idempotency key on the order POST.
Scenario 3: Multi-Platform Content App
Requirements: web, iOS, and Android apps render feeds, profiles, and detail pages from the same entities, but each screen needs different fields, and mobile clients are slow to update.
Choice: a GraphQL gateway in front of internal gRPC or REST services, with DataLoader in every resolver, persisted queries for first-party apps, and cost limits for anything public. Live counters on a page (new comments) can be a GraphQL subscription over SSE or WebSockets, or a separate SSE stream.
Worked Example: A Live AI Interview App
The product: a candidate talks to an AI interviewer by voice in the browser. The interview moves through timed phases (introduction, coding, system design, wrap-up), the candidate writes code in an in-browser editor, and the transcript and code are evaluated asynchronously afterward.
Assume a target of 2,000 concurrent sessions at peak, each about 45 minutes long.
Voice: WebRTC Through an SFU or Hosted Provider
Voice is the only real-time media path, so it is WebRTC. The browser joins a room on an SFU or hosted provider. The AI interviewer is a server-side participant: an agent worker joins the room, runs speech-to-text on the candidate's audio, sends text to a language model, and publishes synthesized speech back.
Why not stream audio over a WebSocket? You can, but you then own jitter buffering, echo cancellation, loss concealment, and bitrate adaptation, which the browser's WebRTC stack provides, and TCP head-of-line blocking stalls audio behind each lost packet.
Back-of-envelope: two voice streams per session at about 50 kbps each, in and out of the SFU, is about 200 kbps per session; 2,000 sessions is about 400 Mbps of SFU traffic. That is modest, and far below what video would cost, so camera video can stay optional. The browser starts with POST /v1/sessions, which validates access, creates the room, returns a join token, and dispatches an agent worker; signaling uses the SFU's or provider's own channel.
Phases and Timers: SSE
Phase changes and timer warnings flow server to browser at low frequency: SSE's sweet spot. The session service publishes to a per-session pub/sub channel, and whichever instance holds the EventSource connection forwards events. Three details matter:
- The server owns time. Send the phase's absolute end time, not a tick per second; the browser counts down locally, which survives tab throttling.
- Events are replayable. Keep recent events per session so
Last-Event-IDreplays what a briefly disconnected candidate missed. - Re-fetch on reconnect. The client reads the authoritative phase from
GET /v1/sessions/42; replay is the fast path, REST the source of truth.
Why not the WebRTC data channel? It couples control events to the media path and pushes business logic into the agent or SFU. Why not a WebSocket? Nothing here needs client-to-server streaming; submit and end are REST calls.
Code Editor: REST With Versioning
The editor autosaves with a debounced PUT /v1/sessions/42/code carrying If-Match, so two tabs cannot silently overwrite each other; final submission is a POST with an Idempotency-Key. Autosave load is easy to bound: at one save every few seconds per active editor, 2,000 sessions produce at most several hundred writes per second, well within a single relational primary.
If a human later watched or co-edited live, this path would move to a WebSocket with OT or CRDTs; for one author, REST is simpler.
Backend Service Calls: REST or gRPC, Plus a Queue
Internal choices:
- Agent worker to session service: gRPC fits, with a unary call for the interview plan, a client stream for transcript segments, and deadlines everywhere. REST with JSON is a fine choice for a small team.
- Session service to question service: unary gRPC or REST with a short timeout and a cached fallback.
- Session end to evaluation: not synchronous. The session service writes the completion and an outbox event in one transaction; a relay publishes to a queue; evaluation workers score it and write results, which the UI fetches over REST.
GraphQL is optional here: one web client does not need it, but a later mobile app and dashboard slicing the same data differently might.
Putting the Failure Modes Together
- Network drops for 20 seconds. WebRTC restarts ICE, SSE reconnects and replays, the client re-fetches state, and unsaved code is retried with its version precondition. The agent pauses instead of talking to an empty room.
- An SSE instance is deployed. Clients reconnect elsewhere with
Last-Event-ID; no instance is special. - The agent worker crashes. The session service sees the missing heartbeat and starts a new worker with the plan and transcript so far.
- Evaluation is slow. The queue absorbs it; the live session never depends on it.
This design uses four of the six protocols (WebRTC, SSE, REST, and gRPC), each for one stated reason.
In the Interview
How to Justify a Choice
Use the same four steps for every path in your design:
- Name the path and its direction: "Phase events flow server to client, a few per minute per session."
- Name the latency and reliability need: "A second of delay is fine; missing an event is not, because the UI would show the wrong phase."
- Name the protocol and the mechanism that meets the need: "SSE, with event IDs and
Last-Event-IDreplay, plus a REST re-fetch after reconnect." - Name the cost you accept: "Long-lived connections mean we scale on concurrent connections and need a pub/sub backbone; we serve it over HTTP/2 to avoid the browser's per-origin connection limit."
Step four is the one most candidates skip.
Common Traps
- "WebSockets for everything real time." Many real-time needs are one-directional, where SSE is simpler, or fine with polling.
- "gRPC is faster, so use it publicly." Browsers need gRPC-Web and a proxy, and at the edge network latency dominates anyway.
- "GraphQL solves over-fetching" with no mention of N+1, caching, or cost limits. Raise them yourself.
- Forgetting reconnects. Every streaming protocol drops connections. Say what the client does on reconnect and how it catches up.
- Forgetting the load balancer. Mention idle timeouts, L7 balancing for gRPC, and draining on deploy.
- WebRTC without TURN. Saying "peer to peer" without NAT traversal and relays is half an answer.
- Unsourced numbers. Derive capacity from stated requirements ("2,000 sessions at 200 kbps is 400 Mbps") rather than quoting benchmarks you cannot defend.
Likely Follow-Up Questions
- "How do you scale WebSockets to a million connections?" Shard connections across many gateway instances, scale on memory and file descriptors, use a pub/sub backbone for cross-instance fan-out, and drain gradually on deploy with client-side jittered reconnects.
- "SSE or WebSockets for notifications?" SSE, unless the client must also stream to the server. SSE gives automatic reconnect and resume, works over HTTP/2, and reuses HTTP authentication and tooling.
- "Why does gRPC need special load balancing?" HTTP/2 multiplexes many calls on one long-lived connection, so per-connection L4 balancing pins a client to one backend. Balance per call at L7 or on the client.
- "How do you version a gRPC API?" Add fields with new numbers, never reuse or retype a number, reserve removed ones, and create a new package version (
v2) only for breaking changes. - "What if TURN costs explode?" Measure what fraction of sessions relay, make sure direct and STUN paths are tried first, place TURN in-region, and price it in; some users can only connect through it.
- "How would you add live typing visibility for a human observer?" Move the editor path to a WebSocket carrying operations with sequence numbers, using operational transforms or a CRDT if more than one person can edit.
- "How do you handle auth on long-lived connections?" Authenticate at connect, bind to a revocable session, re-check periodically, and close on expiry.
Related reading: Designing a Chat Platform like Discord, Designing a Collaborative Editor like Google Docs, and Message Queues vs Logs vs Pub/Sub.