# Client-Server Communication Protocols: REST, GraphQL, gRPC, WebSockets, SSE, and WebRTC

## Blog Details

- **Author**: Naveen R.
- **Date**: October 7, 2026
- **Tags**: system design, api design, grpc, websockets, webrtc
- **Read Time**: 20 mins

## Introduction

"How does the client talk to the server?" sounds small in a system design interview, but the answer shapes load balancing, caching, authentication, reconnection, and capacity. Candidates who answer "WebSockets, because it is real time" or "gRPC, because it is fast" without naming the trade-off lose points on the first follow-up.

This post gives senior backend engineers the depth to choose and defend a protocol. We cover six options: REST, GraphQL, gRPC, WebSockets, Server-Sent Events (SSE), and WebRTC. For each: when to use it, how it works on the wire, an example, how it scales, and how it fails. The running example is a live interview platform with sessions, questions, a code editor, phase changes, and a voice conversation, which we assemble into one worked design at the end.

These are not six competitors for one slot. A real product uses several at once, each on the path it suits.

![A protocol map showing a browser or mobile client reaching the server through six options: REST and GraphQL for request and response, gRPC for typed calls over HTTP/2 (through gRPC-Web from browsers), WebSockets for two-way frames, SSE for server push, and WebRTC for real-time media.](https://d5osvdbc8um23.cloudfront.net/static-asset/blog_images/api-and-realtime-protocols-compared/01-high-level-architecture.png)

## Request/Response vs Streaming

Before comparing protocols, separate two shapes of communication.

**Request/response** means the client asks, the server answers, and the exchange is over. REST, GraphQL queries and mutations, and unary gRPC calls all fit here. The server holds no per-client state between requests, so any healthy instance can serve the next one, and a failed request is simply retried if it is safe to retry.

**Streaming** means a connection stays open and messages flow over time. WebSockets, SSE, gRPC streaming, and WebRTC all fit here. The server holds state per connection, so you scale on concurrent connections rather than requests per second, deploys must drain connections, and every client needs a reconnection strategy.

Between the two sit **short polling** (ask on a timer, wasting requests and adding up to one interval of latency) and **long polling** (hold the request open until there is news, then ask again). Long polling works through nearly any proxy and remains a reasonable fallback, but costs a full HTTP request per message.

Two further distinctions matter:

- **Direction.** SSE is server to client only. WebSockets and gRPC bidirectional streams carry both directions.
- **Transport.** REST, GraphQL, SSE, and WebSockets run over HTTP on TCP (HTTP/1.1 or HTTP/2), and REST, GraphQL, and SSE can also run over HTTP/3 on QUIC. gRPC requires HTTP/2. WebRTC media runs over UDP when it can, because for live audio a late packet is as useless as a lost one, and TCP's in-order retransmission turns one lost packet into a stall for every packet behind it.

Most protocol mistakes in interviews come from forcing a streaming need into request/response (polling every 500 ms for a timer) or forcing a request/response need into streaming (sending form saves over a WebSocket and then reinventing acknowledgements, retries, and status codes).

The diagram below previews how the three request/response options look in the running example: REST and GraphQL at the edge, gRPC between services.

![Request/response paths in the interview app: the web app fetches an interview over REST with an ETag and sends a GraphQL query to resolvers that batch calls with DataLoader, both reaching the interview service, which calls a question service over unary gRPC and reads the interviews database.](https://d5osvdbc8um23.cloudfront.net/static-asset/blog_images/api-and-realtime-protocols-compared/02-request-response-rest-graphql-grpc.png)

## REST

### When to Use It

REST is the default for public APIs, CRUD over business entities, and anything that benefits from HTTP's machinery: caching, status codes, idempotent methods, and a vast ecosystem of gateways and CDNs. If you cannot name a reason to use something else, use REST.

### How It Works

REST (from Roy Fielding's 2000 dissertation) models the system as **resources** identified by URLs and manipulated with a small set of HTTP methods. In practice it usually means pragmatic JSON over HTTP rather than full hypermedia, which is fine as long as you use the semantics correctly:

- `GET` is safe and idempotent; `PUT` and `DELETE` are idempotent; `POST` is neither; `PATCH` is not guaranteed to be idempotent.
- **Status codes** carry meaning that intermediaries understand: `304 Not Modified`, `409 Conflict`, `412 Precondition Failed`, `429 Too Many Requests` with `Retry-After`.
- **Conditional requests** use `ETag` with `If-None-Match` (for cache revalidation) and `If-Match` (for optimistic concurrency).
- **Cache-Control** headers let browsers and CDNs cache responses without application code.

Idempotency is the property that makes retries safe. A client that times out on a `PUT` can resend it. A client that times out on a `POST` cannot know whether the first attempt landed, which is why payment-style APIs accept an `Idempotency-Key` header: the server stores the key with the result and returns the stored result on a retry. This is a widely used convention and is being standardized as an IETF draft.

### Example: Interview Resources

Fetching an interview and saving the candidate's code with optimistic concurrency:

```http
GET /v1/interviews/42 HTTP/1.1
Authorization: Bearer eyJ...
If-None-Match: "v17"

HTTP/1.1 304 Not Modified
ETag: "v17"
```

```http
PUT /v1/interviews/42/code HTTP/1.1
Content-Type: application/json
If-Match: "v17"

{"language": "python", "source": "def two_sum(nums, target): ..."}

HTTP/1.1 412 Precondition Failed
```

The `412` tells the client another tab saved a newer version, so it merges instead of silently overwriting. Final submission is a `POST`, so it carries an idempotency key:

```http
POST /v1/interviews/42/submissions HTTP/1.1
Idempotency-Key: 7f3c2a9e-1b4d-4e8a-9c55-0d2f6b1a7e10
Content-Type: application/json

{"language": "python", "source": "..."}
```

### Scaling

Stateless REST services scale horizontally behind a load balancer. Reads scale further with HTTP caching: a CDN absorbs public `GET`s, and `ETag` revalidation turns many private reads into cheap `304`s. Use cursor pagination (`?after=<opaque cursor>&limit=50`) rather than offsets for large or changing lists, because `OFFSET 100000` forces the database to walk and discard every skipped row. Rate limit per client and return `429` with `Retry-After` so well-behaved clients back off.

### Failure Modes and Anti-Patterns

- **Chatty clients.** A screen needing an interview, its questions, and the candidate becomes several round trips, or N+1 calls over a list. Fix it with `?include=` expansion or a backend-for-frontend.
- **RPC in REST clothing.** `POST /doEverything` with an `action` field throws away caching, idempotency, and meaningful status codes.
- **Unsafe retries.** Retrying non-idempotent `POST`s without an idempotency key creates duplicates.
- **Breaking changes without versioning.** Removing or renaming fields breaks old mobile clients that you cannot force to update. Add fields freely; remove them only behind a version.

## GraphQL: Over-Fetching, N+1, and Caching

### When to Use It

GraphQL fits when many different clients (web, iOS, Android, partner integrations) need different slices of the same connected data, and the cost of building a REST endpoint per screen has become the bottleneck. It was developed at Facebook for their mobile apps and open-sourced in 2015. It is a weaker fit for simple CRUD or cache-heavy public APIs.

### How It Works

A GraphQL server publishes a typed **schema**. Clients send a query document, usually as a `POST` to a single endpoint such as `/graphql`, describing exactly the fields they want. The server parses and validates the query against the schema, then executes it by calling a **resolver** function for each field. Mutations change data; subscriptions stream updates and are usually delivered over a WebSocket (the `graphql-ws` protocol) or SSE.

This solves over-fetching and under-fetching for the client by moving complexity to the server, which creates three problems you should raise yourself.

**N+1 queries.** Resolvers run per field, per object. Fetching 50 interviews with candidate names naively runs 1 query plus 50. The fix is the **DataLoader** pattern: collect keys requested in the same tick, issue one batched `WHERE id IN (...)` query, and cache per request.

**Caching.** GraphQL `POST` requests to one URL defeat CDNs and browser caches. The mitigations are **persisted queries** (the client sends a hash of a query registered at build time, often with `GET`, which makes it cacheable again), **normalized client caches** (Apollo Client and Relay store objects by type and ID so a mutation's response updates every view of that object), and server-side caching inside resolvers.

**Cost control.** A client can send a deeply nested or very wide query that fans out into thousands of database calls. Production GraphQL servers enforce depth limits, complexity or cost scoring, timeouts, and, for first-party apps, an allowlist of persisted queries so arbitrary queries are not accepted at all.

### Example: One Query for the Session Page

```graphql
type Interview {
  id: ID!
  title: String!
  status: InterviewStatus!
  candidate: Candidate!
  questions: [Question!]!
}

query SessionPage($id: ID!) {
  interview(id: $id) {
    title
    status
    candidate { name }
    questions { id title difficulty }
  }
}
```

The candidate resolver, batched with DataLoader:

```javascript
// Created once per request, so the cache never leaks across users.
const candidateLoader = new DataLoader(async (ids) => {
  const rows = await db.query(
    "SELECT id, name FROM candidates WHERE id = ANY($1)", [ids]);
  const byId = new Map(rows.map((r) => [r.id, r]));
  return ids.map((id) => byId.get(id) ?? null); // same order as ids
});

const resolvers = {
  Interview: {
    candidate: (interview, _args, ctx) =>
      ctx.loaders.candidate.load(interview.candidateId),
  },
};
```

### Scaling

The GraphQL layer is stateless; the real work is in resolvers: batching, per-request caching, and bounded fan-out. Larger organizations split the graph across teams with a federated gateway, which adds a query-planning hop. Measure per-resolver timing, because one slow field delays the whole response.

### Failure Modes and Anti-Patterns

- **Errors with 200 OK.** Servers often return 200 with an `errors` array, so 5xx-only monitoring misses failures.
- **Exposing the database schema.** One type per table makes every schema change a breaking API change.
- **Unbounded public queries.** Without depth and cost limits, a public GraphQL endpoint is an easy denial-of-service target.
- **GraphQL for internal service calls.** Between backend services you control both sides; gRPC or REST is usually simpler.

## gRPC: HTTP/2, Protobuf, and Streaming

### When to Use It

gRPC is the strong default for **internal service-to-service** communication: typed contracts, generated clients in many languages, compact binary encoding, deadlines, and streaming in both directions. It is a poor fit for public browser-facing APIs, because browsers cannot speak native gRPC.

### How It Works

You define services and messages in a `.proto` file. The compiler generates client stubs and server interfaces. On the wire:

- **HTTP/2** carries each call as a stream on a long-lived connection. Calls are multiplexed, with no per-call handshake, and headers are HPACK-compressed.
- **Protocol Buffers** encode messages in a compact binary format where each field is identified by its **field number**, not its name. Renaming a field is safe for the binary wire format (though it breaks JSON mapping and field masks), but you must never reuse or change a number; mark removed ones `reserved`. Unknown fields are ignored, so versions coexist during rolling deploys.
- **Status is sent in HTTP/2 trailers** (`grpc-status`, `grpc-message`) after the response body. Browser `fetch` does not expose trailers, which is the concrete reason browsers need **gRPC-Web**: a variant that encodes trailers in the body, translated by a proxy such as Envoy.

gRPC has four call types: **unary** (one request, one response), **server streaming**, **client streaming**, and **bidirectional streaming**. It also has first-class **deadlines**: the client says "answer within 300 ms", the deadline propagates downstream, and work is cancelled when it expires. Mention deadlines in an interview; missing timeouts are how one slow dependency becomes a cascading outage.

### Example: Question and Transcript Services

```protobuf
syntax = "proto3";
package interview.v1;

service QuestionService {
  rpc GetQuestion(GetQuestionRequest) returns (Question);
}

service TranscriptService {
  // Server streaming: push transcript segments as they are produced.
  rpc StreamTranscript(StreamTranscriptRequest) returns (stream Segment);
}

message GetQuestionRequest { string question_id = 1; }

message Question {
  string id = 1;
  string title = 2;
  Difficulty difficulty = 3;
  reserved 4;            // removed field "legacy_hint"; never reuse 4
  repeated string tags = 5;
}

enum Difficulty { DIFFICULTY_UNSPECIFIED = 0; EASY = 1; MEDIUM = 2; HARD = 3; }

message StreamTranscriptRequest { string session_id = 1; }
message Segment { string speaker = 1; string text = 2; int64 ts_ms = 3; }
```

A Python client call with a deadline:

```python
try:
    q = stub.GetQuestion(GetQuestionRequest(question_id="q-77"), timeout=0.3)
except grpc.RpcError as e:
    if e.code() == grpc.StatusCode.DEADLINE_EXCEEDED:
        q = cached_question("q-77")  # degrade instead of hanging
    else:
        raise
```

### Scaling

Because HTTP/2 multiplexes many calls on one long-lived connection, gRPC breaks naive **layer 4 load balancing**. An L4 balancer picks a backend per TCP connection, so a client with one connection sends all its calls to one backend while new backends sit idle. The fixes are an **L7 proxy** that balances per call (Envoy, a service mesh sidecar, or a load balancer with gRPC support; AWS Application Load Balancer, for example, supports gRPC target groups) or **client-side load balancing**, where the client resolves all backend addresses and spreads calls across them.

Configure keepalive pings so idle connections are not silently dropped, and set `max_connection_age` so clients reconnect and rebalance after a scale-out. Many gRPC implementations default to a 4 MB maximum received message size; if you hit it, you probably want streaming rather than a bigger limit.

### Failure Modes and Anti-Patterns

- **No deadlines.** A call without a deadline can wait forever, tying up threads and connections upstream.
- **Retry storms.** Retrying at every layer multiplies load during an incident (three layers each making three attempts is up to 27 calls). Retry at one layer, with backoff, jitter, and a retry budget.
- **Field number reuse.** Reusing a number for a different type corrupts data silently between versions.
- **gRPC for a public browser API**, followed by a gRPC-Web proxy and a JSON transcoder. For browsers and third parties, REST is usually less work.

## WebSockets

### When to Use It

WebSockets fit **bidirectional, low-latency, high-frequency** messaging between a browser and a server: chat, multiplayer games, collaborative editing, live cursors, trading screens. If the client mostly listens and only occasionally sends, SSE plus REST is often simpler.

### How It Works

A WebSocket (RFC 6455) starts as an HTTP/1.1 request with an `Upgrade` header. The server replies `101 Switching Protocols`, and from then on the TCP connection carries WebSocket **frames** in both directions instead of HTTP.

```http
GET /ws/sessions/42 HTTP/1.1
Host: app.example.com
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==
Sec-WebSocket-Version: 13

HTTP/1.1 101 Switching Protocols
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Accept: s3pPLMBiTxaQ9kYGzzhZRbK+xOo=
```

The accept value is derived from the client's key, proving the server understood the handshake. Frames carry text or binary with little overhead; client-to-server frames are masked to protect intermediaries from cache poisoning. Control frames (`ping`, `pong`, `close`) handle liveness and shutdown. RFC 8441 defines WebSockets over HTTP/2, but support across servers, proxies, and load balancers varies, so most deployments still use one TCP connection per WebSocket.

WebSockets give you **no** message IDs, acknowledgements, automatic reconnection, resume, or request/response correlation. Everything above "a pipe of frames" is yours to design.

### Example: Editor Presence Channel

A client that reconnects with jittered exponential backoff and resumes from the last sequence number it saw:

```javascript
let lastSeq = 0, attempt = 0;

function connect() {
  const ws = new WebSocket(`wss://app.example.com/ws/sessions/42?since=${lastSeq}`);
  ws.onopen = () => { attempt = 0; };
  ws.onmessage = (e) => {
    const msg = JSON.parse(e.data);       // {seq, type, payload}
    if (msg.seq <= lastSeq) return;       // drop duplicates after resume
    lastSeq = msg.seq;
    handle(msg);
  };
  ws.onclose = () => {
    const delay = Math.min(30000, 500 * 2 ** attempt++) * Math.random();
    setTimeout(connect, delay);           // full jitter avoids a thundering herd
  };
}
connect();
```

### Scaling

WebSocket servers are stateful. Each open connection holds a socket, buffers, and usually a subscription, so you scale on **concurrent connections and memory**, not requests per second. Back-of-envelope: if each connection costs around 50 KB of memory in your server (measure your own; it varies widely by runtime and buffer settings), 100,000 connections is about 5 GB, which fits a few large instances with headroom.

The harder problem is **fan-out across instances**: a room's members are spread over many gateways. The usual design is a **pub/sub backbone** (Redis Pub/Sub, NATS, Kafka): each instance subscribes to its clients' rooms and publishes what they send. Alternatively, consistent hashing on room ID puts a room on one instance, which removes the hop but turns hot rooms into hot instances.

Load balancers must allow long-lived connections. AWS Application Load Balancer's idle timeout defaults to 60 seconds (configurable up to 4,000), so send application pings more often than that. Managed services add limits: Amazon API Gateway WebSocket APIs document a 10-minute idle timeout and 2-hour maximum connection duration.

Restarting an instance drops all its connections at once. Drain gradually (stop accepting, then close in batches) and rely on client jitter to spread the reconnect wave.

### Failure Modes and Anti-Patterns

- **No backpressure.** A slow client whose send buffer keeps growing will exhaust server memory. Bound per-connection queues and disconnect, or drop non-critical messages, when they fill.
- **No resume.** After a reconnect the client has missed messages. Use sequence numbers and a replay buffer, or re-fetch state over REST.
- **Authentication only at connect.** A token checked once at handshake stays "valid" for hours. Re-validate periodically or close connections when sessions are revoked.
- **Cross-site WebSocket hijacking.** Browsers send cookies on the upgrade request, and WebSockets are not restricted by CORS. Check the `Origin` header.
- **Replacing all REST with WebSockets.** You lose caching, status codes, idempotency, and standard tooling, and you rebuild them badly.

## Server-Sent Events (SSE)

### When to Use It

SSE fits **server-to-client** streams where the client does not need to send much back: notifications, progress updates, live scores, dashboards, phase changes in a session, and token-by-token output from an LLM. The client sends anything it needs to through ordinary REST calls.

### How It Works

SSE is just an HTTP response that never ends. The server responds with `Content-Type: text/event-stream` and writes UTF-8 text events separated by blank lines:

```text
retry: 5000

id: 1041
event: phase
data: {"phase": "CODING", "endsAt": "2026-10-05T14:30:00Z"}

: heartbeat

id: 1042
event: timer
data: {"remainingSec": 120}
```

The browser's `EventSource` API parses this format, dispatches events by type, and **reconnects automatically** when the connection drops. On reconnect it sends a `Last-Event-ID` header with the last `id` it received, so the server can replay what was missed. The `retry:` field tells the browser how long to wait before reconnecting. Lines starting with `:` are comments, which make cheap heartbeats.

Because it is plain HTTP, SSE reuses your cookies, HTTP/2, proxies, and load balancers.

### Example: Session Event Stream

Server side (Node with Express):

```javascript
app.get("/v1/sessions/:id/events", async (req, res) => {
  res.set({
    "Content-Type": "text/event-stream",
    "Cache-Control": "no-cache",
    "X-Accel-Buffering": "no",          // tell nginx not to buffer
  });
  res.flushHeaders();

  const since = Number(req.get("Last-Event-ID") || 0);
  let lastSent = since, replaying = true;
  const pending = [];
  const deliver = (e) => { if (e.seq > lastSent) { lastSent = e.seq; send(res, e); } };

  // Subscribe before replaying so events published during the replay query are not lost.
  const unsubscribe = bus.subscribe(`session:${req.params.id}`,
    (e) => (replaying ? pending.push(e) : deliver(e)));
  const hb = setInterval(() => res.write(": hb\n\n"), 20000);
  req.on("close", () => { clearInterval(hb); unsubscribe(); });

  for (const e of await recentEvents(req.params.id, since)) deliver(e);
  pending.forEach(deliver);
  replaying = false;
});

function send(res, e) {
  res.write(`id: ${e.seq}\nevent: ${e.type}\ndata: ${JSON.stringify(e.data)}\n\n`);
}
```

Client side:

```javascript
const es = new EventSource("/v1/sessions/42/events", { withCredentials: true });
es.addEventListener("phase", (e) => showPhase(JSON.parse(e.data)));
es.addEventListener("timer", (e) => syncTimer(JSON.parse(e.data)));
```

### Scaling

SSE has the same stateful-connection profile as WebSockets, and the same answer: a pub/sub backbone so any instance can deliver any session's events, plus a small replay store (a Redis stream or a capped list per session) keyed by event ID so reconnects can resume. Heartbeats every 15 to 30 seconds keep load balancers and NAT devices from closing idle connections.

![Two long-lived connection styles through one layer 7 load balancer: an editor client exchanges WebSocket frames in both directions with a WebSocket gateway, while a session page opens an SSE stream with Last-Event-ID, served by an SSE endpoint that receives events through a pub/sub fan-out and keeps recent events for replay.](https://d5osvdbc8um23.cloudfront.net/static-asset/blog_images/api-and-realtime-protocols-compared/03-long-lived-websocket-vs-sse.png)

The HTTP version matters. Under **HTTP/1.1**, browsers limit concurrent connections per origin (six in major browsers), and each open `EventSource` permanently occupies one, so a user with several tabs open can starve their own REST calls. Under **HTTP/2**, each stream is multiplexed on one connection and the limit is the negotiated maximum concurrent streams, which the HTTP/2 specification recommends should be no smaller than 100. Serve SSE over HTTP/2 in production.

### Failure Modes and Anti-Patterns

- **Proxy buffering.** nginx and some CDNs buffer responses by default, so events arrive in bursts or not at all. Disable buffering on the SSE route (`proxy_buffering off;` or `X-Accel-Buffering: no`) and test through the full production path.
- **No custom headers in `EventSource`.** The standard API cannot set an `Authorization` header. Use cookie authentication, or a short-lived single-use token in the query string (and keep it out of access logs), or a `fetch`-based SSE client that reads the stream manually.
- **Binary data.** SSE is text only. Base64 works but inflates payloads by about a third; if you need binary, use WebSockets.
- **Events without IDs.** Without `id:` lines, reconnection resumes with no idea what was missed.
- **Treating SSE as reliable delivery.** Replay covers short disconnects. For state that must be correct, the client should re-fetch the authoritative state over REST after reconnecting.

## WebRTC: Signaling, STUN/TURN, and SFUs

### When to Use It

WebRTC is for **real-time audio and video** (calls, meetings, screen sharing, voice AI agents) and low-latency peer data where losing a message beats waiting for it. Nothing else here is built for interactive media.

### How It Works

**Signaling.** Endpoints exchange a **session description** (SDP: codecs, tracks, encryption fingerprints) and **ICE candidates** (possible addresses) before media flows. WebRTC does not specify how; you build the channel, usually a WebSocket. One side sends an **offer**, the other an answer, and candidates trickle in as discovered.

**NAT traversal with ICE, STUN, and TURN.** **ICE** gathers candidate addresses and tests pairs until one works. A **STUN** server tells a client the public IP and port its packets appear from, which is enough for many home networks. Behind restrictive NATs or UDP-blocking firewalls, a **TURN** server relays all media. TURN is expensive, since every byte crosses your bandwidth bill, but without it some users cannot connect. Offer TURN over TLS on port 443 for locked-down corporate networks.

**Media transport.** Media is encrypted with **SRTP**, keyed via **DTLS**; the WebRTC standards make encryption mandatory. Audio commonly uses the Opus codec and video uses VP8, H.264, or newer codecs. Media flows over UDP where possible, with jitter buffers, loss concealment, and congestion control. **Data channels** carry arbitrary messages over SCTP, reliable or unreliable, ordered or not.

### Topologies: Mesh, MCU, SFU

In a **mesh**, every participant sends its stream to every other participant. That works for two or three people and collapses quickly: with five participants each sending video at 1 Mbps, each participant uploads 4 Mbps, and home upload links are often the tightest constraint.

An **MCU** decodes, mixes, and re-encodes all streams into one, which is CPU-heavy and adds latency.

An **SFU** (selective forwarding unit) receives each participant's stream once and forwards it to the others without decoding. Each client uploads one stream and downloads what it needs. With **simulcast**, senders upload two or three quality layers and the SFU forwards the layer each receiver can handle. SFUs are the standard for group calls and server-side participants such as recorders or AI agents. An SFU decrypts hop by hop, so it can see media unless you add end-to-end encryption with encoded transforms.

### Example: Connecting to an SFU

```javascript
const pc = new RTCPeerConnection({
  iceServers: [
    { urls: "stun:stun.example.com:3478" },
    { urls: "turns:turn.example.com:443?transport=tcp",
      username: tempUser, credential: tempPass },   // short-lived credentials
  ],
});

const mic = await navigator.mediaDevices.getUserMedia({ audio: true });
mic.getTracks().forEach((t) => pc.addTrack(t, mic));

pc.onicecandidate = (e) => e.candidate && signaling.send({ type: "ice", candidate: e.candidate });
pc.ontrack = (e) => { remoteAudio.srcObject = e.streams[0]; };

const offer = await pc.createOffer();
await pc.setLocalDescription(offer);
signaling.send({ type: "offer", sdp: offer.sdp });
signaling.on("answer", (a) => pc.setRemoteDescription(a));
signaling.on("ice", (c) => pc.addIceCandidate(c.candidate));
```

Issue TURN credentials from your backend with a short expiry (a common pattern is an HMAC-based time-limited username and password) so leaked credentials cannot be used as a free relay.

![WebRTC signaling and media: Browser A exchanges SDP and ICE candidates with a signaling server over a WebSocket and asks a STUN server for its public address, then sends SRTP media to a selective forwarding unit that forwards it to Browser B, while Browser C behind a strict NAT reaches the SFU through a TURN relay.](https://d5osvdbc8um23.cloudfront.net/static-asset/blog_images/api-and-realtime-protocols-compared/04-webrtc-signaling-and-media.png)

### Scaling

Signaling is light; media is **bandwidth-bound**. Opus voice is commonly configured in the 24 to 64 kbps range; take 50 kbps per stream. A one-to-one call through an SFU has two streams in and two out, roughly 200 kbps of SFU traffic. At 5,000 concurrent calls that is about 1 Gbps through the SFU fleet, before TURN relays. Video is an order of magnitude more per stream. Scale SFUs horizontally by assigning each room to one SFU (or cascading SFUs for very large rooms), place them in regions close to users, and route each user to the nearest region.

Many teams use a hosted WebRTC provider instead of running SFU and TURN fleets, which is usually right when media infrastructure is not the differentiator: you buy NAT traversal and edge presence and pay per participant-minute.

### Failure Modes and Anti-Patterns

- **No TURN.** Everything works in the office and fails for users on strict corporate or mobile networks.
- **Ignoring ICE restarts.** A user switching from Wi-Fi to cellular changes address; without ICE restart handling the call just freezes.
- **WebRTC as a general data protocol.** For client-server messages, a WebSocket is simpler to operate and load balance.
- **Device issues.** Denied microphone permission or the wrong input device are real failures; run a device check before the session.

## Head-to-Head Comparison

| Dimension | REST | GraphQL | gRPC | WebSockets | SSE | WebRTC |
|---|---|---|---|---|---|---|
| Communication shape | Request/response | Request/response; subscriptions | Unary and streaming (all four modes) | Bidirectional stream | Server to client stream | Bidirectional media and data |
| Transport | HTTP/1.1, HTTP/2, HTTP/3 over TCP or QUIC | Usually HTTP | HTTP/2 | TCP after HTTP/1.1 upgrade | HTTP/1.1 or HTTP/2 | UDP preferred (SRTP, SCTP over DTLS), TCP/TLS via TURN |
| Payload format | Usually JSON | JSON | Protobuf (binary) | Text or binary frames | UTF-8 text | Encoded audio/video, arbitrary data |
| Contract and typing | OpenAPI (optional) | Schema (required) | `.proto` (required) | None built in | None built in | SDP negotiation; app data untyped |
| Browser support | Native | Native (via HTTP) | Needs gRPC-Web proxy | Native | Native (`EventSource`) | Native |
| HTTP caching | Excellent | Poor without persisted queries | None | None | None | None |
| Reconnect and resume | Not applicable | Not applicable | Retries per call; streams restart | Build it yourself | Built in (`Last-Event-ID`) | ICE restart; app-level session rejoin |
| Server state per client | None | None | Per stream while open | Per connection | Per connection | Per peer connection, media buffers |
| Load balancing | Any L4/L7 | Any L7 | L7 or client-side | Long-lived connections (L4, or L7 with Upgrade support) | Long-lived connections | Room-to-SFU assignment |
| Latency profile | One round trip per call | One round trip, resolver fan-out behind it | Low; multiplexed connections | Very low once open | Low once open | Lowest for media |
| Operational complexity | Low | Medium (cost limits, N+1, caching) | Medium (proto management, L7 LB) | Medium to high (fan-out, drain, backpressure) | Low to medium (buffering, replay) | High (TURN, SFU, NAT, codecs) |
| Typical workloads | Public APIs, CRUD, payments | Multi-client product APIs | Internal microservices, streaming pipelines | Chat, games, collaboration | Notifications, progress, LLM token streams | Calls, meetings, voice agents |

The table hides one important nuance: these are not mutually exclusive layers. GraphQL subscriptions run over WebSockets or SSE. WebRTC needs a signaling channel, which is usually a WebSocket. gRPC streams can feed the service that produces SSE events. The choice is per path, not per product.

## When to Pick Which

Start from the shape of the communication on each path, not from the protocol name. The flowchart below captures the order of questions that settles most cases.

![A decision flowchart: live audio or video leads to WebRTC; otherwise, if the server pushes updates, frequent client streaming leads to WebSockets and occasional client sends lead to SSE; without server push, internal service-to-service calls lead to gRPC, many clients with varied data needs lead to GraphQL, and everything else leads to REST.](https://d5osvdbc8um23.cloudfront.net/static-asset/blog_images/api-and-realtime-protocols-compared/05-decision-flowchart.png)

Media comes first because nothing else carries it; push comes next because it changes the connection model.

### Scenario 1: Ride-Hailing Driver Location on a Rider's Map

**Requirements**: a rider watches the assigned driver move on a map. The driver's app sends a location every few seconds; the rider's app only displays it.

**Choice**: the driver app sends locations with plain REST `POST`s (or a small batch every few seconds, which tolerates flaky mobile networks well), and the rider app receives them over **SSE** or a WebSocket. SSE is enough because the rider sends nothing on that path, and it reconnects automatically on a flaky network. If the same connection also carries chat between rider and driver, a WebSocket becomes the simpler single channel.

### Scenario 2: Internal Order Pipeline

**Requirements**: an order service calls inventory, pricing, and fraud services on every checkout. Latency budget for the whole call chain is tight, teams ship independently, and the services are in four languages.

**Choice**: **gRPC** with deadlines propagated across every hop, generated clients per language, and an L7 mesh or client-side load balancing. Asynchronous steps (emails, analytics) go to a queue, not a synchronous call. The public checkout API in front of it stays **REST**, with an idempotency key on the order `POST`.

### Scenario 3: Multi-Platform Content App

**Requirements**: web, iOS, and Android apps render feeds, profiles, and detail pages from the same entities, but each screen needs different fields, and mobile clients are slow to update.

**Choice**: a **GraphQL** gateway in front of internal gRPC or REST services, with DataLoader in every resolver, persisted queries for first-party apps, and cost limits for anything public. Live counters on a page (new comments) can be a GraphQL subscription over SSE or WebSockets, or a separate SSE stream.

## Worked Example: A Live AI Interview App

The product: a candidate talks to an AI interviewer by voice in the browser. The interview moves through timed phases (introduction, coding, system design, wrap-up), the candidate writes code in an in-browser editor, and the transcript and code are evaluated asynchronously afterward.

Assume a target of 2,000 concurrent sessions at peak, each about 45 minutes long.

### Voice: WebRTC Through an SFU or Hosted Provider

Voice is the only real-time media path, so it is WebRTC. The browser joins a room on an SFU or hosted provider. The AI interviewer is a **server-side participant**: an agent worker joins the room, runs speech-to-text on the candidate's audio, sends text to a language model, and publishes synthesized speech back.

Why not stream audio over a WebSocket? You can, but you then own jitter buffering, echo cancellation, loss concealment, and bitrate adaptation, which the browser's WebRTC stack provides, and TCP head-of-line blocking stalls audio behind each lost packet.

Back-of-envelope: two voice streams per session at about 50 kbps each, in and out of the SFU, is about 200 kbps per session; 2,000 sessions is about 400 Mbps of SFU traffic. That is modest, and far below what video would cost, so camera video can stay optional. The browser starts with `POST /v1/sessions`, which validates access, creates the room, returns a join token, and dispatches an agent worker; signaling uses the SFU's or provider's own channel.

### Phases and Timers: SSE

Phase changes and timer warnings flow server to browser at low frequency: SSE's sweet spot. The session service publishes to a per-session pub/sub channel, and whichever instance holds the `EventSource` connection forwards events. Three details matter:

1. **The server owns time.** Send the phase's absolute end time, not a tick per second; the browser counts down locally, which survives tab throttling.
2. **Events are replayable.** Keep recent events per session so `Last-Event-ID` replays what a briefly disconnected candidate missed.
3. **Re-fetch on reconnect.** The client reads the authoritative phase from `GET /v1/sessions/42`; replay is the fast path, REST the source of truth.

Why not the WebRTC data channel? It couples control events to the media path and pushes business logic into the agent or SFU. Why not a WebSocket? Nothing here needs client-to-server streaming; submit and end are REST calls.

### Code Editor: REST With Versioning

The editor autosaves with a debounced `PUT /v1/sessions/42/code` carrying `If-Match`, so two tabs cannot silently overwrite each other; final submission is a `POST` with an `Idempotency-Key`. Autosave load is easy to bound: at one save every few seconds per active editor, 2,000 sessions produce at most several hundred writes per second, well within a single relational primary.

If a human later watched or co-edited live, this path would move to a WebSocket with OT or CRDTs; for one author, REST is simpler.

### Backend Service Calls: REST or gRPC, Plus a Queue

Internal choices:

- **Agent worker to session service**: **gRPC** fits, with a unary call for the interview plan, a client stream for transcript segments, and deadlines everywhere. REST with JSON is a fine choice for a small team.
- **Session service to question service**: unary gRPC or REST with a short timeout and a cached fallback.
- **Session end to evaluation**: not synchronous. The session service writes the completion and an outbox event in one transaction; a relay publishes to a queue; evaluation workers score it and write results, which the UI fetches over REST.

GraphQL is optional here: one web client does not need it, but a later mobile app and dashboard slicing the same data differently might.

### Putting the Failure Modes Together

- **Network drops for 20 seconds.** WebRTC restarts ICE, SSE reconnects and replays, the client re-fetches state, and unsaved code is retried with its version precondition. The agent pauses instead of talking to an empty room.
- **An SSE instance is deployed.** Clients reconnect elsewhere with `Last-Event-ID`; no instance is special.
- **The agent worker crashes.** The session service sees the missing heartbeat and starts a new worker with the plan and transcript so far.
- **Evaluation is slow.** The queue absorbs it; the live session never depends on it.

This design uses four of the six protocols (WebRTC, SSE, REST, and gRPC), each for one stated reason.

## In the Interview

### How to Justify a Choice

Use the same four steps for every path in your design:

1. **Name the path and its direction**: "Phase events flow server to client, a few per minute per session."
2. **Name the latency and reliability need**: "A second of delay is fine; missing an event is not, because the UI would show the wrong phase."
3. **Name the protocol and the mechanism that meets the need**: "SSE, with event IDs and `Last-Event-ID` replay, plus a REST re-fetch after reconnect."
4. **Name the cost you accept**: "Long-lived connections mean we scale on concurrent connections and need a pub/sub backbone; we serve it over HTTP/2 to avoid the browser's per-origin connection limit."

Step four is the one most candidates skip.

### Common Traps

- **"WebSockets for everything real time."** Many real-time needs are one-directional, where SSE is simpler, or fine with polling.
- **"gRPC is faster, so use it publicly."** Browsers need gRPC-Web and a proxy, and at the edge network latency dominates anyway.
- **"GraphQL solves over-fetching" with no mention of N+1, caching, or cost limits.** Raise them yourself.
- **Forgetting reconnects.** Every streaming protocol drops connections. Say what the client does on reconnect and how it catches up.
- **Forgetting the load balancer.** Mention idle timeouts, L7 balancing for gRPC, and draining on deploy.
- **WebRTC without TURN.** Saying "peer to peer" without NAT traversal and relays is half an answer.
- **Unsourced numbers.** Derive capacity from stated requirements ("2,000 sessions at 200 kbps is 400 Mbps") rather than quoting benchmarks you cannot defend.

### Likely Follow-Up Questions

- **"How do you scale WebSockets to a million connections?"** Shard connections across many gateway instances, scale on memory and file descriptors, use a pub/sub backbone for cross-instance fan-out, and drain gradually on deploy with client-side jittered reconnects.
- **"SSE or WebSockets for notifications?"** SSE, unless the client must also stream to the server. SSE gives automatic reconnect and resume, works over HTTP/2, and reuses HTTP authentication and tooling.
- **"Why does gRPC need special load balancing?"** HTTP/2 multiplexes many calls on one long-lived connection, so per-connection L4 balancing pins a client to one backend. Balance per call at L7 or on the client.
- **"How do you version a gRPC API?"** Add fields with new numbers, never reuse or retype a number, reserve removed ones, and create a new package version (`v2`) only for breaking changes.
- **"What if TURN costs explode?"** Measure what fraction of sessions relay, make sure direct and STUN paths are tried first, place TURN in-region, and price it in; some users can only connect through it.
- **"How would you add live typing visibility for a human observer?"** Move the editor path to a WebSocket carrying operations with sequence numbers, using operational transforms or a CRDT if more than one person can edit.
- **"How do you handle auth on long-lived connections?"** Authenticate at connect, bind to a revocable session, re-check periodically, and close on expiry.

Related reading: [Designing a Chat Platform like Discord](/blog/system-design-discord), [Designing a Collaborative Editor like Google Docs](/blog/system-design-collaborative-editor), and [Message Queues vs Logs vs Pub/Sub](/blog/message-queues-vs-logs-vs-pubsub).
