# Load Balancers, Reverse Proxies, API Gateways, and Service Meshes: A Complete Guide

## Blog Details

- **Author**: Naveen R.
- **Date**: October 7, 2026
- **Tags**: load balancing, system design, API gateway, service mesh, reverse proxy
- **Read Time**: 20 mins

## Introduction

Draw any system design on a whiteboard and there is a box between the client and the service. Candidates usually label it "load balancer" and move on. Interviewers rarely let them. Does that box work at the connection level or the request level? Does it terminate TLS? How does the service learn the caller's real IP? What happens to in-flight requests during a deploy? Who checks the JWT, who enforces the rate limit, and who retries when the payment service times out?

The answers depend on which component sits in the box. L4 load balancers, L7 load balancers, reverse proxies, API gateways, and service meshes all forward traffic, but they see different things, cost different amounts, and fail in different ways. Production systems often run three or four of them in series.

This post follows one request, a mobile app placing an order, from the edge to an order service and on to a payment service. For each layer it covers when to use it, how it works, real configuration, scaling, and failure modes. Then it compares them head to head, gives a decision flowchart with three scenarios, and closes with the follow-up questions interviewers ask most.

![The request path for a mobile order: the app resolves the API hostname through latency-based DNS, connects over HTTPS to a nearby edge location, passes through an API gateway that enforces auth, quotas, and validation, reaches an application load balancer that routes the orders path to the order service, which calls the payment service over mutual TLS and writes to the orders database while a mesh control plane pushes configuration and certificates to its sidecar.](https://d5osvdbc8um23.cloudfront.net/static-asset/blog_images/load-balancers-gateways-proxies/01-high-level-architecture.png)

## What Sits Between Client and Service

### The Running Example

The mobile app sends `POST https://api.shop.example/v2/orders` with a JWT and an API key identifying the app build. The path:

1. **DNS** resolves `api.shop.example` to the closest healthy region.
2. **Edge**: TCP and TLS terminate at a nearby edge location (a CDN or anycast front door), keeping handshake round trips short on a mobile network.
3. **API gateway** validates the JWT, checks the API key against a usage plan, enforces a rate limit, validates the body, and maps `/v2/orders` to a backend.
4. **L7 load balancer** inside the VPC routes `/orders/*` to the order service and spreads requests across healthy instances.
5. **Service mesh**: the order service's call to the payment service leaves through a local sidecar proxy that adds mutual TLS, a timeout, safe retries, and metrics.

### North-South and East-West

**North-south** traffic enters from outside; **east-west** traffic flows between services. Load balancers and gateways are mostly north-south, meshes east-west, and NGINX and Envoy appear in both.

### Global Load Balancing Before the First Packet

Something has to pick a region first. **DNS-based** global load balancing returns different addresses for the same name based on resolver location, latency, weights, or health; Amazon Route 53 offers latency, geolocation, weighted, and failover policies. Its weakness is caching: resolvers hold answers for the TTL and some ignore it, so failover is only as fast as the slowest cache. **Anycast** advertises one IP from many locations over BGP, and internet routing delivers each packet to the nearest one, so failover happens in routing rather than DNS caches. AWS Global Accelerator and several large CDNs, such as Cloudflare and Fastly, use anycast at the edge.

## L4 Load Balancers: NLB

### When to Reach for L4

A Layer 4 load balancer works on **connections**. It sees IPs, ports, and the protocol (TCP or UDP), but does not parse HTTP, so it cannot route on paths, headers, or cookies. In exchange it is protocol-agnostic and does very little work per packet. Reach for it when you need:

- **Non-HTTP protocols**: raw TCP (databases, MQTT, custom binary), UDP, or TLS you want passed through untouched.
- **Static IPs**: AWS Network Load Balancer (NLB) gives one static IP per Availability Zone and can use Elastic IPs, which matters when partners allowlist your addresses.
- **TLS passthrough**, where the balancer never holds the private key.
- **A front door for your own proxies**: an NLB in front of an Envoy or NGINX fleet (or a Kubernetes ingress controller) supplies static IPs and spreads connections while the proxies do L7 work.

### How It Works

When a client opens a TCP connection, NLB picks a target with a flow hash (over protocol, source and destination IP and port, and TCP sequence number) and forwards every packet of that flow to the same target. The decision is made once per connection, so **every request on a long-lived connection goes to the same backend**.

That is harmless for short HTTP/1.1 connections and a real problem for gRPC, which multiplexes many requests over one long-lived HTTP/2 connection. Ten clients with one connection each, through an L4 balancer to twenty backends, keep at most ten backends busy. The fixes are an L7 balancer that balances HTTP/2 streams, client-side load balancing (the gRPC client resolves all backends and balances per request), or a maximum connection age on the server to force periodic reconnection.

### TLS Termination vs Passthrough

A **TCP listener** passes encrypted bytes straight through; the target holds the certificate and the balancer sees nothing. A **TLS listener** decrypts with a certificate from AWS Certificate Manager and opens a new connection to the target. Termination centralizes certificates and offloads handshakes; passthrough keeps the key on the backend for true end-to-end encryption or client certificate checks. Re-encryption (terminate, then TLS again to the backend) gives inspection plus encryption on the wire, at the cost of a second handshake.

### Client IP Preservation and Proxy Protocol

An L4 balancer cannot add `X-Forwarded-For`, so it uses one of two mechanisms:

- **Client IP preservation**: packets reach the target with the client's source IP. On NLB it is on by default for instance targets, and for IP targets the default depends on protocol, so check the `preserve_client_ip.enabled` attribute rather than assuming. With it on, target security groups must allow client addresses, not just the balancer's.
- **Proxy protocol**: a header, defined by HAProxy, prepended to the TCP stream with the original addresses. Version 1 is text, version 2 binary; NLB supports version 2. The backend must expect it or it will reject the connection.

### Concrete Config: NLB in Front of the Ingress Proxies

```yaml
IngressTargetGroup:
  Type: AWS::ElasticLoadBalancingV2::TargetGroup
  Properties:
    Protocol: TCP
    Port: 8443
    TargetType: ip
    VpcId: !Ref Vpc
    HealthCheckProtocol: HTTP
    HealthCheckPort: "8081"                # plain HTTP; NLB health checks send no TLS or proxy header
    HealthCheckPath: /ready
    TargetGroupAttributes:
      - Key: proxy_protocol_v2.enabled
        Value: "true"
      - Key: preserve_client_ip.enabled
        Value: "false"
      - Key: deregistration_delay.timeout_seconds
        Value: "60"
      - Key: load_balancing.cross_zone.enabled
        Value: "true"
```

### Scaling and Cross-Zone Load Balancing

NLB runs a node in each enabled zone. With **cross-zone load balancing** off, each node sends only to targets in its own zone: if zone A has two targets and zone B eight, each zone A target takes a quarter of all traffic. With it on, load evens out but traffic crosses zones, and AWS bills that inter-zone data on NLB. The defaults differ: cross-zone is **off by default on NLB** and **on by default on ALB**. Keep target counts balanced across zones either way.

NLB's TCP idle timeout was historically fixed at 350 seconds and is now configurable per listener. Clients holding idle connections longer need TCP keepalives below it or they see silent resets.

### Failure Modes and Anti-Patterns

1. **gRPC behind L4 with no client-side balancing**: a few backends take all the load.
2. **Proxy protocol enabled on one side only**: every connection fails to parse. Change both sides together.
3. **Security groups allowing only balancer IPs** while client IP preservation is on: health checks pass, real traffic drops.
4. **Unbalanced zones with cross-zone off**.
5. **Expecting L7 features**: no path routing, header injection, or per-request retries.

## L7 Load Balancers: ALB, Path Routing, and Sticky Sessions

### When to Reach for L7

A Layer 7 load balancer terminates HTTP, parses each request, and decides **per request**. That unlocks routing on host, path, headers, query string, and method; per-request balancing across HTTP/2 connections; header injection; and HTTP-aware health checks. AWS Application Load Balancer (ALB) is the managed example; Google Cloud's Application Load Balancer and Azure Application Gateway fill the same role. For our flow, ALB routes `/orders/*` to the order service and `/track/*` to the tracking service.

![An L4 load balancer pins each TCP connection from the mobile app to one ingress proxy for the life of the connection, while an L7 load balancer parses every request and routes the orders path to the order service and the tracking path to the tracking service.](https://d5osvdbc8um23.cloudfront.net/static-asset/blog_images/load-balancers-gateways-proxies/02-l4-vs-l7.png)

### How It Works

An ALB terminates client connections and keeps its own reusable pool of connections to targets. Each request is matched against listener rules in priority order; the first match picks a target group, whose algorithm picks a target. ALB offers **round robin** (default), **least outstanding requests**, and **weighted random** (optionally with automatic target weights that shift traffic away from targets returning anomalous errors). Least outstanding requests helps when request cost varies, as with an order service where some calls hit a cache and others run a multi-table transaction.

ALB adds `X-Forwarded-For`, `X-Forwarded-Proto`, and `X-Forwarded-Port`, supports WebSockets, and supports gRPC end to end with a target group whose protocol version is GRPC.

### X-Forwarded-For Done Right

`X-Forwarded-For` is a list to which each proxy appends the address it received the connection from. The leftmost entry is whatever the client claimed, so it can be forged. Find the real client by counting from the **right**, skipping the hops you control: behind edge, gateway, and ALB, trust only that many rightmost entries. Using the leftmost value for rate limiting or fraud checks is a classic bug. The standardized `Forwarded` header (RFC 7239) carries the same data in structured form but is less common.

### Health Checks and Connection Draining

An ALB health check is an HTTP request to a path you choose, with configurable interval, timeout, and thresholds. It should report **readiness** of the instance, not the health of every dependency. If the check calls the payment service and payment blips, every order instance fails at once and the balancer has nothing to route to. Check what is local; let dependency failures surface as errors callers can handle.

**Connection draining**, which AWS calls the **deregistration delay**, governs a target being removed in a deploy or scale-in: no new requests, but in-flight ones get up to the delay to finish. The default is 300 seconds. For an order API whose requests finish in under a second, that slows every deploy for nothing; 30 seconds is plenty. For WebSockets the delay is a ceiling after which connections are cut, so clients must reconnect gracefully regardless. **Slow start** is the mirror image: a new target's share ramps up over a window, protecting JVM services that need to warm caches. Note that ALB does not allow slow start together with the least outstanding requests algorithm, so the config below uses the latter.

### Concrete Config: Path Routing and Stickiness

```yaml
OrdersRule:
  Type: AWS::ElasticLoadBalancingV2::ListenerRule
  Properties:
    ListenerArn: !Ref HttpsListener
    Priority: 10
    Conditions:
      - Field: path-pattern
        PathPatternConfig:
          Values: ["/orders", "/orders/*"]
    Actions:
      - Type: forward
        TargetGroupArn: !Ref OrdersTargetGroup

OrdersTargetGroup:
  Type: AWS::ElasticLoadBalancingV2::TargetGroup
  Properties:
    Protocol: HTTP
    ProtocolVersion: HTTP2
    Port: 8080
    TargetType: ip
    VpcId: !Ref Vpc
    HealthCheckPath: /ready
    TargetGroupAttributes:
      - Key: load_balancing.algorithm.type
        Value: least_outstanding_requests
      - Key: deregistration_delay.timeout_seconds
        Value: "30"
      - Key: stickiness.enabled
        Value: "false"
```

### Sticky Sessions and Their Trade-Offs

Sticky sessions pin a client to one target. ALB supports a **duration-based cookie** it generates (`AWSALB`) and an **application-based cookie** your service sets; NLB supports stickiness by source IP. Stickiness helps when a target holds expensive per-client state such as a warmed cache. The costs:

- **Uneven load**: heavy clients pin to a few targets that least-outstanding-requests can no longer relieve.
- **Scaling lag**: new targets only receive new clients.
- **Lost state** when a target dies.
- **NAT pile-ups** with source IP stickiness, since many mobile users share a carrier NAT address.

The better default is stateless instances with session state in Redis or DynamoDB. Stickiness is an optimization, never the source of correctness.

### Scaling and Failure Modes

ALB scales itself, but not instantly; a step change such as a flash sale launch can outrun it, and ALB now offers capacity reservation for planned events. Its IP addresses change as it scales, so point DNS at it with an alias record and never hard-code IPs.

1. **Idle timeout mismatches**: ALB's idle timeout defaults to 60 seconds. If the backend's keepalive timeout is shorter, the backend can close a connection just as the ALB reuses it, causing intermittent 502s. Make the backend's timeout longer.
2. **Deep health checks** that take the whole fleet out during a dependency outage.
3. **Trusting the leftmost X-Forwarded-For entry**.
4. **Rule sprawl** on one listener, which also runs into documented rule quotas.
5. **Stickiness as architecture**, which turns every deploy into a small outage for pinned users.

## Reverse Proxies: NGINX and Envoy

### When to Reach for a Self-Hosted Proxy

Managed load balancers are reverse proxies too, but in an interview "reverse proxy" usually means NGINX, HAProxy, or Envoy running on your own instances or as Kubernetes ingress. Choose one when you need what the managed balancer lacks: response caching, buffering, custom routing, cluster-edge rate limiting, identical config across clouds, or deep per-route telemetry.

### NGINX

NGINX runs a few worker processes, each an event loop over non-blocking sockets, so one worker holds many thousands of connections. Configuration is a static file; `nginx -s reload` starts new workers on the new config while old ones finish their connections. Upstream algorithms include round robin (default), `least_conn`, `ip_hash`, `hash` with an optional `consistent` flag, and `random two least_conn`. Open source NGINX does **passive** health checks (`max_fails`, `fail_timeout`); **active** checks are a commercial NGINX Plus feature.

For our order API, an NGINX tier behind the NLB absorbs slow mobile clients and caches the catalog:

```nginx
upstream order_service {
    least_conn;
    server 10.0.1.10:8080 max_fails=3 fail_timeout=10s;
    server 10.0.2.10:8080 max_fails=3 fail_timeout=10s;
    keepalive 64;
}

server {
    listen 8081;                            # health check port for the NLB
    location = /ready { return 200; }
}

proxy_cache_path /var/cache/nginx keys_zone=catalog:50m max_size=2g inactive=10m;
limit_req_zone $binary_remote_addr zone=per_ip:10m rate=20r/s;

server {
    listen 8443 ssl proxy_protocol;         # NLB in front sends proxy protocol v2
    ssl_certificate     /etc/nginx/tls/api.crt;
    ssl_certificate_key /etc/nginx/tls/api.key;
    set_real_ip_from 10.0.0.0/16;           # trust only our NLB subnet
    real_ip_header proxy_protocol;

    location /orders {
        limit_req zone=per_ip burst=40 nodelay;
        proxy_pass http://order_service;
        proxy_http_version 1.1;
        proxy_set_header Connection "";
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
        proxy_connect_timeout 1s;
        proxy_read_timeout 5s;
    }

    location /catalog/ {
        proxy_cache catalog;
        proxy_cache_valid 200 60s;
        proxy_cache_use_stale error timeout updating;
        proxy_pass http://order_service;
    }
}
```

### Caching and Buffering at the Proxy

- **Buffering**: with `proxy_buffering` on (the NGINX default), the proxy reads the upstream response quickly and trickles it to a slow client, freeing the order service's worker instead of tying it to a 3G connection. Streaming responses (server-sent events, long polling) need buffering off on those routes or clients see nothing until the buffer fills.
- **Caching**: a short-lived cache for anonymous, idempotent responses collapses many identical requests into one upstream fetch, and `proxy_cache_use_stale` turns a backend outage into stale data instead of errors. Never cache responses that depend on `Authorization` unless the user is in the cache key.

### Envoy

Envoy, built at Lyft and now a CNCF graduated project, is an L4 and L7 proxy designed for dynamic environments. Three things set it apart:

1. **Dynamic configuration through xDS**: listeners (LDS), routes (RDS), clusters (CDS), endpoints (EDS), and secrets (SDS) stream from a management server over gRPC, so routes and endpoints change with no reload. This is how a mesh control plane reprograms thousands of proxies as pods come and go.
2. **Observability by default**: per-cluster and per-route stats (latency histograms, retries, circuit breaker trips, ejections), tracing header propagation, and structured access logs. Much of what people credit to "the mesh" is Envoy's telemetry.
3. **Resilience primitives**: per-route timeouts and retries, retry budgets, circuit breakers on connections and pending or concurrent requests, and outlier detection that ejects hosts returning consecutive 5xx.

```yaml
clusters:
  - name: order_service
    type: EDS
    eds_cluster_config:
      eds_config: { ads: {} }          # endpoints pushed by the control plane
    lb_policy: LEAST_REQUEST            # power of two choices by default
    typed_extension_protocol_options:
      envoy.extensions.upstreams.http.v3.HttpProtocolOptions:
        "@type": type.googleapis.com/envoy.extensions.upstreams.http.v3.HttpProtocolOptions
        explicit_http_config: { http2_protocol_options: {} }
    circuit_breakers:
      thresholds:
        - max_connections: 1000
          max_pending_requests: 200
          max_requests: 1000
          max_retries: 50
    outlier_detection:
      consecutive_5xx: 5
      interval: 10s
      base_ejection_time: 30s
      max_ejection_percent: 50

routes:
  - match: { path_separated_prefix: "/orders" }
    route:
      cluster: order_service
      timeout: 3s
      retry_policy:
        retry_on: "connect-failure,refused-stream,unavailable"
        num_retries: 1
        per_try_timeout: 1s
```

Connect failures and refused streams mean the request never reached application code; gRPC `unavailable` usually means the same but is not guaranteed, so non-idempotent calls still need an idempotency key. Retrying `POST /orders` on any 5xx risks a duplicate order unless the API is idempotent.

### NGINX vs Envoy

NGINX is simpler, battle-tested as a web server, strong at static content, caching, and buffering, and configured by files humans edit. Envoy is built to be configured by machines, changes live through xDS, speaks HTTP/2 and gRPC natively on both sides, and emits richer telemetry. One team hand-managing an edge tier often picks NGINX; a control plane managing hundreds of proxies usually means Envoy.

### Failure Modes and Anti-Patterns

1. **A single proxy instance**: run at least two per zone behind an L4 balancer.
2. **Retries at every layer** (see the retry storm section).
3. **Caching personalized responses** with a key that omits the user.
4. **Buffering streaming endpoints**, so server-sent events appear to hang.
5. **Hand-edited config drift** across a fleet: generate config from one source or use xDS.

## API Gateways: Auth, Rate Limits, and Transforms

### When to Reach for a Gateway

An API gateway is an L7 proxy specialized for **API management**. A load balancer asks which healthy instance gets the request. A gateway asks who the caller is, whether they are allowed, whether they are over quota, whether the request is well-formed, and which backend version should handle it. Add one when you expose APIs to mobile apps, partners, or third-party developers and want those concerns handled once instead of in every service.

![Inside an API gateway, a request from the mobile app carrying a JWT and an API key passes through authentication and authorization, then usage plan and rate limit checks, then request validation, then transformation and version routing, before a VPC link forwards it to the internal load balancer, which sends the v1 path to order service v1 and the v2 path to order service v2.](https://d5osvdbc8um23.cloudfront.net/static-asset/blog_images/load-balancers-gateways-proxies/03-api-gateway-responsibilities.png)

### Responsibilities

- **Authentication and authorization**: validate JWTs (signature, issuer, audience, expiry) or call a custom authorizer. Amazon API Gateway supports IAM, Cognito user pools, Lambda authorizers, and native JWT authorizers on HTTP APIs. Coarse checks ("token has `orders:write`") belong here; fine-grained ones ("this user owns this order") belong in the service that has the data.
- **Rate limiting**: token-bucket limits per key, route, or client, protecting backends and enforcing commercial tiers.
- **API keys and usage plans**: identify the calling application and attach a quota. API keys identify, they do not authenticate; AWS's documentation says not to rely on them alone for authorization.
- **Request validation** against a schema before the request costs backend compute.
- **Transformation**: rename fields, add or strip headers, or convert protocols.
- **Versioning and routing**: `/v1/orders` to the old service, `/v2/orders` to the new, or canary a share of traffic to a new stage.

### Managed: Amazon API Gateway

Amazon API Gateway offers **REST APIs** (usage plans, API keys, request validation, mapping templates, caching, private endpoints), **HTTP APIs** (cheaper and simpler, native JWT authorizers, fewer features), and **WebSocket APIs**. Documented limits that shape designs: 10 MB payloads and a default 29-second integration timeout. For REST APIs, AWS now allows raising that timeout for Regional and private APIs through a quota increase, possibly at the cost of a lower throttle limit. Accounts also have a default Region-level throttle quota shared across APIs; check yours rather than assume.

A free-tier usage plan for the mobile app:

```bash
aws apigateway create-usage-plan \
  --name mobile-free \
  --throttle burstLimit=20,rateLimit=10 \
  --quota limit=10000,period=DAY \
  --api-stages apiId=a1b2c3d4e5,stage=prod

aws apigateway create-usage-plan-key \
  --usage-plan-id <plan-id> --key-id <api-key-id> --key-type API_KEY
```

Managed buys zero servers and built-in scaling; you give up portability, deep customization, and cost control at very high volume.

### Self-Hosted: Kong and Envoy-Based Gateways

**Kong** is built on NGINX and OpenResty and adds plugins for auth, rate limiting, transformations, and logging:

```yaml
_format_version: "3.0"
services:
  - name: order-service-v2
    url: http://orders-v2.internal:8080
    routes:
      - name: orders-v2
        paths: ["/v2/orders"]
    plugins:
      - name: jwt
      - name: rate-limiting
        config:
          minute: 60
          policy: redis
          redis:
            host: ratelimit.internal
      - name: request-transformer
        config:
          add:
            headers: ["X-Api-Version:2"]
```

**Envoy-based gateways** (Envoy Gateway, which implements the Kubernetes Gateway API, plus products such as Gloo and Emissary) put a gateway control plane over Envoy.

### Rate Limiting Is a Distributed Problem

One gateway node can count in memory. Twenty nodes either each enforce a twentieth of the limit (inaccurate when traffic is uneven) or share counters in Redis (a network call per request and a new dependency). Decide whether a limit must be exact, like a paid quota, or approximate, like abuse protection, and say so.

### Failure Modes and Anti-Patterns

1. **Business logic in the gateway**: transformations grow into order rules, and gateway deploys become application deploys.
2. **API keys as authentication**: keys are trivially extracted from mobile binaries.
3. **Gateway as the only defense**: internal callers can bypass it, so services still validate input and ownership.
4. **Ignoring the integration timeout**: long work must be asynchronous (return 202 with an order ID, then poll or push).
5. **Undecided failure mode for the limiter store**: choose fail open or fail closed in advance.

## Service Mesh: Sidecars and mTLS

### When to Reach for a Mesh

With dozens of services, every team re-implements service-to-service TLS, retries, timeouts, circuit breakers, metrics, and tracing. A mesh moves those concerns into the network layer, uniformly, and lets a platform team change policy without redeploying services. Order calling payment is the canonical case: encrypted and authenticated both ways, a short timeout, retries only when safe, and visible on a dashboard.

### Data Plane and Control Plane

The **data plane** is the proxies carrying traffic: Envoy in Istio, a lightweight Rust proxy in Linkerd. The **control plane** pushes desired configuration to every proxy. Istio's `istiod` turns Kubernetes resources (VirtualService, DestinationRule, PeerAuthentication) into xDS and acts as a certificate authority for short-lived workload certificates.

![In a sidecar mesh, traffic from the mesh ingress gateway reaches the Envoy sidecar in the order pod over mutual TLS, the sidecar forwards to the order container on localhost, outbound calls leave through the same sidecar with mutual TLS and retries to the payment pod's sidecar, the istiod control plane pushes xDS configuration and certificates to both sidecars, and the sidecars emit metrics and traces to a telemetry backend.](https://d5osvdbc8um23.cloudfront.net/static-asset/blog_images/load-balancers-gateways-proxies/04-service-mesh-sidecars.png)

### Sidecar vs Sidecarless (Ambient) Modes

In the **sidecar** model each pod gets a proxy container, and iptables rules or a CNI plugin redirect its traffic through it. The app speaks plain HTTP to localhost. Costs: a proxy per pod, two extra proxy hops per call, mesh upgrades that require pod restarts, and startup races if the app starts before its sidecar.

**Sidecarless** designs reduce that. Istio's **ambient mode**, generally available since Istio 1.24, uses a per-node **ztunnel** for L4 work (mTLS, identity, L4 authorization) and optional **waypoint proxies** (Envoy) for L7 features only where needed. Cilium offers an eBPF-based mesh with per-node proxies. The trade is lower resource use and easier upgrades against a shared per-node component and a shorter operational track record.

### mTLS and Workload Identity

With mutual TLS both sides present certificates. The mesh issues each workload a certificate encoding its identity (in Istio, a SPIFFE identity from namespace and service account), rotates it automatically, and lets you write policy against identities instead of pod IPs that change constantly:

```yaml
apiVersion: security.istio.io/v1
kind: PeerAuthentication
metadata:
  name: default
  namespace: payments
spec:
  mtls:
    mode: STRICT
---
apiVersion: security.istio.io/v1
kind: AuthorizationPolicy
metadata:
  name: allow-order-to-charge
  namespace: payments
spec:
  selector:
    matchLabels: { app: payment }
  rules:
    - from:
        - source:
            principals: ["cluster.local/ns/orders/sa/order-service"]
      to:
        - operation:
            methods: ["POST"]
            paths: ["/charge"]
```

Roll out in `PERMISSIVE` mode (accept plaintext and mTLS), confirm every caller is meshed, then switch to `STRICT`.

### Retries, Timeouts, and Circuit Breaking

```yaml
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
  name: payment
  namespace: payments
spec:
  hosts: ["payment.payments.svc.cluster.local"]
  http:
    - route:
        - destination: { host: payment.payments.svc.cluster.local }
      timeout: 2s
      retries:
        attempts: 1
        perTryTimeout: 900ms
        retryOn: connect-failure,refused-stream,unavailable
---
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata:
  name: payment
  namespace: payments
spec:
  host: payment.payments.svc.cluster.local
  trafficPolicy:
    connectionPool:
      http: { http2MaxRequests: 500, maxRequestsPerConnection: 100 }
    outlierDetection:
      consecutive5xxErrors: 5
      interval: 10s
      baseEjectionTime: 30s
      maxEjectionPercent: 50
```

Charging is not naturally idempotent, so retries cover only failures where the request never reached the payment container, and the order service sends an idempotency key so even a duplicate charges once. Outlier detection acts as a host-level circuit breaker, ejecting a payment pod that returns repeated 5xx.

### The Cost of the Extra Hop

Each proxy adds latency, CPU, and memory. Per hop the latency is usually small next to network and application time, but a sidecar mesh adds two proxies per call, so a chain of five service-to-service calls passes ten. Measure it. The larger costs are operational: another control plane to upgrade, certificate expiry as a failure mode, and debugging across two sets of logs.

### Failure Modes and Anti-Patterns

1. **A mesh for three services**: a shared client library is cheaper. Meshes pay off at scale and across languages.
2. **Misunderstanding control plane outages**: proxies keep their last config, so traffic flows, but new pods and certificate rotation break.
3. **Mesh retries stacked on application retries**.
4. **Strict mTLS before every caller is meshed**: batch jobs and legacy VMs suddenly fail.
5. **Treating mTLS as user authorization**: it proves which workload called, not which user.

## Load-Balancing Algorithms: Round-Robin, Least Connections, Consistent Hashing

### Round-Robin and Weighted Round-Robin

Round robin sends each request or connection to the next backend in turn. It is simple, needs no shared state, and is fair when requests and backends are uniform. **Weighted** round robin gives bigger instances more turns and is how canaries work: weight 95 to stable, 5 to canary. Its weakness is ignoring current load, so a slow backend keeps its full share and builds a queue.

### Least Connections and Least Outstanding Requests

**Least connections** picks the backend with the fewest open connections; **least outstanding requests** (Envoy calls it least request) counts in-flight requests. Both adapt to slow backends, which accumulate work and receive less. They need per-backend state, which is exact on one balancer node and approximate across many, since each node sees only its own traffic.

### Power of Two Choices

Scanning every backend per request is expensive, and when many balancer nodes act on the same stale load data they cause **herding**, where many balancer nodes see the same least-loaded backend and all pile onto it. **Power of two choices** samples two backends at random and picks the less loaded. A well-known result in randomized load balancing is that this sharply reduces maximum load compared with one random choice, and different nodes sample different pairs, so herding disappears. Envoy's least request policy uses it by default; NGINX offers `random two least_conn`.

```python
import random

def pick_backend(backends, in_flight):
    a, b = random.sample(backends, 2)
    return a if in_flight[a] <= in_flight[b] else b
```

### Consistent Hashing and Maglev

Sometimes the **same key must reach the same backend**: a cache shard for a product ID, a user's WebSocket session, or in-memory rate limit counters. `hash(key) % N` reshuffles almost every key when `N` changes. **Consistent hashing** places backends (with many virtual points each) and keys on a ring; a key maps to the next backend clockwise, so adding or removing one moves only about `1/N` of keys.

**Maglev**, from Google's 2016 paper on its network load balancer, builds a fixed-size lookup table from per-backend permutations, giving single-index lookups, very even load, and small disruption on backend changes. Envoy supports both `RING_HASH` and `MAGLEV`. Both give affinity without cookies but risk hot spots on very popular keys.

```python
import bisect, hashlib

class Ring:
    def __init__(self, nodes, vnodes=100):
        self.points = sorted(
            (int(hashlib.md5(f"{n}#{i}".encode()).hexdigest(), 16), n)
            for n in nodes for i in range(vnodes)
        )
        self.keys = [p for p, _ in self.points]

    def lookup(self, key):
        h = int(hashlib.md5(key.encode()).hexdigest(), 16)
        i = bisect.bisect(self.keys, h) % len(self.points)
        return self.points[i][1]
```

### Retry Storms and Timeout Budgets

Algorithms decide where traffic goes; retries decide how much there is. If the mobile client, gateway, order service, and sidecar each make up to three attempts, one failure at payment becomes `3 x 3 x 3 x 3 = 81` attempts at the bottom, arriving exactly when payment can least cope. That retry storm turns a partial outage into a total one. Defenses:

- **Retry at one layer**, ideally the one nearest the failure that knows whether the operation is idempotent.
- **Retry budgets** that cap retries as a share of normal traffic; Envoy and gRPC both support this style of limit.
- **Exponential backoff with jitter** so clients do not retry in lockstep.
- **Idempotency keys** on `POST /orders` and `/charge`.
- **Deadline propagation**: each timeout shorter than its caller's. If the app waits 10 seconds, the gateway allows 8, the order service gives payment 2 seconds with one 900 ms retry, and payment's database call gets a few hundred milliseconds. gRPC carries the deadline on the wire, and passing the incoming context to outbound calls propagates it; for HTTP, pass the remaining budget in a header. An inner timeout longer than the outer one wastes work the caller already abandoned.

## Head-to-Head Comparison

| Dimension | L4 LB (NLB) | L7 LB (ALB) | Reverse proxy (NGINX, Envoy) | API gateway | Service mesh |
|---|---|---|---|---|---|
| Unit of decision | Connection (flow) | Request | Request (or connection in L4 mode) | Request plus caller identity | Request, per service-to-service call |
| Protocols | TCP, UDP, TLS | HTTP/1.1, HTTP/2, gRPC, WebSocket | HTTP, gRPC, TCP; Envoy also UDP | HTTP, REST, WebSocket; some gRPC | HTTP, gRPC, TCP inside the cluster |
| TLS | Passthrough or termination | Termination (re-encrypt optional) | Termination, passthrough, or re-encrypt | Termination | Automatic mTLS between workloads |
| Client IP to backend | Source IP preservation or proxy protocol | X-Forwarded-For | X-Forwarded-For, proxy protocol | X-Forwarded-For or request context | Peer workload identity; end-user IP via X-Forwarded-For |
| Routing | Port only | Host, path, header, query, method | Arbitrary, scriptable | Path, version, stage, consumer | Service, header, weight, subset |
| Auth | None | Basic OIDC or Cognito integration | Via modules or filters | Core feature (JWT, keys, custom) | Workload identity (mTLS), not users |
| Rate limiting | None | None built in (pair with a WAF) | Local or external service | Core feature, per key or plan | Local or global via external service |
| Config model | Managed API | Managed API | Files (NGINX) or xDS (Envoy) | Managed API or declarative plugins | Control plane resources |
| Observability | Flow logs, connection metrics | Access logs, request metrics | Rich, especially Envoy | Per-API and per-consumer metrics | Uniform golden signals and traces |
| Placement | Edge or in front of proxies | Edge or internal, north-south | Edge, ingress, or sidecar | Public API edge | East-west between services |
| Added cost | Low per hop | Low per hop | You run and patch the fleet | Per-request (managed) or fleet | Proxy per pod or node, control plane |
| Main failure risk | Pinned long-lived connections | Timeout mismatches, deep health checks | Config drift, single point of failure | Logic creep, limiter dependency | Retry amplification, cert and upgrade issues |

The rows that matter most are **unit of decision** and **placement**: an L4 balancer decides once per connection (cheap, protocol-agnostic, blind to gRPC streams), while a mesh decides per internal call (uniform identity and resilience, extra hops everywhere).

## When to Pick Which

These components compose rather than compete. The flowchart asks the question that usually decides the primary component at each boundary; real systems answer "yes" more than once and stack the results.

![A decision flowchart: if the traffic is non-HTTP or needs static IPs or TLS passthrough, use an L4 load balancer; otherwise if it is a public API needing keys and quotas, put an API gateway in front; otherwise if many internal services need mutual TLS, add a service mesh for east-west traffic; otherwise if you need caching or custom routing, run a reverse proxy such as NGINX or Envoy; otherwise use a managed L7 load balancer.](https://d5osvdbc8um23.cloudfront.net/static-asset/blog_images/load-balancers-gateways-proxies/05-decision-flowchart.png)

The last box is deliberately the default: a managed L7 load balancer in front of stateless services fits most HTTP systems. Add everything else because a specific requirement forced it.

### Scenario 1: Mobile Commerce API (the Running Example)

**Requirements**: a public REST API for iOS and Android, JWT auth, per-app quotas, v1 and v2 in parallel during a migration, a dozen services, PCI scope around payments.

**Choice**: DNS latency routing to two regions; a CDN edge for nearby TLS termination and catalog caching; Amazon API Gateway (REST API) for JWT validation, usage plans, request validation, and v1/v2 routing; a VPC link to an internal ALB using least outstanding requests and a 30-second deregistration delay. East-west, start with mTLS and timeouts in a shared client library, and adopt a mesh when service count and language mix make that impractical, using it first to ensure only the order service can call `/charge`.

### Scenario 2: Real-Time Order Tracking Over WebSockets and gRPC

**Requirements**: hundreds of thousands of concurrent WebSocket connections from courier and customer apps, plus internal gRPC streaming between tracking and location services.

**Choice**: an L7 balancer with WebSocket support (ALB, or Envoy behind an NLB), idle timeouts above the client heartbeat interval, and clients that reconnect with jittered backoff, since deploys and scale-in will cut connections. For internal gRPC, use client-side balancing or an Envoy sidecar so HTTP/2 requests are balanced per request; a plain L4 path, including a basic Kubernetes ClusterIP service, balances connections and leaves backends uneven. If session state lives on one tracking node, route by consistent hash on user ID rather than cookie stickiness.

### Scenario 3: Partner TCP Integration

**Requirements**: a logistics partner connects over a custom TCP protocol with mutual TLS and needs static IPs to allowlist.

**Choice**: an NLB with Elastic IPs per zone and a TCP listener (TLS passthrough) so the partner's client certificate reaches the service intact, proxy protocol v2 so the service logs the partner's real address, and health checks on a separate HTTP readiness port. No gateway or L7 balancer, because there is no HTTP to inspect.

## In the Interview

### How to Justify a Choice

For every box between client and service, state five things:

1. **Layer and unit of decision**: "An L7 load balancer, so it routes per request, which matters because the app uses HTTP/2."
2. **What it terminates**: "TLS ends at the edge and is re-encrypted to the ALB; inside the cluster the mesh uses mTLS."
3. **What it owns**: "The gateway owns JWT validation and per-app quotas; ownership checks stay in the order service."
4. **Failure behavior**: "Readiness-only health checks, a 30-second deregistration delay, and only the sidecar retries payment calls, with an idempotency key."
5. **The cost**: "Each proxy is a hop and a system to run; I would not add a mesh until a shared library stops scaling."

### Common Traps

- **"The load balancer handles it"** without naming the layer. Expect a gRPC or WebSocket follow-up.
- **Sticky sessions as the state strategy** instead of stateless services with an external session store.
- **Retries everywhere**: pick one layer, add budgets and idempotency keys, and show the multiplication.
- **Inner timeouts longer than outer ones.**
- **Trusting the leftmost X-Forwarded-For entry.**
- **API keys as authentication.**
- **A single instance of anything** in the path: every tier spans zones, and DNS or anycast handles regions.

### Likely Follow-Up Questions

- **"Why not an NLB for everything?"** No path or header routing, no `X-Forwarded-For`, and long-lived HTTP/2 connections pin to one backend. It suits non-HTTP traffic, static IPs, and fronting your own proxies.
- **"How do you deploy without dropping requests?"** Readiness checks, a deregistration delay sized to the longest normal request, graceful shutdown that finishes in-flight work, and slow start for new instances.
- **"How does the backend know the client IP?"** `X-Forwarded-For` counted from the right through trusted hops at L7; source IP preservation or proxy protocol at L4.
- **"Where do you terminate TLS?"** At the edge for latency and certificate management, re-encrypt inside if compliance requires, passthrough only when the backend must see the client certificate or the balancer must never hold the key.
- **"API gateway or service mesh?"** Different directions: the gateway handles external callers (user identity, quotas, versions), the mesh handles internal calls (workload identity, mTLS, retries). Large systems use both.
- **"What if the mesh control plane goes down?"** Proxies keep their last config and traffic flows; new pods and certificate rotation break, so run the control plane highly available.
- **"A downstream is slow and the site is degrading. What do you check?"** Retry amplification, overlong timeouts, connection pools exhausted by slow calls, and whether outlier detection or circuit breakers are ejecting bad hosts.

Related reading: [Designing a Distributed Rate Limiter](/blog/system-design-distributed-rate-limiter), [Designing a Content Delivery Network](/blog/system-design-cdn), and [REST, GraphQL, gRPC, WebSockets, SSE, and WebRTC Compared](/blog/api-and-realtime-protocols-compared).
