Load Balancers, Reverse Proxies, API Gateways, and Service Meshes: A Complete Guide

    20 min read
    load balancing
    system design
    API gateway
    service mesh
    reverse proxy

    Introduction

    Draw any system design on a whiteboard and there is a box between the client and the service. Candidates usually label it "load balancer" and move on. Interviewers rarely let them. Does that box work at the connection level or the request level? Does it terminate TLS? How does the service learn the caller's real IP? What happens to in-flight requests during a deploy? Who checks the JWT, who enforces the rate limit, and who retries when the payment service times out?

    The answers depend on which component sits in the box. L4 load balancers, L7 load balancers, reverse proxies, API gateways, and service meshes all forward traffic, but they see different things, cost different amounts, and fail in different ways. Production systems often run three or four of them in series.

    This post follows one request, a mobile app placing an order, from the edge to an order service and on to a payment service. For each layer it covers when to use it, how it works, real configuration, scaling, and failure modes. Then it compares them head to head, gives a decision flowchart with three scenarios, and closes with the follow-up questions interviewers ask most.

    The request path for a mobile order: the app resolves the API hostname through latency-based DNS, connects over HTTPS to a nearby edge location, passes through an API gateway that enforces auth, quotas, and validation, reaches an application load balancer that routes the orders path to the order service, which calls the payment service over mutual TLS and writes to the orders database while a mesh control plane pushes configuration and certificates to its sidecar.

    What Sits Between Client and Service

    The Running Example

    The mobile app sends POST https://api.shop.example/v2/orders with a JWT and an API key identifying the app build. The path:

    1. DNS resolves api.shop.example to the closest healthy region.
    2. Edge: TCP and TLS terminate at a nearby edge location (a CDN or anycast front door), keeping handshake round trips short on a mobile network.
    3. API gateway validates the JWT, checks the API key against a usage plan, enforces a rate limit, validates the body, and maps /v2/orders to a backend.
    4. L7 load balancer inside the VPC routes /orders/* to the order service and spreads requests across healthy instances.
    5. Service mesh: the order service's call to the payment service leaves through a local sidecar proxy that adds mutual TLS, a timeout, safe retries, and metrics.

    North-South and East-West

    North-south traffic enters from outside; east-west traffic flows between services. Load balancers and gateways are mostly north-south, meshes east-west, and NGINX and Envoy appear in both.

    Global Load Balancing Before the First Packet

    Something has to pick a region first. DNS-based global load balancing returns different addresses for the same name based on resolver location, latency, weights, or health; Amazon Route 53 offers latency, geolocation, weighted, and failover policies. Its weakness is caching: resolvers hold answers for the TTL and some ignore it, so failover is only as fast as the slowest cache. Anycast advertises one IP from many locations over BGP, and internet routing delivers each packet to the nearest one, so failover happens in routing rather than DNS caches. AWS Global Accelerator and several large CDNs, such as Cloudflare and Fastly, use anycast at the edge.

    L4 Load Balancers: NLB

    When to Reach for L4

    A Layer 4 load balancer works on connections. It sees IPs, ports, and the protocol (TCP or UDP), but does not parse HTTP, so it cannot route on paths, headers, or cookies. In exchange it is protocol-agnostic and does very little work per packet. Reach for it when you need:

    • Non-HTTP protocols: raw TCP (databases, MQTT, custom binary), UDP, or TLS you want passed through untouched.
    • Static IPs: AWS Network Load Balancer (NLB) gives one static IP per Availability Zone and can use Elastic IPs, which matters when partners allowlist your addresses.
    • TLS passthrough, where the balancer never holds the private key.
    • A front door for your own proxies: an NLB in front of an Envoy or NGINX fleet (or a Kubernetes ingress controller) supplies static IPs and spreads connections while the proxies do L7 work.

    How It Works

    When a client opens a TCP connection, NLB picks a target with a flow hash (over protocol, source and destination IP and port, and TCP sequence number) and forwards every packet of that flow to the same target. The decision is made once per connection, so every request on a long-lived connection goes to the same backend.

    That is harmless for short HTTP/1.1 connections and a real problem for gRPC, which multiplexes many requests over one long-lived HTTP/2 connection. Ten clients with one connection each, through an L4 balancer to twenty backends, keep at most ten backends busy. The fixes are an L7 balancer that balances HTTP/2 streams, client-side load balancing (the gRPC client resolves all backends and balances per request), or a maximum connection age on the server to force periodic reconnection.

    TLS Termination vs Passthrough

    A TCP listener passes encrypted bytes straight through; the target holds the certificate and the balancer sees nothing. A TLS listener decrypts with a certificate from AWS Certificate Manager and opens a new connection to the target. Termination centralizes certificates and offloads handshakes; passthrough keeps the key on the backend for true end-to-end encryption or client certificate checks. Re-encryption (terminate, then TLS again to the backend) gives inspection plus encryption on the wire, at the cost of a second handshake.

    Client IP Preservation and Proxy Protocol

    An L4 balancer cannot add X-Forwarded-For, so it uses one of two mechanisms:

    • Client IP preservation: packets reach the target with the client's source IP. On NLB it is on by default for instance targets, and for IP targets the default depends on protocol, so check the preserve_client_ip.enabled attribute rather than assuming. With it on, target security groups must allow client addresses, not just the balancer's.
    • Proxy protocol: a header, defined by HAProxy, prepended to the TCP stream with the original addresses. Version 1 is text, version 2 binary; NLB supports version 2. The backend must expect it or it will reject the connection.

    Concrete Config: NLB in Front of the Ingress Proxies

    IngressTargetGroup:
      Type: AWS::ElasticLoadBalancingV2::TargetGroup
      Properties:
        Protocol: TCP
        Port: 8443
        TargetType: ip
        VpcId: !Ref Vpc
        HealthCheckProtocol: HTTP
        HealthCheckPort: "8081"                # plain HTTP; NLB health checks send no TLS or proxy header
        HealthCheckPath: /ready
        TargetGroupAttributes:
          - Key: proxy_protocol_v2.enabled
            Value: "true"
          - Key: preserve_client_ip.enabled
            Value: "false"
          - Key: deregistration_delay.timeout_seconds
            Value: "60"
          - Key: load_balancing.cross_zone.enabled
            Value: "true"
    

    Scaling and Cross-Zone Load Balancing

    NLB runs a node in each enabled zone. With cross-zone load balancing off, each node sends only to targets in its own zone: if zone A has two targets and zone B eight, each zone A target takes a quarter of all traffic. With it on, load evens out but traffic crosses zones, and AWS bills that inter-zone data on NLB. The defaults differ: cross-zone is off by default on NLB and on by default on ALB. Keep target counts balanced across zones either way.

    NLB's TCP idle timeout was historically fixed at 350 seconds and is now configurable per listener. Clients holding idle connections longer need TCP keepalives below it or they see silent resets.

    Failure Modes and Anti-Patterns

    1. gRPC behind L4 with no client-side balancing: a few backends take all the load.
    2. Proxy protocol enabled on one side only: every connection fails to parse. Change both sides together.
    3. Security groups allowing only balancer IPs while client IP preservation is on: health checks pass, real traffic drops.
    4. Unbalanced zones with cross-zone off.
    5. Expecting L7 features: no path routing, header injection, or per-request retries.

    L7 Load Balancers: ALB, Path Routing, and Sticky Sessions

    When to Reach for L7

    A Layer 7 load balancer terminates HTTP, parses each request, and decides per request. That unlocks routing on host, path, headers, query string, and method; per-request balancing across HTTP/2 connections; header injection; and HTTP-aware health checks. AWS Application Load Balancer (ALB) is the managed example; Google Cloud's Application Load Balancer and Azure Application Gateway fill the same role. For our flow, ALB routes /orders/* to the order service and /track/* to the tracking service.

    An L4 load balancer pins each TCP connection from the mobile app to one ingress proxy for the life of the connection, while an L7 load balancer parses every request and routes the orders path to the order service and the tracking path to the tracking service.

    How It Works

    An ALB terminates client connections and keeps its own reusable pool of connections to targets. Each request is matched against listener rules in priority order; the first match picks a target group, whose algorithm picks a target. ALB offers round robin (default), least outstanding requests, and weighted random (optionally with automatic target weights that shift traffic away from targets returning anomalous errors). Least outstanding requests helps when request cost varies, as with an order service where some calls hit a cache and others run a multi-table transaction.

    ALB adds X-Forwarded-For, X-Forwarded-Proto, and X-Forwarded-Port, supports WebSockets, and supports gRPC end to end with a target group whose protocol version is GRPC.

    X-Forwarded-For Done Right

    X-Forwarded-For is a list to which each proxy appends the address it received the connection from. The leftmost entry is whatever the client claimed, so it can be forged. Find the real client by counting from the right, skipping the hops you control: behind edge, gateway, and ALB, trust only that many rightmost entries. Using the leftmost value for rate limiting or fraud checks is a classic bug. The standardized Forwarded header (RFC 7239) carries the same data in structured form but is less common.

    Health Checks and Connection Draining

    An ALB health check is an HTTP request to a path you choose, with configurable interval, timeout, and thresholds. It should report readiness of the instance, not the health of every dependency. If the check calls the payment service and payment blips, every order instance fails at once and the balancer has nothing to route to. Check what is local; let dependency failures surface as errors callers can handle.

    Connection draining, which AWS calls the deregistration delay, governs a target being removed in a deploy or scale-in: no new requests, but in-flight ones get up to the delay to finish. The default is 300 seconds. For an order API whose requests finish in under a second, that slows every deploy for nothing; 30 seconds is plenty. For WebSockets the delay is a ceiling after which connections are cut, so clients must reconnect gracefully regardless. Slow start is the mirror image: a new target's share ramps up over a window, protecting JVM services that need to warm caches. Note that ALB does not allow slow start together with the least outstanding requests algorithm, so the config below uses the latter.

    Concrete Config: Path Routing and Stickiness

    OrdersRule:
      Type: AWS::ElasticLoadBalancingV2::ListenerRule
      Properties:
        ListenerArn: !Ref HttpsListener
        Priority: 10
        Conditions:
          - Field: path-pattern
            PathPatternConfig:
              Values: ["/orders", "/orders/*"]
        Actions:
          - Type: forward
            TargetGroupArn: !Ref OrdersTargetGroup
    
    OrdersTargetGroup:
      Type: AWS::ElasticLoadBalancingV2::TargetGroup
      Properties:
        Protocol: HTTP
        ProtocolVersion: HTTP2
        Port: 8080
        TargetType: ip
        VpcId: !Ref Vpc
        HealthCheckPath: /ready
        TargetGroupAttributes:
          - Key: load_balancing.algorithm.type
            Value: least_outstanding_requests
          - Key: deregistration_delay.timeout_seconds
            Value: "30"
          - Key: stickiness.enabled
            Value: "false"
    

    Sticky Sessions and Their Trade-Offs

    Sticky sessions pin a client to one target. ALB supports a duration-based cookie it generates (AWSALB) and an application-based cookie your service sets; NLB supports stickiness by source IP. Stickiness helps when a target holds expensive per-client state such as a warmed cache. The costs:

    • Uneven load: heavy clients pin to a few targets that least-outstanding-requests can no longer relieve.
    • Scaling lag: new targets only receive new clients.
    • Lost state when a target dies.
    • NAT pile-ups with source IP stickiness, since many mobile users share a carrier NAT address.

    The better default is stateless instances with session state in Redis or DynamoDB. Stickiness is an optimization, never the source of correctness.

    Scaling and Failure Modes

    ALB scales itself, but not instantly; a step change such as a flash sale launch can outrun it, and ALB now offers capacity reservation for planned events. Its IP addresses change as it scales, so point DNS at it with an alias record and never hard-code IPs.

    1. Idle timeout mismatches: ALB's idle timeout defaults to 60 seconds. If the backend's keepalive timeout is shorter, the backend can close a connection just as the ALB reuses it, causing intermittent 502s. Make the backend's timeout longer.
    2. Deep health checks that take the whole fleet out during a dependency outage.
    3. Trusting the leftmost X-Forwarded-For entry.
    4. Rule sprawl on one listener, which also runs into documented rule quotas.
    5. Stickiness as architecture, which turns every deploy into a small outage for pinned users.

    Reverse Proxies: NGINX and Envoy

    When to Reach for a Self-Hosted Proxy

    Managed load balancers are reverse proxies too, but in an interview "reverse proxy" usually means NGINX, HAProxy, or Envoy running on your own instances or as Kubernetes ingress. Choose one when you need what the managed balancer lacks: response caching, buffering, custom routing, cluster-edge rate limiting, identical config across clouds, or deep per-route telemetry.

    NGINX

    NGINX runs a few worker processes, each an event loop over non-blocking sockets, so one worker holds many thousands of connections. Configuration is a static file; nginx -s reload starts new workers on the new config while old ones finish their connections. Upstream algorithms include round robin (default), least_conn, ip_hash, hash with an optional consistent flag, and random two least_conn. Open source NGINX does passive health checks (max_fails, fail_timeout); active checks are a commercial NGINX Plus feature.

    For our order API, an NGINX tier behind the NLB absorbs slow mobile clients and caches the catalog:

    upstream order_service {
        least_conn;
        server 10.0.1.10:8080 max_fails=3 fail_timeout=10s;
        server 10.0.2.10:8080 max_fails=3 fail_timeout=10s;
        keepalive 64;
    }
    
    server {
        listen 8081;                            # health check port for the NLB
        location = /ready { return 200; }
    }
    
    proxy_cache_path /var/cache/nginx keys_zone=catalog:50m max_size=2g inactive=10m;
    limit_req_zone $binary_remote_addr zone=per_ip:10m rate=20r/s;
    
    server {
        listen 8443 ssl proxy_protocol;         # NLB in front sends proxy protocol v2
        ssl_certificate     /etc/nginx/tls/api.crt;
        ssl_certificate_key /etc/nginx/tls/api.key;
        set_real_ip_from 10.0.0.0/16;           # trust only our NLB subnet
        real_ip_header proxy_protocol;
    
        location /orders {
            limit_req zone=per_ip burst=40 nodelay;
            proxy_pass http://order_service;
            proxy_http_version 1.1;
            proxy_set_header Connection "";
            proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
            proxy_set_header X-Forwarded-Proto $scheme;
            proxy_connect_timeout 1s;
            proxy_read_timeout 5s;
        }
    
        location /catalog/ {
            proxy_cache catalog;
            proxy_cache_valid 200 60s;
            proxy_cache_use_stale error timeout updating;
            proxy_pass http://order_service;
        }
    }
    

    Caching and Buffering at the Proxy

    • Buffering: with proxy_buffering on (the NGINX default), the proxy reads the upstream response quickly and trickles it to a slow client, freeing the order service's worker instead of tying it to a 3G connection. Streaming responses (server-sent events, long polling) need buffering off on those routes or clients see nothing until the buffer fills.
    • Caching: a short-lived cache for anonymous, idempotent responses collapses many identical requests into one upstream fetch, and proxy_cache_use_stale turns a backend outage into stale data instead of errors. Never cache responses that depend on Authorization unless the user is in the cache key.

    Envoy

    Envoy, built at Lyft and now a CNCF graduated project, is an L4 and L7 proxy designed for dynamic environments. Three things set it apart:

    1. Dynamic configuration through xDS: listeners (LDS), routes (RDS), clusters (CDS), endpoints (EDS), and secrets (SDS) stream from a management server over gRPC, so routes and endpoints change with no reload. This is how a mesh control plane reprograms thousands of proxies as pods come and go.
    2. Observability by default: per-cluster and per-route stats (latency histograms, retries, circuit breaker trips, ejections), tracing header propagation, and structured access logs. Much of what people credit to "the mesh" is Envoy's telemetry.
    3. Resilience primitives: per-route timeouts and retries, retry budgets, circuit breakers on connections and pending or concurrent requests, and outlier detection that ejects hosts returning consecutive 5xx.
    clusters:
      - name: order_service
        type: EDS
        eds_cluster_config:
          eds_config: { ads: {} }          # endpoints pushed by the control plane
        lb_policy: LEAST_REQUEST            # power of two choices by default
        typed_extension_protocol_options:
          envoy.extensions.upstreams.http.v3.HttpProtocolOptions:
            "@type": type.googleapis.com/envoy.extensions.upstreams.http.v3.HttpProtocolOptions
            explicit_http_config: { http2_protocol_options: {} }
        circuit_breakers:
          thresholds:
            - max_connections: 1000
              max_pending_requests: 200
              max_requests: 1000
              max_retries: 50
        outlier_detection:
          consecutive_5xx: 5
          interval: 10s
          base_ejection_time: 30s
          max_ejection_percent: 50
    
    routes:
      - match: { path_separated_prefix: "/orders" }
        route:
          cluster: order_service
          timeout: 3s
          retry_policy:
            retry_on: "connect-failure,refused-stream,unavailable"
            num_retries: 1
            per_try_timeout: 1s
    

    Connect failures and refused streams mean the request never reached application code; gRPC unavailable usually means the same but is not guaranteed, so non-idempotent calls still need an idempotency key. Retrying POST /orders on any 5xx risks a duplicate order unless the API is idempotent.

    NGINX vs Envoy

    NGINX is simpler, battle-tested as a web server, strong at static content, caching, and buffering, and configured by files humans edit. Envoy is built to be configured by machines, changes live through xDS, speaks HTTP/2 and gRPC natively on both sides, and emits richer telemetry. One team hand-managing an edge tier often picks NGINX; a control plane managing hundreds of proxies usually means Envoy.

    Failure Modes and Anti-Patterns

    1. A single proxy instance: run at least two per zone behind an L4 balancer.
    2. Retries at every layer (see the retry storm section).
    3. Caching personalized responses with a key that omits the user.
    4. Buffering streaming endpoints, so server-sent events appear to hang.
    5. Hand-edited config drift across a fleet: generate config from one source or use xDS.

    API Gateways: Auth, Rate Limits, and Transforms

    When to Reach for a Gateway

    An API gateway is an L7 proxy specialized for API management. A load balancer asks which healthy instance gets the request. A gateway asks who the caller is, whether they are allowed, whether they are over quota, whether the request is well-formed, and which backend version should handle it. Add one when you expose APIs to mobile apps, partners, or third-party developers and want those concerns handled once instead of in every service.

    Inside an API gateway, a request from the mobile app carrying a JWT and an API key passes through authentication and authorization, then usage plan and rate limit checks, then request validation, then transformation and version routing, before a VPC link forwards it to the internal load balancer, which sends the v1 path to order service v1 and the v2 path to order service v2.

    Responsibilities

    • Authentication and authorization: validate JWTs (signature, issuer, audience, expiry) or call a custom authorizer. Amazon API Gateway supports IAM, Cognito user pools, Lambda authorizers, and native JWT authorizers on HTTP APIs. Coarse checks ("token has orders:write") belong here; fine-grained ones ("this user owns this order") belong in the service that has the data.
    • Rate limiting: token-bucket limits per key, route, or client, protecting backends and enforcing commercial tiers.
    • API keys and usage plans: identify the calling application and attach a quota. API keys identify, they do not authenticate; AWS's documentation says not to rely on them alone for authorization.
    • Request validation against a schema before the request costs backend compute.
    • Transformation: rename fields, add or strip headers, or convert protocols.
    • Versioning and routing: /v1/orders to the old service, /v2/orders to the new, or canary a share of traffic to a new stage.

    Managed: Amazon API Gateway

    Amazon API Gateway offers REST APIs (usage plans, API keys, request validation, mapping templates, caching, private endpoints), HTTP APIs (cheaper and simpler, native JWT authorizers, fewer features), and WebSocket APIs. Documented limits that shape designs: 10 MB payloads and a default 29-second integration timeout. For REST APIs, AWS now allows raising that timeout for Regional and private APIs through a quota increase, possibly at the cost of a lower throttle limit. Accounts also have a default Region-level throttle quota shared across APIs; check yours rather than assume.

    A free-tier usage plan for the mobile app:

    aws apigateway create-usage-plan \
      --name mobile-free \
      --throttle burstLimit=20,rateLimit=10 \
      --quota limit=10000,period=DAY \
      --api-stages apiId=a1b2c3d4e5,stage=prod
    
    aws apigateway create-usage-plan-key \
      --usage-plan-id <plan-id> --key-id <api-key-id> --key-type API_KEY
    

    Managed buys zero servers and built-in scaling; you give up portability, deep customization, and cost control at very high volume.

    Self-Hosted: Kong and Envoy-Based Gateways

    Kong is built on NGINX and OpenResty and adds plugins for auth, rate limiting, transformations, and logging:

    _format_version: "3.0"
    services:
      - name: order-service-v2
        url: http://orders-v2.internal:8080
        routes:
          - name: orders-v2
            paths: ["/v2/orders"]
        plugins:
          - name: jwt
          - name: rate-limiting
            config:
              minute: 60
              policy: redis
              redis:
                host: ratelimit.internal
          - name: request-transformer
            config:
              add:
                headers: ["X-Api-Version:2"]
    

    Envoy-based gateways (Envoy Gateway, which implements the Kubernetes Gateway API, plus products such as Gloo and Emissary) put a gateway control plane over Envoy.

    Rate Limiting Is a Distributed Problem

    One gateway node can count in memory. Twenty nodes either each enforce a twentieth of the limit (inaccurate when traffic is uneven) or share counters in Redis (a network call per request and a new dependency). Decide whether a limit must be exact, like a paid quota, or approximate, like abuse protection, and say so.

    Failure Modes and Anti-Patterns

    1. Business logic in the gateway: transformations grow into order rules, and gateway deploys become application deploys.
    2. API keys as authentication: keys are trivially extracted from mobile binaries.
    3. Gateway as the only defense: internal callers can bypass it, so services still validate input and ownership.
    4. Ignoring the integration timeout: long work must be asynchronous (return 202 with an order ID, then poll or push).
    5. Undecided failure mode for the limiter store: choose fail open or fail closed in advance.

    Service Mesh: Sidecars and mTLS

    When to Reach for a Mesh

    With dozens of services, every team re-implements service-to-service TLS, retries, timeouts, circuit breakers, metrics, and tracing. A mesh moves those concerns into the network layer, uniformly, and lets a platform team change policy without redeploying services. Order calling payment is the canonical case: encrypted and authenticated both ways, a short timeout, retries only when safe, and visible on a dashboard.

    Data Plane and Control Plane

    The data plane is the proxies carrying traffic: Envoy in Istio, a lightweight Rust proxy in Linkerd. The control plane pushes desired configuration to every proxy. Istio's istiod turns Kubernetes resources (VirtualService, DestinationRule, PeerAuthentication) into xDS and acts as a certificate authority for short-lived workload certificates.

    In a sidecar mesh, traffic from the mesh ingress gateway reaches the Envoy sidecar in the order pod over mutual TLS, the sidecar forwards to the order container on localhost, outbound calls leave through the same sidecar with mutual TLS and retries to the payment pod's sidecar, the istiod control plane pushes xDS configuration and certificates to both sidecars, and the sidecars emit metrics and traces to a telemetry backend.

    Sidecar vs Sidecarless (Ambient) Modes

    In the sidecar model each pod gets a proxy container, and iptables rules or a CNI plugin redirect its traffic through it. The app speaks plain HTTP to localhost. Costs: a proxy per pod, two extra proxy hops per call, mesh upgrades that require pod restarts, and startup races if the app starts before its sidecar.

    Sidecarless designs reduce that. Istio's ambient mode, generally available since Istio 1.24, uses a per-node ztunnel for L4 work (mTLS, identity, L4 authorization) and optional waypoint proxies (Envoy) for L7 features only where needed. Cilium offers an eBPF-based mesh with per-node proxies. The trade is lower resource use and easier upgrades against a shared per-node component and a shorter operational track record.

    mTLS and Workload Identity

    With mutual TLS both sides present certificates. The mesh issues each workload a certificate encoding its identity (in Istio, a SPIFFE identity from namespace and service account), rotates it automatically, and lets you write policy against identities instead of pod IPs that change constantly:

    apiVersion: security.istio.io/v1
    kind: PeerAuthentication
    metadata:
      name: default
      namespace: payments
    spec:
      mtls:
        mode: STRICT
    ---
    apiVersion: security.istio.io/v1
    kind: AuthorizationPolicy
    metadata:
      name: allow-order-to-charge
      namespace: payments
    spec:
      selector:
        matchLabels: { app: payment }
      rules:
        - from:
            - source:
                principals: ["cluster.local/ns/orders/sa/order-service"]
          to:
            - operation:
                methods: ["POST"]
                paths: ["/charge"]
    

    Roll out in PERMISSIVE mode (accept plaintext and mTLS), confirm every caller is meshed, then switch to STRICT.

    Retries, Timeouts, and Circuit Breaking

    apiVersion: networking.istio.io/v1
    kind: VirtualService
    metadata:
      name: payment
      namespace: payments
    spec:
      hosts: ["payment.payments.svc.cluster.local"]
      http:
        - route:
            - destination: { host: payment.payments.svc.cluster.local }
          timeout: 2s
          retries:
            attempts: 1
            perTryTimeout: 900ms
            retryOn: connect-failure,refused-stream,unavailable
    ---
    apiVersion: networking.istio.io/v1
    kind: DestinationRule
    metadata:
      name: payment
      namespace: payments
    spec:
      host: payment.payments.svc.cluster.local
      trafficPolicy:
        connectionPool:
          http: { http2MaxRequests: 500, maxRequestsPerConnection: 100 }
        outlierDetection:
          consecutive5xxErrors: 5
          interval: 10s
          baseEjectionTime: 30s
          maxEjectionPercent: 50
    

    Charging is not naturally idempotent, so retries cover only failures where the request never reached the payment container, and the order service sends an idempotency key so even a duplicate charges once. Outlier detection acts as a host-level circuit breaker, ejecting a payment pod that returns repeated 5xx.

    The Cost of the Extra Hop

    Each proxy adds latency, CPU, and memory. Per hop the latency is usually small next to network and application time, but a sidecar mesh adds two proxies per call, so a chain of five service-to-service calls passes ten. Measure it. The larger costs are operational: another control plane to upgrade, certificate expiry as a failure mode, and debugging across two sets of logs.

    Failure Modes and Anti-Patterns

    1. A mesh for three services: a shared client library is cheaper. Meshes pay off at scale and across languages.
    2. Misunderstanding control plane outages: proxies keep their last config, so traffic flows, but new pods and certificate rotation break.
    3. Mesh retries stacked on application retries.
    4. Strict mTLS before every caller is meshed: batch jobs and legacy VMs suddenly fail.
    5. Treating mTLS as user authorization: it proves which workload called, not which user.

    Load-Balancing Algorithms: Round-Robin, Least Connections, Consistent Hashing

    Round-Robin and Weighted Round-Robin

    Round robin sends each request or connection to the next backend in turn. It is simple, needs no shared state, and is fair when requests and backends are uniform. Weighted round robin gives bigger instances more turns and is how canaries work: weight 95 to stable, 5 to canary. Its weakness is ignoring current load, so a slow backend keeps its full share and builds a queue.

    Least Connections and Least Outstanding Requests

    Least connections picks the backend with the fewest open connections; least outstanding requests (Envoy calls it least request) counts in-flight requests. Both adapt to slow backends, which accumulate work and receive less. They need per-backend state, which is exact on one balancer node and approximate across many, since each node sees only its own traffic.

    Power of Two Choices

    Scanning every backend per request is expensive, and when many balancer nodes act on the same stale load data they cause herding, where many balancer nodes see the same least-loaded backend and all pile onto it. Power of two choices samples two backends at random and picks the less loaded. A well-known result in randomized load balancing is that this sharply reduces maximum load compared with one random choice, and different nodes sample different pairs, so herding disappears. Envoy's least request policy uses it by default; NGINX offers random two least_conn.

    import random
    
    def pick_backend(backends, in_flight):
        a, b = random.sample(backends, 2)
        return a if in_flight[a] <= in_flight[b] else b
    

    Consistent Hashing and Maglev

    Sometimes the same key must reach the same backend: a cache shard for a product ID, a user's WebSocket session, or in-memory rate limit counters. hash(key) % N reshuffles almost every key when N changes. Consistent hashing places backends (with many virtual points each) and keys on a ring; a key maps to the next backend clockwise, so adding or removing one moves only about 1/N of keys.

    Maglev, from Google's 2016 paper on its network load balancer, builds a fixed-size lookup table from per-backend permutations, giving single-index lookups, very even load, and small disruption on backend changes. Envoy supports both RING_HASH and MAGLEV. Both give affinity without cookies but risk hot spots on very popular keys.

    import bisect, hashlib
    
    class Ring:
        def __init__(self, nodes, vnodes=100):
            self.points = sorted(
                (int(hashlib.md5(f"{n}#{i}".encode()).hexdigest(), 16), n)
                for n in nodes for i in range(vnodes)
            )
            self.keys = [p for p, _ in self.points]
    
        def lookup(self, key):
            h = int(hashlib.md5(key.encode()).hexdigest(), 16)
            i = bisect.bisect(self.keys, h) % len(self.points)
            return self.points[i][1]
    

    Retry Storms and Timeout Budgets

    Algorithms decide where traffic goes; retries decide how much there is. If the mobile client, gateway, order service, and sidecar each make up to three attempts, one failure at payment becomes 3 x 3 x 3 x 3 = 81 attempts at the bottom, arriving exactly when payment can least cope. That retry storm turns a partial outage into a total one. Defenses:

    • Retry at one layer, ideally the one nearest the failure that knows whether the operation is idempotent.
    • Retry budgets that cap retries as a share of normal traffic; Envoy and gRPC both support this style of limit.
    • Exponential backoff with jitter so clients do not retry in lockstep.
    • Idempotency keys on POST /orders and /charge.
    • Deadline propagation: each timeout shorter than its caller's. If the app waits 10 seconds, the gateway allows 8, the order service gives payment 2 seconds with one 900 ms retry, and payment's database call gets a few hundred milliseconds. gRPC carries the deadline on the wire, and passing the incoming context to outbound calls propagates it; for HTTP, pass the remaining budget in a header. An inner timeout longer than the outer one wastes work the caller already abandoned.

    Head-to-Head Comparison

    DimensionL4 LB (NLB)L7 LB (ALB)Reverse proxy (NGINX, Envoy)API gatewayService mesh
    Unit of decisionConnection (flow)RequestRequest (or connection in L4 mode)Request plus caller identityRequest, per service-to-service call
    ProtocolsTCP, UDP, TLSHTTP/1.1, HTTP/2, gRPC, WebSocketHTTP, gRPC, TCP; Envoy also UDPHTTP, REST, WebSocket; some gRPCHTTP, gRPC, TCP inside the cluster
    TLSPassthrough or terminationTermination (re-encrypt optional)Termination, passthrough, or re-encryptTerminationAutomatic mTLS between workloads
    Client IP to backendSource IP preservation or proxy protocolX-Forwarded-ForX-Forwarded-For, proxy protocolX-Forwarded-For or request contextPeer workload identity; end-user IP via X-Forwarded-For
    RoutingPort onlyHost, path, header, query, methodArbitrary, scriptablePath, version, stage, consumerService, header, weight, subset
    AuthNoneBasic OIDC or Cognito integrationVia modules or filtersCore feature (JWT, keys, custom)Workload identity (mTLS), not users
    Rate limitingNoneNone built in (pair with a WAF)Local or external serviceCore feature, per key or planLocal or global via external service
    Config modelManaged APIManaged APIFiles (NGINX) or xDS (Envoy)Managed API or declarative pluginsControl plane resources
    ObservabilityFlow logs, connection metricsAccess logs, request metricsRich, especially EnvoyPer-API and per-consumer metricsUniform golden signals and traces
    PlacementEdge or in front of proxiesEdge or internal, north-southEdge, ingress, or sidecarPublic API edgeEast-west between services
    Added costLow per hopLow per hopYou run and patch the fleetPer-request (managed) or fleetProxy per pod or node, control plane
    Main failure riskPinned long-lived connectionsTimeout mismatches, deep health checksConfig drift, single point of failureLogic creep, limiter dependencyRetry amplification, cert and upgrade issues

    The rows that matter most are unit of decision and placement: an L4 balancer decides once per connection (cheap, protocol-agnostic, blind to gRPC streams), while a mesh decides per internal call (uniform identity and resilience, extra hops everywhere).

    When to Pick Which

    These components compose rather than compete. The flowchart asks the question that usually decides the primary component at each boundary; real systems answer "yes" more than once and stack the results.

    A decision flowchart: if the traffic is non-HTTP or needs static IPs or TLS passthrough, use an L4 load balancer; otherwise if it is a public API needing keys and quotas, put an API gateway in front; otherwise if many internal services need mutual TLS, add a service mesh for east-west traffic; otherwise if you need caching or custom routing, run a reverse proxy such as NGINX or Envoy; otherwise use a managed L7 load balancer.

    The last box is deliberately the default: a managed L7 load balancer in front of stateless services fits most HTTP systems. Add everything else because a specific requirement forced it.

    Scenario 1: Mobile Commerce API (the Running Example)

    Requirements: a public REST API for iOS and Android, JWT auth, per-app quotas, v1 and v2 in parallel during a migration, a dozen services, PCI scope around payments.

    Choice: DNS latency routing to two regions; a CDN edge for nearby TLS termination and catalog caching; Amazon API Gateway (REST API) for JWT validation, usage plans, request validation, and v1/v2 routing; a VPC link to an internal ALB using least outstanding requests and a 30-second deregistration delay. East-west, start with mTLS and timeouts in a shared client library, and adopt a mesh when service count and language mix make that impractical, using it first to ensure only the order service can call /charge.

    Scenario 2: Real-Time Order Tracking Over WebSockets and gRPC

    Requirements: hundreds of thousands of concurrent WebSocket connections from courier and customer apps, plus internal gRPC streaming between tracking and location services.

    Choice: an L7 balancer with WebSocket support (ALB, or Envoy behind an NLB), idle timeouts above the client heartbeat interval, and clients that reconnect with jittered backoff, since deploys and scale-in will cut connections. For internal gRPC, use client-side balancing or an Envoy sidecar so HTTP/2 requests are balanced per request; a plain L4 path, including a basic Kubernetes ClusterIP service, balances connections and leaves backends uneven. If session state lives on one tracking node, route by consistent hash on user ID rather than cookie stickiness.

    Scenario 3: Partner TCP Integration

    Requirements: a logistics partner connects over a custom TCP protocol with mutual TLS and needs static IPs to allowlist.

    Choice: an NLB with Elastic IPs per zone and a TCP listener (TLS passthrough) so the partner's client certificate reaches the service intact, proxy protocol v2 so the service logs the partner's real address, and health checks on a separate HTTP readiness port. No gateway or L7 balancer, because there is no HTTP to inspect.

    In the Interview

    How to Justify a Choice

    For every box between client and service, state five things:

    1. Layer and unit of decision: "An L7 load balancer, so it routes per request, which matters because the app uses HTTP/2."
    2. What it terminates: "TLS ends at the edge and is re-encrypted to the ALB; inside the cluster the mesh uses mTLS."
    3. What it owns: "The gateway owns JWT validation and per-app quotas; ownership checks stay in the order service."
    4. Failure behavior: "Readiness-only health checks, a 30-second deregistration delay, and only the sidecar retries payment calls, with an idempotency key."
    5. The cost: "Each proxy is a hop and a system to run; I would not add a mesh until a shared library stops scaling."

    Common Traps

    • "The load balancer handles it" without naming the layer. Expect a gRPC or WebSocket follow-up.
    • Sticky sessions as the state strategy instead of stateless services with an external session store.
    • Retries everywhere: pick one layer, add budgets and idempotency keys, and show the multiplication.
    • Inner timeouts longer than outer ones.
    • Trusting the leftmost X-Forwarded-For entry.
    • API keys as authentication.
    • A single instance of anything in the path: every tier spans zones, and DNS or anycast handles regions.

    Likely Follow-Up Questions

    • "Why not an NLB for everything?" No path or header routing, no X-Forwarded-For, and long-lived HTTP/2 connections pin to one backend. It suits non-HTTP traffic, static IPs, and fronting your own proxies.
    • "How do you deploy without dropping requests?" Readiness checks, a deregistration delay sized to the longest normal request, graceful shutdown that finishes in-flight work, and slow start for new instances.
    • "How does the backend know the client IP?" X-Forwarded-For counted from the right through trusted hops at L7; source IP preservation or proxy protocol at L4.
    • "Where do you terminate TLS?" At the edge for latency and certificate management, re-encrypt inside if compliance requires, passthrough only when the backend must see the client certificate or the balancer must never hold the key.
    • "API gateway or service mesh?" Different directions: the gateway handles external callers (user identity, quotas, versions), the mesh handles internal calls (workload identity, mTLS, retries). Large systems use both.
    • "What if the mesh control plane goes down?" Proxies keep their last config and traffic flows; new pods and certificate rotation break, so run the control plane highly available.
    • "A downstream is slow and the site is degrading. What do you check?" Retry amplification, overlong timeouts, connection pools exhausted by slow calls, and whether outlier detection or circuit breakers are ejecting bad hosts.

    Related reading: Designing a Distributed Rate Limiter, Designing a Content Delivery Network, and REST, GraphQL, gRPC, WebSockets, SSE, and WebRTC Compared.

    Structured data for LLMs, AI agents, and automated crawlers is available at/blog/load-balancers-gateways-proxies.md. Please reviewrobots.txt andllms.txt before crawling. All referenced data must be credited to roundz.ai with a link tohttps://roundz.ai