API Gateway

API Gateway

Definition: A server that acts as a single entry point into a system, sitting between clients and a collection of backend microservices, routing requests and centralizing cross-cutting concerns that would otherwise be duplicated in every service. The pattern grew directly out of the shift to Microservices Architecture in the early-to-mid 2010s — once a system was dozens of independently deployable services instead of one, clients needed a stable façade, and the “Backend for Frontend” (BFF) variant, commonly credited to engineers at SoundCloud around 2015, refined the idea further by giving each client type (web, mobile) its own tailored gateway.

How It Works

  • Accepts all inbound API calls from clients, routes them to the appropriate internal microservices based on path/host/header rules, and returns the aggregated result to the caller
  • Handles cross-cutting concerns globally: SSL/TLS termination, authentication (JWT validation, API key checks), rate limiting and throttling per client, and CORS headers, so individual services don’t each reimplement them
  • Can transform requests and responses — translating public REST calls into internal gRPC calls, or aggregating several backend calls into a single response for the client, the “Backend for Frontend” variant of this pattern tailors that aggregation per client type
  • Managed offerings (AWS API Gateway, Azure API Management, Google Apigee) bill per request and integrate directly with serverless compute, routing straight to a Lambda function without a backend server in between; self-hosted gateways (Kong, Envoy Gateway, Traefik) run as their own deployable service, often as an Ingress controller in front of a Kubernetes (K8s) cluster
  • Supports canary and blue-green routing by directing a configurable percentage of traffic to a new backend version before a full cutover, catching regressions on a small slice of real traffic first
  • Enforces per-client rate limits and quotas at the edge, commonly using a token-bucket or leaky-bucket algorithm, so a misbehaving or over-eager client is throttled before it ever reaches a backend service

Why It Matters

  • Prevents clients from needing to know the exact IP addresses, ports, and network topology of dozens of internal microservices — the gateway is the only address that has to stay stable
  • Centralizes security and traffic control at one auditable choke point instead of duplicating auth checks, rate limits, and TLS config across every individual service
  • Gives operators a single place to apply global rate limits, WAF rules, and observability instrumentation without touching application code in any downstream service
  • Gives API consumers a stable, versioned public contract even while the backend services behind it are refactored, split, merged, or rewritten entirely

Under the Hood: Rate Limiting at the Gateway with the Token Bucket Algorithm

Because every request in the system passes through the gateway, it’s the natural place to enforce rate limits, and the token bucket is the algorithm most gateways implement for it. Each client (identified by API key, IP, or user ID) gets a virtual bucket that holds up to N tokens and refills at a fixed rate — say, 100 tokens, refilling at 10 tokens per second. Every incoming request consumes one token; if the bucket is empty, the request is rejected with a 429 Too Many Requests (often with a Retry-After header) instead of being forwarded to a backend service. The key property that makes this better than a naive fixed-window counter is that it tolerates bursts gracefully — a client that’s been idle can spend its full bucket in one burst — while still enforcing a hard average rate over time, since the bucket only refills at the configured rate rather than resetting instantly at a window boundary. Gateways typically implement the bucket state in a fast shared store (Redis, or an in-memory store synced across gateway replicas) so the limit is enforced consistently even when a client’s requests are load-balanced across multiple gateway instances, see Rate Limiting for the algorithm in more depth.

Comparison: API Gateway vs Reverse Proxy vs Service Mesh

API GatewayReverse ProxyService Mesh
Primary scopeNorth-south (client to backend)North-south, general purposeEast-west (service to service)
Auth/rate limitingBuilt-in, per-clientPossible, usually less feature-richPossible via policy, less common
Request transformationYes (aggregation, protocol translation)Limited (headers, rewrites)No — mostly transparent passthrough
Awareness of API semanticsHigh (routes, versions, contracts)Low (mostly path/host based)None — operates below the API layer
Typical deploymentEdge of the systemEdge of the system, or in front of a single appInside the cluster, between every service

See Service Mesh for the complementary east-west layer — many production systems run both: a gateway at the edge for client traffic, and a mesh internally for service-to-service traffic.

Common Pitfalls

  • Putting heavy business logic or data transformation directly into the API Gateway layer, turning it into a tightly coupled monolithic bottleneck (the “Enterprise Service Bus” anti-pattern) that every team must coordinate through to ship
  • Making the gateway a single point of failure by under-provisioning it or skipping horizontal scaling and Auto-Scaling, since every request in the system now passes through it
  • Setting global timeout and retry policies that don’t account for slow downstream services, causing cascading request pile-ups instead of fast, isolated failures
  • Letting API versioning sprawl unmanaged at the gateway, with no deprecation policy, until dropping an old route silently breaks clients nobody remembers exist
  • Misconfiguring CORS or auth rules broadly across all routes instead of per-route, accidentally exposing internal-only endpoints to public clients
  • Treating the gateway as a substitute for service-level security instead of defense in depth — a compromised internal network shouldn’t be able to skip the gateway’s checks entirely if a backend service trusts any caller unconditionally

Code Example

# Kong declarative config: route + JWT auth + rate limiting for one service
services:
  - name: billing-service
    url: http://billing.internal:4000
    routes:
      - name: billing-route
        paths: ["/api/billing"]
    plugins:
      - name: jwt
      - name: rate-limiting
        config:
          minute: 1000
          policy: redis
          redis_host: redis.internal
      - name: cors
        config:
          origins: ["https://app.example.com"]

Code Example: Backend-for-Frontend Response Aggregation

// A mobile-specific gateway route combining three backend calls into one
// response, so the client makes a single round trip instead of three
app.get('/mobile/api/dashboard', async (req, res) => {
  const userId = req.user.id;

  const [profile, orders, notifications] = await Promise.all([
    userService.get(`/users/${userId}`),
    orderService.get(`/users/${userId}/orders?limit=5`),
    notificationService.get(`/users/${userId}/unread`),
  ]);

  res.json({
    profile: { name: profile.name, avatarUrl: profile.avatarUrl },
    recentOrders: orders.items,
    unreadCount: notifications.count,
  });
});

Best Practices

  • Keep the gateway thin: routing, auth, rate limiting, and TLS termination belong there, business logic and data aggregation beyond simple composition belong in services
  • Version APIs explicitly (via path or header) and publish a deprecation timeline before removing an old version, rather than breaking clients without warning
  • Apply per-client rate limits and quotas at the gateway so one noisy client can’t degrade service for everyone else behind it, see Rate Limiting
  • Monitor the latency the gateway itself adds — an extra network hop and TLS handshake should be near-negligible, but a misconfigured gateway can quietly become the slowest part of every request
  • Use canary or weighted routing at the gateway to validate a new backend version against a small slice of real traffic before a full cutover

FAQ

How is an API Gateway different from a load balancer? A Load Balancer distributes traffic across identical backend instances based on network-level rules; an API Gateway is API-aware — it routes by path/version, terminates auth, transforms payloads, and enforces per-client policy, and commonly sits in front of one or more load balancers rather than replacing them.

Do I need an API Gateway for a small app with two services? Usually not yet — a simple Reverse Proxy or even direct client calls to each service can suffice until the number of services, clients, or cross-cutting concerns (auth, rate limiting) grows enough that duplicating that logic per service becomes real maintenance pain.

Can a system have more than one API Gateway? Yes — the Backend for Frontend variant intentionally runs a separate gateway per client type (web, mobile, partner API), each tailored to that client’s exact needs, rather than forcing one gateway to serve every consumer identically.

History

  • Emerged from Service-Oriented Architecture-era Enterprise Service Buses in the 2000s, though ESBs typically carried far more logic than a modern gateway is meant to hold
  • Apigee, founded in 2004 and later acquired by Google in 2016, was one of the earliest dedicated API management platforms, popularizing the gateway as a standalone product category
  • AWS API Gateway launched in 2015 alongside the rise of serverless computing, pairing naturally with Lambda since a gateway was needed to expose functions as HTTP endpoints
  • Kong, built on top of the Nginx-based OpenResty and open-sourced in 2015, became one of the most widely adopted self-hosted gateways, followed by Kubernetes-native options like Ambassador and Envoy Gateway as the ecosystem matured

Example

AWS API Gateway or Kong sits in front of backend APIs, verifying a client’s JWT and enforcing a 1000 req/min per-client rate limit before allowing traffic to reach the internal Node.js or Python microservices behind it — so a misbehaving client gets throttled at the edge with a 429 instead of overwhelming the Billing service directly, and the Billing service never has to implement rate limiting or JWT validation itself.

Dig deeper