AlgoMaster Logo

Load Balancers

High Priority9 min readUpdated September 19, 2026
AI Mock Interview

Practice this topic in a realistic system design interview

Listen to this chapter
Unlock Audio

Load Balancers

Imagine your application is running on a single server. At first, everything works fine.

But as more users start using the application, that server has to handle more and more requests. Eventually, you hit a limit.

A load balancer is how you get past that limit. This chapter covers how load balancers work, Layer 4 versus Layer 7 load balancing, and how load balancers are used in real-world systems.

1. The Problem with a Single Server

Once a single server hits its limit, the server becomes overloaded, requests slow down, and if that server fails, the entire application goes down.

The obvious solution is to add more servers. But now you have a new problem:

  • When a request arrives, which server should handle it?
  • How do you distribute traffic evenly?
  • What happens if one of those servers becomes unhealthy or goes offline?

This is exactly what a load balancer solves.

2. How a Load Balancer Works

A load balancer is a component that sits between clients and a group of backend servers.

Instead of clients sending requests directly to a specific server, they send requests to the load balancer. The load balancer then decides which backend server should handle each request and forwards the request to that server.

To the client, the system still looks like a single service. But behind the load balancer, you can have many servers working together.

This allows you to scale horizontally by adding more servers as traffic grows. And if one server becomes unavailable, the load balancer can stop sending traffic to it and route requests to the remaining healthy servers.

3. Why Use Load Balancers

There are a few main reasons we use load balancers in distributed systems.

Scalability

A single server can only handle a limited amount of traffic. Once you reach that limit, instead of making the server bigger and bigger, you can add more servers and let the load balancer distribute requests across them.

This is called horizontal scaling.

High Availability

If your entire application depends on one server, that server becomes a single point of failure. If it crashes, your application goes down.

With multiple servers behind a load balancer, traffic can be redirected to healthy servers when one of them fails.

Better Resource Utilization

Without a load balancer, one server might be overloaded while another is barely doing any work. A load balancer helps spread traffic more evenly so that your available capacity is used efficiently.

Adding and Removing Servers

Load balancers make it easier to add and remove servers. As traffic increases, new servers can be added to the backend pool. As traffic drops, servers can be removed.

The client doesn't need to know that any of this is happening. From the client's perspective, it still sends requests to the same endpoint.

4. Choosing a Backend Server

Once the load balancer receives a request, it uses load balancing algorithms like Round Robin and Least Connections to decide which backend server should handle it.

We will look at load balancing algorithms in detail in the next lesson.

5. Stateless Servers and Sticky Sessions

One thing that affects load balancing is whether your backend servers are stateless or stateful.

Stateless Servers

In a stateless system, any request can be handled by any server. The server does not rely on local information from previous requests.

This makes load balancing much easier because the load balancer is free to send each request to whichever healthy server is available.

Stateless systems are also easier to scale. You can add or remove servers without worrying about which users are connected to which server.

Stateful Servers and Sticky Sessions

A stateful system is different. Here, a server may store information locally about a user's session.

For example, suppose a user logs in through Server 1 and Server 1 stores that session in its own memory. If the user's next request is sent to Server 2, Server 2 may not know anything about that session.

In that case, the system may need to keep routing that user's requests back to Server 1. This is called session affinity, or sticky sessions.

In general, stateless application servers are preferred in distributed systems because they make scaling and failure recovery much simpler.

When possible, shared state is usually moved outside the application servers into systems such as databases, caches, or distributed storage.

6. Health Checks

A load balancer should never send traffic to a server that is unhealthy or unavailable. This is where health checks come in.

The load balancer periodically checks each backend server to make sure it is still able to handle requests.

A simple health check might send a request to an endpoint like /health. If the server responds successfully within the expected time, it remains in the pool.

If it fails several checks in a row, the load balancer marks it as unhealthy and stops sending new traffic to it. A typical configuration looks like this:

The load balancer continues checking the failed server in the background. Once it starts responding successfully again, it can be added back into the pool.

Health checks can be very simple, such as checking whether a server responds to an HTTP request, or more detailed, such as verifying that the application can also reach critical dependencies like a database.

CheckWhat It VerifiesExample
SimpleThe server responds to an HTTP requestGET /health returns 200
DetailedThe application can also reach its critical dependencies/health returns 200 only if the database query succeeds

7. Layer 4 vs Layer 7 Load Balancing

Load balancers can operate at different layers of the network stack, but the two most common are Layer 4 and Layer 7.

Loading simulation...

Layer 4 Load Balancers

A Layer 4 load balancer works at the transport layer. It makes routing decisions using information like:

  • Source and destination IP address
  • TCP or UDP protocol
  • Port numbers

It does not need to understand the actual application data inside the request.

Because it works with lower-level network information, Layer 4 load balancing is generally fast and works with many different protocols, not just HTTP.

Layer 7 Load Balancers

A Layer 7 load balancer works at the application layer. It understands application-level protocols such as HTTP and HTTPS, which means it can inspect things like:

  • URL path
  • Hostname
  • Headers
  • Cookies
  • HTTP method

This enables much smarter routing. For example, requests to /images can go to an image service, and requests to /payments can go to a payment service.

The tradeoff is that Layer 7 load balancing requires more processing because the load balancer needs to understand the application protocol.

FeatureLayer 4Layer 7
Works atTransport layerApplication layer
Decision inputSource and destination IP, TCP or UDP, portURL path, hostname, headers, cookies, HTTP method
ProtocolsMany, not just HTTPHTTP and HTTPS
StrengthFast, protocol-agnosticSmarter, content-based routing
CostLess processing per requestMore processing per request

8. TLS Termination

Another common responsibility of a load balancer is TLS termination.

When a client connects to an application over HTTPS, the connection is encrypted using TLS.

Instead of making every backend server handle TLS independently, the load balancer can terminate the encrypted connection. The load balancer handles the TLS handshake, decrypts the request, and then forwards it to the appropriate backend server.

TLS handshakeEncrypted HTTPS requestDecrypt requestForward requestResponseEncrypted responseClientLoad BalancerBackend Server
6 / 6
algomaster.io

This gives us a few benefits:

  1. Simpler certificate management: Certificates can be managed in one place instead of on every application server.
  2. Less work on backend servers: Backend servers no longer need to spend resources handling TLS handshakes and encryption for every connection.
  3. Consistent configuration: TLS configuration becomes easier to update and enforce consistently across the system.

After termination, the load balancer can forward traffic to backend servers over plain HTTP if the internal network is trusted. But in many production systems, traffic is encrypted again between the load balancer and the backend servers.

So TLS termination does not necessarily mean encryption stops at the load balancer. It simply means the load balancer becomes the point where the client-side TLS connection ends, giving us centralized control over HTTPS traffic.

9. Connection Draining

Another important feature is connection draining.

Suppose you want to remove a server from the backend pool because you're deploying a new version, scaling down, or performing maintenance.

If you simply remove that server immediately, any requests or connections already being handled by it may get interrupted.

Connection draining avoids this. When a server enters draining mode, the load balancer stops sending it new requests, but allows existing requests and connections to finish.

For example, imagine Server 2 is handling several active requests:

  1. Instead of shutting it down immediately, the load balancer marks it as draining.
  2. New requests are sent only to Server 1 and Server 3.
  3. Server 2 finishes the work it has already started.
  4. Once those requests complete, Server 2 can be safely removed.

This is especially important for long-running requests and long-lived connections, such as WebSockets.

10. High Availability for Load Balancers

An obvious question is: if the load balancer is responsible for keeping our backend highly available, what happens if the load balancer itself fails?

If we only have one load balancer, then it becomes a new single point of failure. That defeats much of the purpose of having multiple backend servers.

To avoid this, production systems usually run multiple load balancer instances. There are two common setups.

Active-Active

Multiple load balancers handle traffic at the same time. If one instance fails, the remaining instances continue serving requests.

Active-Passive

One load balancer handles traffic while another stays on standby. If the active instance fails, the standby takes over.

With managed cloud load balancers, much of this redundancy is handled automatically by the cloud provider, so you usually interact with a single logical endpoint even though multiple load balancer instances may be running behind it.

11. Common Load Balancing Technologies

In practice, you usually don't build a load balancer from scratch. There are several widely used tools and managed services that provide load balancing out of the box.

  • NGINX is commonly used as a reverse proxy and Layer 7 load balancer for HTTP and HTTPS traffic.
  • HAProxy is another popular option, especially when you need high-performance Layer 4 or Layer 7 load balancing.
  • Envoy is widely used in modern microservice architectures and service meshes, where it can handle both external and internal traffic.
  • Cloud load balancers: AWS provides services such as the Application Load Balancer, which operates at Layer 7, and the Network Load Balancer, which operates at Layer 4. Google Cloud and Azure provide similar managed load-balancing services.
  • Kubernetes: Traffic can be distributed using components such as Services, Ingress controllers, and gateways.

The exact technology depends on your infrastructure and requirements. But the underlying ideas remain the same: receive traffic, choose a healthy backend, and distribute requests efficiently across multiple instances.

Summary

A load balancer sits between clients and a group of backend servers. To the client, the system looks like a single service, while servers can be added, removed, or fail behind it.

Load balancers give you horizontal scaling, high availability, better resource utilization, and an easy way to add and remove servers.

Health checks keep traffic away from servers that cannot handle it, and bring them back once they recover.

Layer 4 load balancers route on IP addresses, protocols, and ports, which makes them fast and protocol-agnostic. Layer 7 load balancers understand HTTP, which enables smarter routing at the cost of more processing.

Stateless servers make load balancing simple. Stateful servers may need sticky sessions, so shared state is usually moved into databases, caches, or distributed storage.

TLS termination centralizes HTTPS handling at the load balancer. Connection draining lets a server finish its work before it is removed.

The load balancer itself must be redundant, usually as an active-active or active-passive pair, or as a managed cloud service.

Quiz

What are Load Balancers? Quiz

10 quizzes