Spread requests across identical instances, and stop sending them to the sick one.
A layer 4 balancer works at the transport level. It sees an IP address, a port and a stream of packets, and it routes purely on that: no idea whether the payload is HTTP, what path is being requested, or what the request even means. That ignorance is what makes it fast, because there is no parsing and no buffering a full request before deciding where it goes.
A layer 7 balancer terminates the connection and reads the application protocol itself, the host, the path, the method, the cookies. That understanding is what lets it route by content rather than by connection, and it is not free: every request is parsed, and in most designs it is proxied rather than forwarded, meaning the balancer opens its own connection to the chosen backend instead of simply redirecting packets. That costs CPU and adds a hop most people never notice until it shows up in a latency breakdown.
TLS termination placement follows the same split. Terminate at the layer 7 balancer and it can see the decrypted request well enough to route on path or host, while backends run plain HTTP internally and carry no certificate management of their own. Terminate at layer 4, or pass the connection through untouched, and every backend needs its own certificate and its own rotation, in exchange for the traffic never appearing as plaintext anywhere the balancer runs. Re-encrypting to the backend after an L7 termination sits between the two, and it earns its cost when the internal network cannot be trusted outright.
Path-based routing is what actually earns layer 7's cost. Splitting one hostname across several backends by path, sending a canary release to a percentage of requests by cookie, or routing by content type are all decisions a layer 4 balancer cannot make, because it never sees that information. I would not reach for a full HTTP proxy for a single backend with a single purpose. Paying for one anyway is added latency and an added certificate surface bought for a feature nothing is using.
A shallow health check confirms that a process is running and responding, usually a static endpoint that returns success without touching anything else. That tells you the process has not crashed. It tells you nothing about whether the instance can do its actual job: an instance whose database connection pool is exhausted will still answer a shallow check in milliseconds while every real request behind it times out.
A deep health check calls into the dependencies that matter: a database, a queue, a cache, and reports healthy only if those respond too. That is a far more honest signal of whether the instance can serve real traffic. It also has a real cost. The check itself consumes a connection against every dependency it touches, on every instance, on whatever interval the balancer polls at, and each dependency it checks becomes something the instance's own health now depends on.
That dependency is also the cascade risk. If a downstream service is merely slow rather than down, a deep check against it can time out on every instance at once, since they are all checking the same downstream on roughly the same schedule. The balancer then sees the entire fleet fail its health check simultaneously and pulls all of it from rotation, turning one degraded dependency into a complete outage the balancer itself caused, which is a worse outcome than the slow downstream would have produced alone.
Separating liveness from readiness stops two different questions from being collapsed into one check. Liveness asks whether the process should be restarted, and failing it instructs the orchestrator to kill and replace it. Readiness asks whether the instance should currently receive traffic, and failing it only pulls the instance from the balancer's rotation without touching the process at all. Wiring a downstream dependency into liveness rather than readiness is how a slow dependency turns into a restart storm across the whole fleet, destroying warm caches and in-flight work that a simple rotation pull would have left alone.
Stickiness buys the simplest possible migration from a single instance. Pin each client to one backend by cookie or by IP hash, and session state can stay exactly where it already lives, in that process's memory, without standing up a shared store first. Nothing about the application has to change, which is why it is usually the first thing teams reach for.
The cost shows up at deploy. A rolling release eventually has to take every instance out of rotation, including the ones holding sessions nothing else knows about, and there is no gentle way to hand that state to the instance replacing it. Either the deploy waits for stuck sessions to expire on their own, which slows every release down, or it accepts that some sessions get dropped exactly when a deploy runs, which is the worst possible moment for that to happen.
Failover costs the same thing without any warning first. When the instance holding a session fails a health check, everything pinned to it loses that state at once, and the number of people affected scales with how much traffic that instance happened to be carrying. The instance serving the most active users is also the one whose failure does the most damage, the opposite of what you want from a component meant to add resilience.
Moving session state into a shared store, a small cache or database keyed by session id, removes both costs at once. Any instance can then serve any request, so retiring one for a deploy or losing one to failure costs nothing beyond whatever was in flight at that exact moment. The balancer is free to route however is most efficient rather than being constrained to keep every client on the instance it started on.
Draining is the mechanism that makes a rolling replacement actually safe. Before an instance is terminated, it is marked as failing readiness so the balancer stops sending it new connections, while requests already in flight are left to finish normally. Killing a process outright instead cuts every one of those requests off mid-response, turning a routine deploy into a burst of errors for whoever was unlucky enough to be mid-request.
That requires an explicit budget: a drain timeout stating how long to wait for in-flight work before terminating the instance regardless. Too short and slow requests get cut off anyway, which defeats the point of draining at all. Too long and a single long-lived connection, a websocket or a long poll that may never finish on its own, holds up the whole deploy. I set that number from my own measured request duration at a high percentile, not from whatever default shipped with the platform.
Rolling replacement, a few instances at a time behind the same balancer, is what turns a deploy from an event into a background process. Fleet capacity never drops below what traffic needs, because the instances being replaced are drained and removed one small batch at a time while the rest keep serving, rather than every instance going down together and taking capacity with it.
This is where the concept's own caution about taking on that operational complexity only when you need the availability actually gets tested. Standing up a balancer and then deploying by killing every instance at once throws away exactly the availability it was built to provide, so draining and the rolling shape are not optional extras. They are the part of the design that makes the earlier decision to add a balancer worth having made, and I would rather test that draining happens the way the platform claims than assume it because a setting exists somewhere.
Have a product, platform or delivery challenge? Let’s talk about turning it into a structured, scalable solution.
Open to technical leadership, product delivery and senior engineering roles, and available for architecture consulting, technical reviews and mentorship. Engagements run as project-based work, contracts, consulting, freelance engagements, remote collaboration and long-term partnerships.
Based in Cairo, Egypt, working remotely with clients across the MENA region and internationally.