Copy the database so reads can go somewhere the writes are not.
Asynchronous replication is the default because synchronous replication asks the primary to wait for every replica to confirm before it tells the caller a write succeeded, and that wait is added directly to every write's latency, whether or not the write is urgent. Given that trade, most systems replicate asynchronously and accept that a replica sits some distance behind the primary at any given instant, whether that distance is one commit or several seconds during a spike. Lag is the visible cost of that decision. It will never reach exactly zero; it will only be smaller or larger depending on load.
The place this window turns into a support ticket is read-your-own-writes: a user submits a form, the write succeeds against the primary, and the very next request in the same flow reads from a replica that has not applied it yet. From the user's side, the record they just saved has disappeared, not because it was lost but because the read landed on a copy that had not yet been told about it. Anyone who has seen a report of a 'successful' save vanishing minutes later has usually seen exactly this.
Pinning is the fix, and the discipline is in scoping it narrowly. The read that immediately follows a user's own write, typically the confirmation view or the page they are redirected to, should go to the primary or to a replica known to have caught up. Extending that pin to the rest of the session, or to unrelated reads that never touch the record just written, throws away the reason replicas exist in the first place: to keep that traffic off the one node that must also accept every write.
A blanket delay, routing every read to the primary for a fixed few seconds after any write, is a cruder version of the same idea, and it works passably when the write rate is low enough that the primary can absorb the extra reads. It stops working once writes become frequent, because the fixed window then overlaps constantly and most reads end up pinned regardless of whether they touch the record just written. Tracking the actual replication position, and reading from a replica only once it has passed the point the write committed at, scales the same way regardless of how often writes happen, because it pins by need rather than by a schedule.
The question worth asking before a query is routed to a replica is not what table it touches but what the result is used for, because two reads against the same table can carry entirely different tolerance for being a moment out of date. A page showing a product's price to someone browsing can tolerate however long the replication window happens to be. The same table, read at the instant of checkout, where the customer is about to be charged that price, cannot.
Routing by table rather than by intent is common precisely because it is a mechanical rule that touches no application code, but it treats every query against a table as equally safe or equally risky. That is rarely true, and it produces one of two failures: either the routing is conservative and pins the whole table to the primary, losing replication's benefit for that table entirely, or it is permissive and sends the checkout price lookup to a replica alongside the browsing one, because from the table's point of view they look identical.
The classification that holds up is by the read's purpose within its own request. Does the answer feed a decision that must be current, authorising an action, computing a charge, confirming an inventory count before reserving it, or does it feed a display the user will simply refresh if it turns out to be stale, a listing, a dashboard, a report. The former belongs on the primary, or on a replica whose lag is being actively measured against a bound; the latter is exactly what replicas were built to carry.
That classification is a property of the call site, not of the table, which means it has to be made explicit in the code issuing the query rather than inferred later from wherever the query happens to live. A data access layer that defaults every read to a replica unless the caller opts into the primary puts the burden of remembering in the wrong place, on whoever writes the next feature, rather than on the one place the decision is made once.
Promotion, turning a replica into the new primary, sounds mechanical: point the application at a different host, and let the old primary rejoin later as a replica if it comes back at all. In practice it is a sequence with a strict order. The replica has to stop accepting changes from the old primary, catch up any queued writes, and only then start accepting writes of its own. Skipping a step under pressure recreates exactly the inconsistency replication exists to prevent.
Split-brain is the failure mode that appears when that ordering is not enforced. The old primary and the newly promoted replica both believe themselves to be the writable copy at the same moment, usually because a network partition made the old primary unreachable to the monitoring system rather than genuinely dead. Two nodes accepting writes independently is worse than either being briefly unavailable, because the two histories then have to be reconciled by hand, and there is no general rule for whose write wins when both are technically valid.
A failover procedure that has only ever been read, never run against a real cluster, is not one I would call ready. The steps that fail in practice are rarely the ones written down: a connection string cached in a background worker that never re-reads configuration, a monitoring check hard-coded to a hostname instead of a role, an assumption that DNS updates faster than a client that already holds a healthy connection open. A rehearsed failover surfaces those problems before the night they matter; an unrehearsed one surfaces them during the incident, in front of whoever is on call.
Automating promotion removes the delay of a person noticing and running the procedure by hand, and it adds its own failure surface: the coordinator making the promotion decision needs its own view of which node is reachable, and a coordinator confused about that is exactly how automation produces split-brain faster than a person would have. I would rather run a slower, rehearsed manual failover than a fast automated one nobody has tested against a real partition, and treat automation as something to add once the manual path is proven, not as a substitute for proving it.
Adding replicas to a primary struggling under read load fixes the problem precisely when the reads are already efficient and there are simply too many of them for one node. It does nothing for a query that is slow because it lacks an index, or because a loop issues one database call per row instead of one call for the whole set. Sending that same expensive query to three replicas instead of one does not make it cheaper. The same waste simply runs in three places, and the aggregate load on the fleet goes up rather than down.
The tell is in the ratio between load and result. A query that consumes disproportionate database time for what it returns, a handful of rows taking as long as a full report, or a page load that triggers dozens of round trips, is a candidate for an index, a rewrite or a cache, not for a bigger read tier. Reading the slow query log before reaching for replication answers this directly, and it is a far cheaper measurement to take than a migration is to run.
The concept's own caution deserves a sharper edge here: adding replicas over a missing index does not remove the problem. It disperses the symptom across more machines, so each one looks less alarming on a dashboard that measures nodes individually, while the total database time consumed across the whole fleet has not dropped and may well have grown. The metric everyone is watching improves. The metric that matters, total time spent doing needless work, does not.
Before I add a replica for read scaling, I want a specific number in hand: which queries dominate total database time, and whether they are dominated by frequency or by per-execution cost. A query that runs rarely but slowly is a different problem from one that is cheap but constant; the first wants a rewrite or a cache, the second is the more honest case for spreading load across replicas, because there is no per-query fix left to apply once the query is already efficient. Replication earns its place once that distinction has been made, not before it.
Have a product, platform or delivery challenge? Let’s talk about turning it into a structured, scalable solution.
Open to technical leadership, product delivery and senior engineering roles, and available for architecture consulting, technical reviews and mentorship. Engagements run as project-based work, contracts, consulting, freelance engagements, remote collaboration and long-term partnerships.
Based in Cairo, Egypt, working remotely with clients across the MENA region and internationally.