Telecommunications · Pacific · delivered 2025

From standalone MySQL to an always-on, multi-site platform, for a Pacific telecommunications operator

Single points of failure removed from the database, the routing layer and the recovery position at once: group replication with automatic primary election, redundant routers behind a floating connection point, a replica at a second site, and a migration whose validation gate sat before the cutover.

MySQL EnterpriseGroup ReplicationMySQL RouterKeepalived

The problem

For a telecommunications operator, database availability is not an IT statistic. When the database behind a critical application stops answering, the effect reaches operations, customer-facing services and revenue systems within minutes — and it does so in a market where there is often no second provider for customers to fall back on.

This operator was running a business-critical application against a standalone MySQL server. Every failure scenario ended the same way: somebody recovering a database by hand while the service was down.

The brief was not "add a replica". A second copy of the database answers exactly one of the ways this environment could fail, and the others would have been left untouched.

What it had to solve

Eight problems, not one

Availability is not a feature you switch on. Each of the following is a separate way the service could stop, and a design that answers six of them still leaves two ways to take the business offline.

  • Survive the loss of an individual database server without manual recovery
  • Choose a new writable primary automatically, rather than waiting for someone to be woken up
  • Stop applications being tied to one named database server
  • Remove the routing layer as a single point of failure in its own right
  • Keep a current replica at a separate recovery location
  • Move the existing databases into the new cluster without losing anything on the way
  • Keep a real backup and recovery strategy, because replication is not a backup
  • Leave the operations team written procedures for failover, recovery and routine administration

What we built

Group replication in single-primary mode on the InnoDB engine: three members in the primary environment and a fourth at the recovery site. One member accepts writes; the rest stay read-only. If the primary leaves the group unexpectedly, the group elects an eligible member to replace it and that member becomes read/write while the others carry on as secondaries. Database availability stops depending on the health of one server.

That is half a solution. An application pointed directly at a specific host does not benefit from a successful database failover — it reconnects to a server that is no longer the primary and gets errors instead of service. So applications connect through MySQL Router, which knows the role of each member and sends read/write traffic to whichever one currently holds it.

Which moves the problem rather than solving it, unless you deal with the router too. There are two router instances, with Keepalived holding a floating connection point in front of them: if the active router goes, the address moves to the survivor and the application path stays up.

The result is protection at two distinct levels. The database layer elects a new primary. The routing layer moves the application-facing connection between redundant routers. Both have to hold for the service to stay up, and each is now someone else’s single point of failure, not this operator’s.

And a replica outside the building

Local high availability answers server failure. It does nothing about losing the site. A fourth member at the disaster-recovery location takes replicated changes from the group and normally stays read-only, available for promotion if the primary site becomes unavailable.

Those are genuinely two different capabilities, and conflating them is how estates end up discovering that their "high availability" was only ever server-level. This operator has both, and knows which one answers which question.

The part worth copying

A migration whose validation gate came before the cutover

Building the cluster was the straightforward half. Moving live databases into it without losing anything is where these projects actually fail — so the sequence was designed so that the cutover is the least interesting step in it.

Assess, then back up

The existing databases, the MySQL configuration and the application dependencies are inventoried first, including confirming that the source tables meet what group replication requires. Then a full backup of the existing environment, before anything is touched.

Build the cluster and prove it

The group replication environment is stood up and exercised in full before production has anything to do with it, with global transaction identifiers handled deliberately so the cluster knows exactly what it has already applied.

Dry run with test workloads

Real workloads are put through the new cluster to validate its behaviour under something resembling use, rather than inferring it from the fact that the nodes are green.

Validate the data, then cut over

Data is verified across the cluster members and replication health confirmed — and only then are applications redirected through the router. The gate sits before the cutover, which is the whole point.

Prove the service path, end to end

After the move: every member confirmed online, replication health rechecked, monitoring enabled, a fresh backup of the new cluster taken, and a primary-role change tested on purpose. That validates application to router to current primary, not just the data migration.

What the operations team got

An operational service, not a pile of servers

A high-availability platform nobody has been shown how to run is a high-availability platform that gets switched off during the first confusing incident. Backups sit alongside replication here rather than being replaced by it — replicating a database is not the same as being able to restore one.

  • Written procedures for starting and stopping the cluster, and for bringing group replication up and validating it
  • Router operations and floating connection-point validation, so the layer applications depend on can be checked independently
  • Both automatic and deliberate failover: the primary role can be moved between members for maintenance or testing, not only during an emergency
  • Full and incremental backups through MySQL Enterprise Backup, with database and backup storage kept apart
  • Cluster health monitoring, so the state of replication is something the team reads rather than assumes
  • Routine maintenance procedures, written down before they were needed rather than during

The outcome

The valuable thing here was never four MySQL servers running group replication. It was removing, one at a time, every place where a single failure could interrupt application access — the database server, the routing layer, and the site itself — and then building the migration and the operating procedures around that architecture rather than bolting them on afterwards.

For a Pacific telecommunications operator, that is the difference between a server failure and an outage. And when something larger goes wrong, there is a defined path to recovery that somebody has already walked.

How this engagement was run

Delivered as a fixed-price project against written acceptance criteria, to the agreed timeline and without a cost variation. Client not named: the work was done under confidentiality.

Discuss a projectOther engagements