Manufacturing · Sri Lanka · delivered 2024

From legacy replication to tested ERP disaster recovery, for a Sri Lankan manufacturing organisation

A business-critical ERP database needed a new disaster-recovery strategy as its replication platform reached end of life. Now: a multi-site SQL Server distributed availability group, a runbook for a real site loss, and a DR test that runs while production stays open to applications.

SQL ServerAlways On availability groupsWindows Server

The problem

A manufacturing organisation was running its ERP on a database the business could not operate without, and the platform doing the replication and backup underneath it was reaching end of life. Replacing one standalone replication product with another would have solved the immediate problem and recreated the long-term one: another separate tool to license, patch, monitor and eventually migrate off again.

The production database was already protected inside its primary data centre by an Always On availability group. What it did not have was protection against losing the data centre — and, more to the point, any practical way to find out whether recovery would work. A single availability group stretched across both sites would have coupled them: site-level work touches the production replicas, and a disaster-recovery test means interrupting the thing you are protecting.

Which is why so many disaster-recovery plans are never tested. If testing costs an outage, the test gets postponed until it is needed, and the first real failover is also the first failover.

What it had to solve

Seven problems, not one

This was never a database installation. Getting an ERP platform to a recovery position the IT team would actually rely on meant answering all of the following, and a design that answers six of them is not a disaster-recovery capability.

  • High availability inside each data centre, so losing one server is not a site event
  • Database replication between two geographically separate sites, without another standalone replication product
  • Application connectivity after a failover — where the ERP application connects once the primary has moved
  • Monitoring that shows whether replication is actually healthy, rather than the absence of an alert
  • A repeatable disaster-recovery test the business could run on a schedule
  • A documented procedure for a genuine site-level disaster
  • A safe way back to normal operation once a test is finished

What we built

Two availability groups, joined. The existing production group kept doing its job inside the primary data centre. At the recovery site we built a second two-node Windows Server failover cluster with its own availability group, so the DR location has local resilience of its own rather than being one server someone hopes will be enough.

The two groups are then connected by a distributed availability group, which replicates between the groups rather than between individual replicas. Each location has its own listener, so applications connect to a name at whichever site they are meant to be using and never need to know which replica is currently primary.

That gives two distinct levels of protection: server-level resilience within each data centre, and site-level replication between them. The distinction matters more than it sounds. A single replica at another site gives you another copy of the database. A disaster-recovery architecture has to answer what happens when servers, database services, application connections or an entire location go away — and they are four different questions.

It also retired the end-of-life dependency rather than replacing it. The remote replication now runs inside the SQL Server high-availability architecture itself, on the database engine the ERP platform was already built on, so there is one technology to operate instead of two.

The part worth copying

A disaster-recovery test that does not cost an outage

Plenty of businesses have a second copy of the database somewhere and call it disaster recovery. The question that matters is when the applications were last run against it. This procedure is what makes the honest answer to that question "last quarter" rather than "never".

Detach the recovery replica

The designated DR replica is separated from its availability group for the duration of the test, which leaves its database recoverable in isolation. Production is untouched and replication inside the production group carries on.

Open it at the recovery site

The database is recovered and brought up for application access, reachable through the DR-side listener — the same connection point a real failover would use, so the test exercises the real path.

Point real applications at it

Test application instances are directed to the recovery environment and put through functional validation. This is the step that turns an infrastructure check into an answer: can the application connect, operate and complete its work from the recovery site?

Rebuild and rejoin

When testing is finished the recovery database is rebuilt from the production copy and returned to the availability group, so the environment ends the test in the same state it started — which is what makes it safe to run again.

And a written procedure for the real thing

Infrastructure without a recovery procedure still leaves the IT team working out what to do during the outage. So the engagement delivered the runbook as well as the architecture: check distributed availability group health, confirm the state of the recovery environment, fail the database workload across to the DR availability group, validate the new primary, redirect the ERP application to the DR-side listener, and release the system for application and integration testing before the business is let back on.

Every one of those is a step someone would otherwise be inventing at the worst possible moment, with the business watching.

Proof it is still working

What the operations team can check, any day

Disaster recovery degrades quietly. The implementation documents the queries that read SQL Server’s own availability-group and HADR views, so the team can confirm the two sides are connected and healthy rather than infer it from the fact that nothing has gone off.

  • Distributed availability group connectivity between the two sites
  • Which replica currently holds the primary role, at each location
  • Availability mode and operational state of every replica
  • Synchronisation health across the distributed group
  • Per-database synchronisation state, not just the group-level summary
  • The last hardened transaction log position, which is what tells you how far behind the recovery site actually is

The outcome

The client moved off an end-of-life replication platform onto a SQL Server-native architecture spanning two data centres: local high availability at each site, cross-site replication between them, defined connection points for production and recovery, monitoring the operations team can read, and written procedures for both a test and a real disaster.

The value was never the extra nodes. It was turning disaster recovery from a line in a document into something the business can exercise on a schedule and watch succeed — so the conversation changes from "we have a DR environment" to "we know how to use it".

How this engagement was run

Delivered as a fixed-price project against written acceptance criteria, to the agreed timeline and without a cost variation. Client not named: the work was done under confidentiality.

Discuss a projectOther engagements