Telecommunications · Pacific · delivered 2021

Migrating a Mission-Critical SQL Server Platform to Azure SQL Managed Instance

A cloud database migration where losing a transaction was not an option. The databases moved, and so did the jobs, the logins and the encryption objects the applications depend on — rehearsed first, cut over with writes stopped and the final state validated, then geo-replicated to a second region that also takes read traffic.

Azure SQL Managed InstanceSQL ServerGeo-replicationAzure Front DoorExpressRoute

The problem

Moving a critical SQL Server database to the cloud is easy to describe. Moving it without losing a transaction, breaking an application dependency, or discovering a missing SQL Server object the morning after cutover is a different project entirely.

A backup and a restore will move a database. What it will not move is the service around it — the scheduled jobs that run the overnight processing, the logins the applications authenticate with, the certificates without which an encrypted database will not open. Those arrive as absences, discovered one at a time, after production is already on the other side.

So the question was never whether a SQL Server database could be restored into Azure. It was whether the complete database service could move and the applications carry on as though nothing had happened.

What had to move with it

The database is the easy part

Each of the following is a dependency that is invisible while it works and obvious the moment it does not. They were inventoried before the migration rather than discovered after it.

  • The production databases themselves, with their performance characteristics intact
  • SQL Agent jobs — the scheduled operational and application processing that nothing points at until it stops running
  • Logins, users and permissions, so applications and the support team authenticate the same way afterwards
  • Encryption certificates and keys, without which an encrypted database is present but unopenable
  • Application connection configuration, and the integrations that reach the platform from outside it
  • Operational settings and database-level configuration that the applications were built against
  • A rollback route, defined before the cutover rather than improvised during it
  • No lost transactions — which constrains how the final cutover can be sequenced

What we built

Azure SQL Managed Instance as the target: a managed Azure SQL service that keeps the broad SQL Server compatibility an established enterprise application estate depends on. Choosing it was the quick decision. Getting production onto it safely was the project.

That started with an assessment of what actually had to move — at the database level and at the instance level. Jobs, logins, permissions, encryption objects and operational configuration were catalogued as migration items in their own right, not as things to tidy up afterwards. It is the step that separates a database that is online from a database service that works, and the one most often skipped because it produces no visible progress.

The target was then placed inside the wider Azure design, with the application tier in front of it and a second Azure region behind it for recovery.

What the database landed in

The managed instance did not arrive on a flat network. A payment workload has an awkward set of neighbours: public web and mobile traffic on one side, an existing data centre on another, and partner banks and remittance networks on a third. All three have to reach the platform. None of them should be able to reach each other.

So the Azure environment was built as separate virtual networks rather than one estate. Public traffic enters through a global front-door service and a web application firewall, then an application gateway and firewall decide which backend it reaches — no application server is directly exposed. The existing data centre connects over a private circuit instead of the public internet, which is what allowed the database to move without every related system moving with it. Partner banks and remittance providers come in through VPN gateways into a network of their own, each able to reach the services its integration needs and nothing else.

Collapsing that boundary is easy at build time and very hard to unpick once a dozen partners are connected through it. Subscriptions, resource grouping, naming, tagging and policy were settled before the first workload moved for the same reason: they are the least interesting part of a cloud migration and the most expensive to change afterwards.

The part worth copying

Rehearse it, then cut over with the writes stopped

Production was not the first test of the migration process. The sequence below ran more than once before the live window, so the cutover itself was the shortest and least interesting part of the project.

Rehearse the whole migration

The full move was performed and validated ahead of the live window — which is what surfaces the hidden dependency, the job that does not exist on the target and the certificate nobody listed. Problems found in a rehearsal cost an afternoon; the same problems found at cutover cost an outage.

Test the service, not the database

Connectivity, application functionality, integration and load testing, then remediation, then testing again. A successful query proves the database is up. It does not prove a telecommunications platform is ready to take traffic.

Stop the writes at a defined point

This is what makes zero data loss a fact rather than a slogan. Application writes are halted at an agreed moment, so there is no window in which a transaction can be committed to a source that is no longer the one being migrated.

Synchronise and validate the final state

The last of the data is moved and checked against the source before anything is redirected. The decision to proceed is made on evidence, at a point where not proceeding is still an option.

Redirect, then watch

Application connectivity and the dependent integrations are pointed at Azure, end-to-end validation confirms transactions complete, and the environment is monitored through the stabilisation period rather than signed off at the change window’s end.

Or go back, on a written procedure

Defined before the migration began: return services to the on-premises platform, restore application connections, redirect the dependent systems. The team knew both routes before it started, which is the only state in which a cutover decision can be made calmly.

A recovery region that earns its keep

The database was then geo-replicated into a second Azure region, giving the operator a geographically separate copy as part of its recovery position rather than everything depending on one location.

The more interesting decision was what to do with it the rest of the time. Most disaster-recovery infrastructure is paid for monthly and used never — a line item defended once a year. Here, eligible read workloads were directed to the secondary region while transactional writes stayed with the primary.

That buys two things from one spend. The recovery replica stays ready, and it absorbs reporting and query traffic that would otherwise compete with production writes. It also makes the read and write paths explicit rather than accidental, which is the groundwork for scaling either of them later. Disaster recovery stops being an insurance policy and becomes part of the production architecture — and a DR environment that is carrying real traffic is one you find out about quickly if it breaks.

The outcome

The measure of this project was never whether a SQL database existed in Azure at the end of the change window. It was whether the business could keep operating from Azure with its transactions, its security, its scheduled processing, its integrations and its recovery capability all intact.

The databases moved with the service around them. The cutover protected every committed transaction because the writes were stopped before the final synchronisation, not after. The route back existed before the route forward was taken. And the recovery region does real work between the incidents it was bought for.

How this engagement was run

Delivered as a fixed-price project against written acceptance criteria, to the agreed timeline and without a cost variation. Client not named: the work was done under confidentiality.

Discuss a projectOther engagements