Payments · Pacific · delivered 2021

A payment platform going to cloud, with customers, banks and partners attached

Moving the application was only part of the challenge. Public traffic, an existing data centre, banks and remittance partners all needed to reach the new Azure platform — without being able to reach each other. A segmented, multi-region design with private hybrid connectivity, isolated partner access and database disaster recovery.

AzureAzure Front DoorApplication GatewayExpressRouteAzure Site Recovery

The problem

Moving a payment platform to the cloud is not a server migration. Public web and mobile traffic still has to reach the service. Internal systems have to stay connected. Banks and remittance partners need tightly controlled access. The database has to be protected. And if a region goes away, the business still needs somewhere to go.

Payment platforms have an awkward set of neighbours: public users on one side, an existing data centre on another, partner banks and remittance networks on a third. All three need to reach the platform. None of them should be able to reach each other. A lift-and-shift puts them all on one flat network and quietly removes that distinction — which is the one thing a payment environment cannot afford to lose.

What it had to solve

Everything the platform was already connected to

The application was the smallest part of the problem. Each of the following was an existing dependency that had to keep working the day after the migration, and several of them belong to someone else.

  • Public web and mobile users, reaching the service over the internet
  • An existing data centre that was not moving, and systems on it the platform still depends on
  • Banking institutions, each with its own integration
  • Remittance partners, likewise
  • The application and API tier, and the SQL Server databases behind it
  • Production and staging, which needed to stay distinguishable rather than merging in the move
  • Security, governance and operational monitoring, from day one rather than retrofitted
  • A recovery position that survives losing the whole primary region

What we built

Separate virtual networks, not one estate. Public traffic enters through a global front-door service and a web application firewall, then an application gateway and firewall decide which backend it reaches — production or staging, and never an application server exposed directly to the internet. That boundary, between the public internet and the systems processing the transaction, is the highest-risk surface on a payment platform, so it is the one that got the layers.

The existing data centre connects over a private circuit rather than the public internet. That is what made a staged migration possible: the payment platform could move to Azure without every system it integrates with having to move at the same time.

Banks and remittance partners come in through VPN gateways into a network of their own, kept separate from the on-premises path. Each partner reaches the payment services its integration needs and has no route to anything else. The principle is simple and the consequence is not: collapsing that boundary is easy at build time and very hard to unpick once a dozen partners are connected through it.

And somewhere to go if the region fails

Migrating to cloud does not remove the need for disaster recovery; it changes what the recovery unit is. The environment spans a primary and a secondary Azure region. The managed database replicates across them, region-to-region recovery is provisioned for the rest of the infrastructure rather than the database alone, and the front-door service can redirect traffic to the recovery region when it is needed. Private connectivity reaches both, so a recovery does not stall waiting for a network path to be built under pressure.

The outcome is not "the application is now in Azure". It is that the application has somewhere to go when its primary region is not.

The part worth copying

A cutover with a rollback written before it started

A successful VM migration means very little if customers, internal systems or financial partners cannot complete a transaction afterwards. So the sequence proved the whole path before production depended on it — and the way back was documented before anyone needed it.

Build, then prove the connections

Azure resources stood up, the database migration prepared, the application deployed — then connectivity testing to verify that every system and network that has to talk, can.

Test it as a service, not as infrastructure

Functional and integration testing across the application and its connected services, load testing against expected demand, and compliance validation. All of it resolved before the production move rather than discovered during it.

A controlled outage for the data

Application writes stopped, a final backup captured, and the production database moved to the managed service. This is the only part the business feels, and it is short because everything else has already been proven.

Cut over the traffic, then the partners

DNS redirected to the new environment, banking and remittance connectivity repointed to the Azure endpoints, and the remaining on-premises systems and APIs sent to their new destinations.

Validate the whole transaction path

End-to-end verification that a transaction completes, then watch the environment. A green resource list is not the same as a healthy service, and on a payment platform the difference is the whole job.

Or roll back, on a written procedure

Documented before the migration, not during it: stop the Azure application layer, return the database on-premises, restore application connectivity, redirect partners and DNS, validate again. A defined decision path instead of one invented at three in the morning.

Decided before the first workload moved

Governance, because it is the expensive thing to change later

Subscription layout, resource grouping, naming and policy are the least interesting part of a cloud migration and the most expensive to retrofit. Deciding them first is what keeps a cloud estate auditable; deciding them afterwards means re-homing live resources.

  • Management groups and subscriptions separating production from staging, rather than letting one undifferentiated estate grow
  • Resource groups organised by function — networking, security, database, application, management, storage — so operational ownership is obvious
  • A standard naming model and a tagging strategy covering business criticality, owner, application, cost centre, budget, DR classification and environment
  • Identity services, network security groups, a managed secrets store and resource locks applied as part of the build
  • Monitoring, logging, alerting, service health, application insights and network traffic analysis included in the migration scope, not booked as later work
  • A governance framework oriented to ISO 27001 and PCI DSS expectations, because a payment workload will be asked about both

The outcome

The valuable part of this engagement was not moving workloads from a data centre into Azure. It was deciding how a payment platform should be shaped once it got there.

Customers needed access. Banks and remittance partners needed access. The existing data centre needed access. None of those networks needed unrestricted access to each other — and keeping that true, while adding regional protection for the database, monitoring the operations team can read and a governed estate that can be audited, is the difference between a migration and a platform.

How this engagement was run

Delivered as a fixed-price project against written acceptance criteria, to the agreed timeline and without a cost variation. Client not named: the work was done under confidentiality.

Discuss a projectOther engagements