Telecommunications · Australia · delivered 2018

Migrating a Business-Critical Oracle RAC Estate with Minimal Cutover Risk

Replacing the storage under a business-critical Oracle RAC estate put databases, applications, failover and disaster recovery at risk at once. The replacement cluster was built alongside production and tested against real applications first, with Data Guard closing the gap at cutover and the old cluster kept as the way back.

Oracle DatabaseOracle RACOracle ASMOracle Data GuardOracle Enterprise Manager

The problem

An Australian telecommunications business was running a business-critical Oracle RAC estate on NFS shared storage, and a new enterprise SAN was arriving. On paper this is a storage refresh. In practice, changing the storage underneath a RAC estate changes almost everything that sits on top of it.

The databases move. The cluster layer moves with them. Application connectivity has to be repointed. The Data Guard relationships protecting the estate have to be rebuilt. Monitoring and the management repository have to follow, or they end up stranded on storage that is being decommissioned. And every one of those is a place where a migration becomes an outage.

So the question was never "how do we move the files". It was how to move the platform while proving the applications, RAC failover, database performance and the recovery architecture all still work — before the business depends on the answer.

What it had to solve

Everything the storage was sitting under

The estate carried production, non-production, monitoring and disaster-recovery workloads. Each of the following had to survive the move, and several of them are only discovered to be dependencies when they break.

  • Oracle RAC and the grid infrastructure underneath it
  • Shared storage, restructured through ASM rather than carried across as-is
  • The production databases themselves, with their performance characteristics intact
  • Application connectivity, which is pointed at the cluster rather than at a database
  • Backup and recovery, reconfigured on infrastructure that did not exist yet
  • Data Guard replication, and the disaster-recovery site in a second Australian data centre
  • The Oracle Enterprise Manager repository, which monitors everything else and is easily forgotten
  • Patch-level consistency, so a failover never becomes an unplanned version change
  • Oracle licensing exposure, which infrastructure placement can quietly increase

What we built

A new cluster, not an in-place conversion. New virtual machines, grid infrastructure, RAC and ASM installed fresh, and the database software deployed at patch levels matching the existing environment — so the migration moves the data without also moving the version.

The new SAN was presented as shared raw devices and organised into ASM disk groups by function: database files, redo, archive and the recovery area, cluster voting storage, and dedicated groups for the larger workloads. That is the part that makes the difference between new storage and a storage architecture.

Because the replacement was built beside production rather than on top of it, the existing cluster stayed available and untouched throughout. The estate was never in a half-migrated state, and the project could be stopped at any point before cutover with nothing lost. That single decision is what turned a high-risk storage change into a controlled platform transition.

One more thing shaped the design: where the new virtual machines were placed. Oracle licensing follows infrastructure, and a technically valid architecture is not the right architecture if it creates commercial exposure nobody budgeted for. Placement was specified to keep the new machines within the existing RAC licensing boundary.

The part worth copying

Test first. Cut over second.

Most database migrations find out whether the new platform works after production is on it. This one was sequenced to answer that question before the cutover window opened — which is also what makes a rollback plan realistic rather than theoretical.

Restore production copies into the new cluster

Copies of the live databases are restored onto the new RAC and ASM platform and opened for the client’s own application teams. Not a smoke test: real applications against real data on the real target.

Prove five things, not one

Migrated data, SQL and query performance, application performance, application connectivity, and RAC failover behaviour. The platform does not progress until the client confirms all five have passed.

Close the gap with Data Guard

Only then are fresh production backups restored and Data Guard configured, so transaction changes flow from the live database to the new platform. The new cluster tracks production right up to the switchover instead of being a copy that is already stale.

Cut over on the business’s terms

Databases move individually or as a coordinated group, so the sequence follows business risk and application readiness rather than forcing every workload into one all-or-nothing event. Applications are redirected to the new SCAN listeners, which is the only interruption users see.

Keep the way back

After a database switches over, the original cluster becomes its standby. The migration was not built on the assumption that nothing would go wrong — it was built so the recovery path existed before the change started.

And disaster recovery, rebuilt as part of the same job

The estate had Data Guard protection between two Australian data centres, and moving production breaks it. Rather than leaving that as a follow-on project, rebuilding the standby databases on the new storage and re-establishing Data Guard was planned into the migration itself.

Including the uncomfortable part: the plan identified the window after cutover during which disaster-recovery protection would not be available, until the new standby was rebuilt. That exposure was known, sized and sequenced in advance. The alternative is discovering it afterwards, which is how an organisation ends up unprotected without having decided to be.

The management layer moved too. The Oracle Enterprise Manager repository was relocated onto the new storage under its own controlled procedure — snapshot, new filesystem, shutdown, file movement, restart, agent connectivity confirmed, fresh backup — so monitoring did not stay behind on storage that was being retired.

Staged, so one thing changes at a time

Four stages with a validation point between each

Delivered as an operational platform rather than as installed binaries: RMAN backups, monitoring and alerts, management agents, OEM integration, RAC verification and as-built documentation were all part of the scope.

  • Stage 1 — production RAC: build the new cluster and ASM storage, restore, synchronise with Data Guard, validate, cut over
  • Stage 2 — Oracle Enterprise Manager: move the repository database and its storage across
  • Stage 3 — disaster recovery: rebuild the standby environment and restore Data Guard protection
  • Stage 4 — non-production: migrate the remaining workloads and their storage
  • Backups configured and monitoring live on the new platform before it carried production, not after
  • As-built documentation handed over, so the estate is operable by the people who own it

The outcome

At a glance this was a move from NFS to SAN. In practice the storage changed, the RAC platform changed, the cluster connection point changed, the production databases moved, the Data Guard relationships changed, disaster recovery had to be rebuilt and monitoring had to follow — and none of it was allowed to become an unacceptable outage.

What made that possible was sequence. The replacement was built and proven alongside production, the final data was synchronised rather than copied, the rollback path was designed before the change began, and recovery was rebuilt as part of the project instead of after it. The client did not migrate and then find problems; they found the problems while production was still running somewhere else.

How this engagement was run

Delivered as a fixed-price project against written acceptance criteria, to the agreed timeline and without a cost variation. Client not named: the work was done under confidentiality.

Discuss a projectOther engagements