Skip to main content

Command Palette

Search for a command to run...

Data Migration: Moving Records Is Only Part of the Work

Updated
•5 min read•View as Markdown
Data Migration: Moving Records Is Only Part of the Work

When we talk about data migration, it’s easy to jump straight to the mechanics.

Extract the data. Transform it. Load it into the new system.

Those steps matter. But before choosing how to move the data, I think we need to agree on what it means in the replacement.

A successful load tells us the destination accepted the records. It doesn’t tell us whether the application will interpret them correctly.

The migration approach shapes the data plan

In Part 1, I talked about the constraints of keeping the legacy application running alongside its replacement.

Data is one of the places where that constraint becomes concrete.

A database redesign might make sense for the new application. But if the legacy application still needs the current structure, we have to decide how both systems will continue to operate during the transition.

Will they share a database? Will the replacement have its own data model? Which system can update which information? How will changes reach the other system?

Each choice creates work. I’d want those responsibilities agreed before treating the data migration as a separate implementation task.

Map meaning before mapping columns

A source-to-target mapping should explain more than where a value goes.

Consider a hypothetical legacy field called status. A value might mean “active” in one workflow but “available for a specific group” in another. Moving it into a new boolean field could lose a distinction that the business still needs.

Before simplifying that model, I’d ask:

  • What does each source value mean?
  • Which workflows use it?
  • Are there exceptions or historical values?
  • What should happen when a value doesn’t fit the new model?
  • Who can approve that interpretation?

The transformation rule should follow those answers.

This is also where business analysts and stakeholders are useful. A developer can identify how code uses a value; the business can help decide whether that behavior should continue.

Decide what deserves to move

Preserving behavior doesn’t necessarily mean moving every historical record into the same place.

I’d separate the decisions about active data, historical data, and information that no longer serves the application. That is a business decision as much as a technical one.

For example, an old record might no longer appear in a user workflow but still be needed for reporting. Another might be obsolete, duplicated, or incomplete.

The important thing is to make the treatment explicit. Migrate it, retain it elsewhere, resolve it, or exclude it through an approved decision.

I wouldn’t want a migration script to make that decision accidentally because a record failed validation.

Give synchronization an owner

If users keep working in the legacy application while migration is underway, the source continues to change.

That means an initial bulk load is a starting point.

We need a plan for changes made afterward: newly created records, updates, and deletions. We also need to define when the old system stops accepting changes and when the replacement becomes authoritative.

For me, the most useful questions are simple:

Who owns writes during each stage? How do we identify changes? What happens if a synchronization step fails? How do we know the destination has caught up?

The answers may differ by migration approach. The responsibility for answering them should be clear.

Validate what the business will see

Record counts are useful, but I wouldn’t use matching counts as the definition of success.

Two systems can contain the same number of records and produce different results.

I’d validate at several levels:

Validation What it helps establish
Counts and identifiers Expected records arrived without unintended duplication
Field mappings and relationships Values and links were transformed as intended
Business rules The migrated data produces the agreed outcomes
User workflows Users can find, use, and update the information they need
Exceptions Rejected or ambiguous records are visible and accounted for

For a hypothetical training platform, that could mean checking whether a user’s completion history produces the expected course eligibility. The presence of the completion record alone wouldn’t be enough.

That is why I value testers who understand the business and the legacy behavior. They can help validate the meaning of the migrated data.

Rehearse the transition, including what happens if it fails

I’d want the migration rehearsed with representative data before the final cutover.

A rehearsal should test more than whether the script runs. It should help establish how long the process takes, what fails, how exceptions are handled, and how we verify readiness.

It should also address recovery.

If the replacement has already accepted new writes, switching users back to the legacy application can create another data problem. Where do those new changes go? Can the old system interpret them? Who reconciles the differences?

I’d rather resolve those questions while designing the transition than discover that our rollback plan only covers deployment.

What I would call a successful data migration

For me, success means the replacement contains the information it needs, interprets it correctly, and has a clear owner for future changes.

The load process is part of that outcome. So are the mapping decisions, validation, synchronization, and cutover plan.

That is why I would bring the data discussion into the architecture conversation early. It affects how we build the replacement and how we get users there.


This is Part 3 of the Legacy Application Migration series. Start with Part 1. The final article will cover testing, user readiness, and reducing risk through a soft launch.