Migration Considerations
In brief: The most common source of MDM migration budget overruns is remediating data quality problems that were not fully surfaced in pre-migration profiling. Profiling typically samples a subset of records — often the most structured and well-maintained subset, which understates the quality problems of the full population. When migration encounters the tail of the…
Module support notes
The Hidden Cost of Poor Migration Planning
The most common source of MDM migration budget overruns is remediating data quality problems that were not fully surfaced in pre-migration profiling. Profiling typically samples a subset of records — often the most structured and well-maintained subset, which understates the quality problems of the full population. When migration encounters the tail of the quality distribution — the oldest records, the records from the least-maintained source systems, the records created before any quality standards were in place — the remediation effort expands significantly beyond the planned estimate.
Organizations that conduct exhaustive pre-migration profiling across the full record population, including the oldest and least-accessed records, produce significantly more accurate migration cost estimates and avoid the budget crises that derail implementation timelines.
Warning
Remediation is the largest migration cost component — and the one most frequently underestimated in initial planning. In a typical Phase 1 migration, remediation of pre-migration quality issues accounts for approximately 35% of total migration cost, more than ETL development, entity resolution, and cutover combined. Accurate migration cost estimation requires an honest, exhaustive data quality assessment before planning begins — not a sample-based profile of the best-maintained records.
Entity Resolution for Migrated Historical Data
Resolving entities across historical migration periods is significantly more complex than resolving current records because entities change over time — customers move, change names, change email addresses, and change employment status. Standard entity resolution configurations optimized for current records frequently mishandle historical records where the identifying attributes no longer match the current golden record.
Organizations migrating historical data for AI training purposes need to invest specifically in historical entity resolution — extending matching rules to account for temporal variations in identifying attributes, and validating resolution accuracy against known historical entity relationships rather than assuming that current-record matching configurations will generalize to historical data.
Note
Historical entity resolution is more complex than current-record resolution — plan for it explicitly. A churn prediction model trained on fragmented historical records that were not correctly linked to their golden entity will learn the wrong lifetime value patterns, producing systematically biased predictions. Historical entity resolution is not a data housekeeping task — it is a prerequisite for AI training data that correctly represents customer history.
Migration is the transformation that makes AI training on governed data possible — treating it as a first-class implementation workstream rather than a background loading exercise determines whether the data that emerges is AI-ready or merely present.
Lesson progress
0% watched