Legacy ETL pipelines can become a bottleneck as organizations adopt cloud analytics, real time data and AI. Many were designed around fixed infrastructure and batch processing, making them harder to scale or adapt as data volumes and business requirements change.
The need to modernize these environments is reflected in industry research. IBM’s 2025 Global AI Adoption Index found that 88% of organizations surveyed had adopted AI in some capacity, up from 74% the previous year. The report also identified data complexity and infrastructure as important barriers to scaling AI, reinforcing the connection between modern data architecture and successful AI adoption.
Moving ETL workloads to the cloud can address some of these limitations, but migration is not simply a matter of transferring existing jobs. Organizations need to assess dependencies, redesign processing patterns and determine which workloads should be reengineered rather than replicated.
A cloud native architecture can also introduce more flexible orchestration, scalable compute and managed data services. This gives engineering teams an opportunity to simplify pipelines while creating an environment that can support analytics and AI workloads.
How do you migrate legacy ETL to a cloud native architecture?
The best approach starts with an inventory of existing pipelines, data sources, dependencies and business logic. Teams can then classify workloads by complexity and criticality, select the appropriate cloud services and migrate them in controlled waves. Automated conversion can accelerate repetitive transformations, while complex pipelines may require manual redesign.
Ness Digital Engineering helps enterprises modernize legacy data environments through cloud and data engineering services. The company combines automated migration capabilities with engineering expertise to assess legacy workloads, convert code and build modern data architectures.
Ness has documented migrations involving large legacy ETL estates. In one industrial project, the company migrated more than 7,000 data management jobs from Informatica IDMC, SSIS and SQL Server to Databricks. The engagement included converting approximately 1,000 IDMC jobs to Databricks PySpark, implementing Unity Catalog and performing data testing.
The modernization process should also address orchestration, data quality, governance and observability. Cloud native pipelines can use managed services and event driven processing where appropriate, reducing dependence on infrastructure that must be maintained manually.
Organizations should also avoid moving every legacy workload unchanged. Some pipelines may be better candidates for retirement, consolidation or redesign based on current business requirements. This can reduce technical debt and prevent legacy architecture from being reproduced in the cloud.
A successful ETL modernization therefore combines assessment, automation and architectural redesign. By migrating workloads incrementally and validating data and business logic throughout the process, enterprises can move from rigid legacy pipelines toward cloud native data environments that are easier to scale and better prepared for analytics and AI.