ETL Modernization Architect - 1646
Summary
Design and modernize legacy ETL pipelines using Databricks Spark on AWS, replacing IBM DataStage jobs and ensuring scalable, governed data solutions.
This is a remote position.
Onsite/Hybrid/Remote: Remote
Duration: 12 months
Rate Range: $75 in W2
Work Authorization: GC and US Citizens Only
Must Have:
- ETL/ELT architecture and modernization
- IBM DataStage
- Databricks and Apache Spark
- Delta Lake, Unity Catalog, and Photon
- AWS Glue, Redshift, and Lambda
- Python and Unix/Linux scripting
- Terraform and CI/CD
- Large-scale database and data warehouse migrations
- Data governance, quality, and observability
Responsibilities:
- Define the architecture and migration strategy for modernizing legacy ETL and ELT pipelines.
- Assess IBM DataStage jobs, databases, and data warehouses for migration readiness.
- Design scalable data solutions using Databricks Spark on AWS.
- Establish architecture standards, reusable frameworks, and governance controls.
- Design CI/CD pipelines for automated builds, testing, and deployment.
- Lead technical design reviews, migration planning, and artifact validation.
- Define parity, functional, UAT, regression, and performance testing strategies.
- Ensure schema validation, data quality, lineage, security, and production readiness.
- Guide cutover, go-live, hypercare, and operational stabilization activities.
- Oversee the decommissioning of legacy DataStage jobs and related components.
- Create operational documentation and conduct knowledge-transfer sessions.
- Provide technical direction to developers and engineering teams.
Qualifications:
- Extensive experience designing enterprise ETL/ELT architectures.
- Hands-on experience with IBM DataStage and Databricks modernization projects.
- Strong experience with Databricks, Delta Lake, Unity Catalog, Photon, and Spark.
- Strong knowledge of AWS data services, including Glue, Redshift, and Lambda.
- Experience designing large-scale database and data warehouse migration programs.
- Proficiency in Python and Unix/Linux scripting.
- Experience with Terraform and automated deployment pipelines.
- Knowledge of data governance, observability, lineage, and operational readiness.
- Experience defining testing and production-readiness standards.
- Experience working in Agile/Scrum environments and PI planning.
Nice to Have:
- GitLab or Azure DevOps experience
- JIRA experience
- Automated data-testing framework experience
- Experience leading enterprise cutovers and hypercare activities
- Experience mentoring data engineering teams