Data Engineer
Summary
Designs and maintains data pipelines and analytics-ready datasets using Python, Azure Databricks, and SQL, integrating enterprise data sources to support reporting and business decisions.
We are seeking an experienced and highly motivated Data Engineer to design, build, and operate reliable data pipelines and analytics-ready datasets that power reporting and business decision-making. In this role, you will develop and optimize ETL/ELT processes using Python, Azure Databricks, and SQL, integrating data across platforms including Microsoft SQL Server and PostgreSQL. You will partner with analysts, data consumers, and engineering teams to translate business requirements into scalable, well-governed data products, applying strong engineering practices around testing, monitoring, performance, and data quality, while using AI-assisted development tools and generative AI to enhance productivity and solution quality. We are looking for someone with hands-on experience integrating AI capabilities into enterprise data workflows and a strong understanding of responsible AI principles, including data privacy, fairness, and explainability.
Key duties and responsibilities:
Data Pipeline Development (ETL/ELT): Design, build, and maintain batch and/or incremental pipelines using Python, Azure Databricks, and SQL to ingest, transform, and curate data for analytics and downstream applications.
Orchestration, Scheduling & Reliability: Automate and schedule pipeline execution, implement idempotent processing, and build monitoring/alerting and operational runbooks to ensure reliable, observable data products.
Production Operations Support: Own day-to-day production support for data pipelines and reporting workflows, including incident triage and resolution, root-cause analysis, and proactive monitoring to meet SLAs and ensure data reliability.
Databricks Engineering: Develop and optimize notebooks and jobs; implement reusable libraries, parameterized workflows, and cluster/job configurations; apply performance tuning techniques (partitioning, caching, query optimization) as appropriate.
SQL performance optimization: Optimize schemas, tables, views, and SQL logic across Microsoft SQL Server and PostgreSQL. Troubleshoot performance issues and ensure data integrity through constraints, indexing, and query tuning.
Data Quality & Controls: Define and implement validation rules, reconciliation checks, and automated tests to detect anomalies, enforce SLAs, and improve trust in reporting and analytics outputs.
AI Enablement & Productivity: Use AI-assisted development tools and generative AI to improve engineering productivity, data pipeline quality, and documentation while maintaining strong validation and governance practices.
Required experience & competencies
- 5+ years in Data engineering roles.
- Strong Python for ETL, automation, and data transformations.
- Hands-on SQL Server development: schemas, procedures, indexing, optimization.
- Experience with Azure data services and Databricks operations.
- Proven experience with AI-assisted software development tools and workflows.
- Demonstrated ability to use generative AI to improve developer productivity and solution quality.
- Solid understanding of responsible AI principles, including data privacy, fairness, and explainability.
- Hands-on experience integrating AI capabilities into enterprise data or analytics applications.
- Degree in Computer Science, Management of Information Systems, or a related analytical field or equivalent experience.
Nice-to-haves:
- Azure Data Factory
- Experience with data quality testing and validation frameworks.
- Familiarity with Lakehouse concepts and Delta tables.
- Experience with PostgreSQL, including query and schema optimization.
- CI/CD for data pipelines using Azure DevOps or Gitlab or similar.
- Ability to integrate AI/ML model endpoints into data platforms and pipelines using Azure ML or similar.
- Advanced SQL for analytics, modeling, and performance tuning.