Data Engineer (Databricks/AWS)
Summary
Build and maintain ETL/ELT pipelines on Databricks and AWS for a pharma-focused data team, ensuring reliable, governed data for analytics and science.
- Design, build, and maintain scalable ETL/ELT pipelines (batch and streaming) using Databricks, AWS, and related orchestration tools
- Write and optimize advanced SQL, and build data transformations in Python or Scala
- Integrate external data sources via APIs and manage pipeline orchestration (Airflow, Databricks Workflows, AWS Glue)
- Apply data quality, governance, cataloging, and lineage practices aligned with regulated-industry standards
- Work within GxP-regulated data environments and apply awareness of data privacy/compliance considerations (e.g., 21 CFR Part 11, GDPR where applicable)
- Partner with business stakeholders across the pharma value chain (R&D, Manufacturing & Quality, Commercial, Drug Development) to gather and translate requirements into technical specifications
- Present technical work and data strategy to executive-level audiences
- Prioritize high-impact data initiatives and proactively identify and avoid duplicated data efforts
- Support change management and adoption of new data solutions across business teams
- Help stand up new data domains from scratch (green-field build), not just maintain existing ones
Requirements
- ETL/ELT development (batch and streaming)
- Advanced SQL (joins, window functions, query optimization)
- Python or Scala for data transformation
- Data pipeline orchestration (Airflow, Databricks Workflows, AWS Glue)
- API integration for external data source ingestion
- Databricks (Delta Lake, Unity Catalog, Genie)
- Cloud platforms — AWS (S3, Glue, Athena) and/or Azure/GCP equivalents
- Data warehousing concepts (dimensional modeling, star schema)
- BI/visualization tools (Tableau, Power BI, or similar) to understand downstream consumption
- Data profiling and cleansing techniques
- Metadata management and data cataloging
- Master data management (MDM) principles
- Data lineage tracking
- Data governance frameworks (especially regulated-industry standards)
- Familiarity with GxP-regulated data environments
- Understanding of the pharma value chain (R&D, Manufacturing & Quality, Commercial, Drug Development)
- Awareness of data privacy/compliance considerations (21 CFR Part 11, GDPR where applicable)
- Knowledge of common pharma data domains (clinical, manufacturing, quality, commercial)
- Requirements gathering and translation (business need → technical spec)
- Cross-functional communication (Business ↔ IT)
- Executive-level presentation skills (given EC visibility)
- Change management / adoption support
- Prioritization frameworks (identifying high-impact vs. low-value data asks)
- Cost-avoidance mindset (spotting duplication before it happens)
- Ability to work with ambiguity and evolving priorities
- Agile/Scrum familiarity
- Documentation discipline (data dictionaries, source-to-target mappings)
- Vendor/partner coordination (if external data sources are involved)
- Prior consulting or client-facing delivery experience
- Experience standing up new data domains from scratch (green-field vs. maintenance)
- Familiarity with AI/GenAI-enabled analytics tools