Associate Data Engineer
Summary
Maintains and troubleshoots data pipelines, validates data quality, and supports ETL workflows using SQL and Python for a healthcare-focused product team.
Responsibilities
- Handle support tickets and operational issues reported by internal teams and external partners
- Perform KTLO tasks including monitoring pipeline health and responding to alerts
- Conduct data source discovery, profiling, and documentation of data structures
- Assist with data validation and testing using SQL queries
- Support data quality initiatives by running diagnostics and documenting findings
- Help maintain and improve documentation for existing data systems and pipelines
- Assist senior engineers with debugging data pipeline issues and tracing transformations
- Conduct quality assurance activities and review data outputs
- Perform exploratory data analysis to understand data patterns
- Learn and apply data engineering best practices including Git and code reviews
Requirements
- Bachelor's degree in Computer Science, Engineering, Information Systems, or related field
- Strong SQL proficiency for data exploration and validation
- Proficiency in Python or another programming language
- Basic understanding of data modeling, ETL/ELT concepts, and pipeline architecture
- Familiarity with version control (Git)
- Strong communication and documentation skills
- Analytical mindset and strong problem-solving skills
- Attention to detail and commitment to data accuracy
Preferred Qualifications
- Experience with healthcare data (FHIR, HL7, CCD) or claims data
- Familiarity with cloud platforms like AWS, Databricks, or Snowflake
- Experience with dbt or other transformation frameworks
- Knowledge of orchestration tools like Airflow or Databricks Workflows
- Experience with Apache Spark or distributed computing
- Background in healthcare, pharmaceutical, or regulated industries