Point your AI agent at freehire and let it find you a job.

Get the CLI →

Cohort AI

NewBe an early applicant

Senior Data Engineer

Discussion

Summary

Senior Data Engineer at Cohort AI builds and scales healthcare data infrastructure: designing production ETL/ELT pipelines, data quality/monitoring, and warehousing for clinical and claims data using SQL, Python, PySpark, Apache Spark, and modern data platforms, collaborating with Data Science, Clinical Informatics, and Product teams.

At Cohort AI, We’re looking for a Senior Data Engineer to help build and scale the data infrastructure behind our platform. You’ll work hands-on with large and complex healthcare datasets, develop reliable data pipelines, and collaborate closely with Data Science, Clinical Informatics, Product, and Infrastructure teams.

This is an opportunity for an experienced data engineer who enjoys solving challenging data problems, building production-quality systems, and working in an environment where healthcare data and AI come together.

What You’ll Do

  • Design, build, and maintain scalable, reliable data pipelines.

  • Develop high-performance ETL/ELT workflows using SQL, Python, PySpark, and Apache Spark.

  • Work with complex healthcare datasets, including clinical and claims data.

  • Build ingestion and transformation workflows that are reliable, maintainable, and scalable.

  • Develop and improve data quality, monitoring, alerting, and observability solutions.

  • Troubleshoot and resolve production data pipeline issues using Linux and shell-based tools.

  • Optimize data processing performance and cloud infrastructure costs.

  • Apply strong engineering practices around code quality, testing, version control, and CI/CD.

  • Contribute to data modeling, data warehousing, and distributed data architecture decisions.

  • Work closely with Data Science, Clinical Informatics, Product, Infrastructure, and Commercial teams to translate requirements into effective data solutions.

  • Support customer onboarding and complex data integration initiatives.

  • Participate in design and code reviews and contribute to improving engineering practices.

  • Help identify opportunities to improve the scalability, reliability, and efficiency of our data platform.

What You Bring

  • 5+ years of professional Data Engineering experience.

  • Strong SQL skills and experience working with large datasets.

  • Strong hands-on experience with Python and PySpark.

  • Experience working with Apache Spark and distributed data processing.

  • Hands-on experience with modern data platforms such as Databricks, Snowflake, BigQuery, Redshift, or similar technologies.

  • Experience designing and operating production-grade ETL/ELT pipelines.

  • Strong understanding of data modeling, data warehousing, and distributed data systems.

  • Solid Linux command-line experience, including bash and shell scripting.

  • Experience implementing data quality, monitoring, and alerting frameworks.

  • Experience with Git and CI/CD.

  • Strong problem-solving skills and the ability to independently investigate and resolve complex technical issues.

  • Strong communication skills and the ability to collaborate effectively with cross-functional teams.

Nice to Have

  • Experience working with healthcare data and standards such as OMOP, FHIR, HL7, ICD, CPT, claims, EHR/EMR, or related datasets.

  • Experience with AWS, GCP, or Azure.

  • Experience working with large-scale distributed systems.

  • Familiarity with Airflow, Dagster, Prefect, or similar workflow orchestration tools.

  • Exposure to Generative AI, LLMs, or AI-enabled data applications.

  • Experience working in healthcare, life sciences, health technology, or a data-intensive environment

Skills

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available