freehire launches on Product Hunt on 26 August.

Follow →

Data Engineer

About ParetoHealth

ParetoHealth is redefining the way employers fund healthcare. As the largest and fastest-growing benefits captive in the United States, we help thousands of small and midsize employers take control of healthcare costs through a smarter, more sustainable model.

Our mission is simple: give small and midsize employers the scale and protection they need to eliminate volatility and lower healthcare costs.

By combining data-driven insights, innovative risk management, and the collective purchasing power of our community, we enable employers to reduce volatility, improve long-term outcomes, and reinvest savings into their businesses and their people.

Headquartered in Philadelphia, ParetoHealth is growing rapidly and transforming one of the country's largest industries. Our success is fueled by talented people who are united by our four core values: Fire in the Belly, For the Greater Good, See the Field, and Get It Done Right. These values shape how we innovate, collaborate, make decisions, and deliver exceptional results for our clients and one another.

If you're energized by solving complex challenges, thrive in a high-growth environment, and want to help reshape the future of healthcare, we'd love to meet you.

Please note that ParetoHealth does not provide employment visa sponsorship for this position. Candidates must be authorized to work in the United States without sponsorship both now or in the future.

Position Summary:

The Data Engineer will design, build, and maintain serverless data pipelines and data models on AWS that make high‑quality, analytics‑ready data available to our Analytics and AI teams. This role will own ingestion, transformation, and storage patterns using services such as Lambda, Glue, Athena, and S3, ensuring data is reliable, well‑documented, and aligned with business and model‑training needs.

Key Responsibilities:

  • Design, implement, and maintain scalable, fully serverless data pipelines on AWS using Lambda, Glue, Athena, Step Functions, and S3 to support reporting, analytics, and AI use cases.
  • Build and evolve data models and schemas that enable performant querying and downstream consumption by Analytics, Data Science, and AI engineering teams.
  • Develop ETL/ELT workflows to ingest, cleanse, transform, and load data from internal applications, third‑party sources, and event streams into our data lake and analytical layers.
  • Partner closely with Product, Underwriting, Analytics, and AI stakeholders to understand data requirements and translate them into robust data structures, contracts, and SLAs.
  • Implement data quality controls, monitoring, and alerting to ensure accuracy, completeness, timeliness, and lineage of critical datasets and features used by models.
  • Optimize serverless workloads for cost, performance, and scalability, including query tuning in Athena and efficient storage formats/partitioning in S3.
  • Contribute to and enforce data engineering best practices, including version control, code review, CI/CD for data pipelines, and Infrastructure as Code with AWS CDK.
  • Collaborate with Analytics and AI teams to design and maintain feature stores and other reusable data assets that accelerate experimentation and model deployment.
  • Troubleshoot pipeline issues, resolve data‑related incidents, and provide ongoing support for production data workflows and model‑driven applications.
  • Document data models, pipelines, and data contracts, and help evangelize data literacy and self‑service analytics across the organization.

Required Qualifications / Skills:

  • Strong experience with AWS data and serverless services (Lambda, Glue, Athena, S3, Step Functions or similar orchestration tools).
  • Proficiency in Python and TypeScript or a similar language for data engineering, including building ETL/ELT jobs and reusable libraries.
  • Solid understanding of data modeling principles (dimensional, normalized, wide‑table, and event‑driven designs) for analytics and AI workloads.
  • Experience designing and operating data pipelines at scale, including batch and near‑real‑time ingestion.
  • Familiarity with SQL and query optimization in columnar data stores and engines such as Athena.
  • Knowledge of data quality, governance, and security practices, including handling sensitive healthcare and financial data.
  • Hands‑on experience with Infrastructure as Code, preferably AWS CDK, for provisioning and managing data infrastructure.
  • Ability to collaborate with Analytics and AI teams, understand model data needs, and design data solutions that enable experimentation and productionization.
  • Strong communication skills and the ability to translate complex data concepts into clear, actionable language for business stakeholders.

Minimum Requirements:

  • Bachelor’s degree in Computer Science, Engineering, Mathematics, or related field, or equivalent practical experience.
  • 3+ years of experience in data engineering or software engineering roles focused on data pipelines and analytics platforms.
  • 2+ years of hands‑on experience with AWS data services in a production environment (e.g., Lambda, Glue, Athena, S3).
  • Experience building and maintaining data solutions that support analytics, BI, and/or AI/ML initiatives.

Perks & Benefits:

  • Fully paid medical, dental, and vision benefits.
  • Flexible PTO
  • 401k company contribution
  • Tuition reimbursement
  • Professional development allowance
  • Transportation allowance and daily parking reimbursement
  • Engaging hybrid work environment

We are guided by our values:

Fire in the belly

The drive to learn, to improve, and to deliver outstanding value every day.

See the field

The ability to see the big picture and prepare to meet tomorrow’s needs.

Get it done right

The passion to produce at higher rates and to the highest standards.

For the greater good

A united community creating better health benefit solutions for all.

Please note that any communication from our recruiters and hiring managers at ParetoHealth about a job opportunity will only be made by a ParetoHealth employee with an @paretohealth.com address. ParetoHealth does not conduct text message or chat-based interviews. Any other email addresses, agencies, or forums may be phishing scams designed to obtain your personal information.

We will not ask you to provide personal or financial information, including, but not limited to, your social security number, online account passwords, credit card numbers, passport information, and other related banking information until we begin onboarding activities, which will be coordinated by a member of the ParetoHealth People Ops Team with an @paretohealth.com email address.

Disclosures:
ParetoHealth is an Equal Opportunity Employer and does not discriminate on the basis of race, color, religion (creed), gender, gender expression, age, national origin (ancestry), disability, marital status, sexual orientation, or military status, in any of its activities or operations. These activities include, but are not limited to, hiring and firing of staff, selection of volunteers and vendors, and provision of services. We are committed to providing an inclusive and welcoming environment for all members of our staff, clients, volunteers, subcontractors, vendors, and clients.
California Applicants: See Pareto’s CCPA Notice of Collection for California Employees and Applicants for information about how Pareto Captive Services, LLC, Pareto Health, LLC, and Pareto Underwriting Partners, LLC, together with their respective subsidiaries (collectively, “Pareto”) collects and uses personal information submitted by employment applicants.

What this application asks

greenhouse

First Name, Last Name, Email, Phone, Resume/CV, Cover Letter

  • LinkedIn Profile
  • Are you legally authorized to work in the United States? choose one
  • Will you now or in the future require sponsorship to work in the United States? choose one
  • Do you attest to the following statement? Pareto Captive Services (Pareto") is an equal opportunity employer.  Pareto does not discriminate in employment on account of race, color, religion, national origin, citizenship status, ancestry, age, gender, sexual orientation, marital status, physical or mental disability, military status or unfavorable discharge from military service. I understand that neither the completion of this application nor any other part of my consideration for employment establishes any obligation for Pareto to hire me.  If I am hired, I understand that either Pareto or I can terminate my employment at any time and for any reason, with or without cause and without prior notice.  I understand that no representative of Pareto has the authority to make any assurance to the contrary. I attest with my signature below that I have given to Pareto true and complete information on this application.  No requested information has been concealed.  I authorize Pareto to contact references provided for employment reference checks with prior notice to me.  If any information I have provided is untrue, or if I have concealed material information, I understand that this will constitute cause for the denial of employment or immediate dismissal.  Electronic acknowledgement required. choose one
  • Do you have 3+ years of experience in data engineering or software engineering roles focused on data pipelines and analytics platforms? choose one
  • Do you have 2+ years of hands‑on experience with AWS data services in a production environment (e.g., Lambda, Glue, Athena, S3)? choose one
  • Do you have professional experience using SQL and Python/TypeScript to build, maintain, or support production data pipelines or data models? choose one

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available