Point your AI agent at freehire and let it find you a job.

Get the CLI →

Lead AI Data Engineer (Hybrid)

NewBe an early applicant
About The Position

The Lead AI Data Engineer serves as a scientific, technical, and managerial lead for data-heavy research projects. The Lead AI Data Engineer will overlap with a multidisciplinary team of government and contract researchers, academic experts and consumers. The Lead AI Data Engineer will provide technical and management support, oversee project execution, and provide key guidance on data architecture and infrastructure. The team that the Lead AI Data Engineer oversees is responsible for the curation and maintenance of large data pipelines, developing ETL pipelines, defining schemas, identifying bottlenecks, and deriving variables /features directly used in machine learning / AI model development. The Lead AI Data Engineer serves as a versatile position who designs, builds, and maintains the entire data ecosystem. As project aims evolve, collaborators join, and team processes change, an ideal candidate is flexible and can adapt quickly. The ideal candidate will also be mindful of data privacy / security and will adhere to data governance and data management best practices.

This is a full time hybrid remote position working at the Walter Reed National Military Medical Center in Bethesda, MD that will require working in office/on site at least 2 days per week. Background checks will be administered.

About The Program

Sleep Physiology Modeling Project: Sleep & Wearables Operational Readiness for Research & Defense (SWORD) Lab

Salary Range

$155,000 - $193,000. Salaries are determined based on several factors including external market data, internal equity, and the candidate’s related knowledge, skills, and abilities for the position.

Qualifications

  • PhD in a relevant field (e.g., Computer Science, Engineering, Data Science, Biomedical Engineering) required
  • 8+ years experience with multimodal data analysis and data pipeline engineering required
  • Proven experience with multivariate signal processing (e.g., time-series biosensor data)
  • Hands-on experience with relational (SQL) and non-relational (NoSQL) databases
  • Hands-on experience with version control systems (eg, Git) and demonstrated ability to work in (and lead) a collaborative coding environment
  • Solid problem-solving and analytical skills to address complex technical challenges
  • Hands-on experience using Google Cloud Platform (GCP) cloud infrastructure (or equivalent), including setting up and managing cloud-native data warehouses (eg, BigQuery), storage, and compute resources
  • Ability to translate high-level scientific hypotheses into scalable engineering solutions and data products
  • Ability to work in a fast-paced, multidisciplinary, multi-site (sometimes asynchronous) team environment
  • Preferred qualifications: Strong knowledge of sleep science and hands-on experience with handling data from consumer wearable devices (eg, actigraphy, PPG, EEG); familiarity with machine learning workflows, including model development, model tuning, and deploying models at scale; Leadership and/or project management experience with the ability to oversee a team of people ingesting data

Management Responsibilities

  • Foster a collaborative coding and research environment, driving skill development for junior and mid-level data engineers and analysts across the data pipeline
  • Serve as the primary technical liaison to senior management, translating high-level research aims into actionable objectives
  • Communicate team progress, bottlenecks, and milestones

Responsibilities

  • Produce clean, well-documented, efficient code across the entire stack
  • Design, develop, and deploy robust, scalable applications (both front-end interfaces and back-end data pipelines) to support large-scale research
  • Lead the engineering workflows to acquire, ingest, and clean multimodal datasets, ensuring efficient storage, retrieval, and processing of massive datasets (+1million records)
  • Architect and maintain scalable infrastructure to support advanced machine learning models using physiological features and sleep microarchitectures
  • Optimize application performance and scalability through performance tuning, code refactoring, and database optimization techniques
  • Stay updated with emerging industry trends, academic literature, and technologies in data engineering, cloud architecture, and machine learning to continuously improve the lab’s technical capabilities
  • Ensure data integrity throughout engineering workflows. Assist in the preparation of Standard Operating Procedures (SOPs), analytical frameworks, and technical documentation
  • Maintain open communication with leadership, advise on technical processes, and curate progress reports
  • Provide technical support and oversight to team members with less experience
  • Assist in regulatory support, Data Sharing Agreements, and other project documentation

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available