Senior Data Scientist, Medical Imaging
Summary
Senior Data Scientist at HeartFlow (San Francisco, hybrid) who builds analytical pipelines to quantify how real-world medical imaging variability affects deep learning model accuracy, harmonizes data factors, and develops annotation protocols for training and FDA submissions. Core stack: Python statistics ecosystem, PyTorch, and medical imaging (CT) data.
Heartflow is a medical technology company advancing the diagnosis and management of coronary artery disease, the #1 cause of death worldwide, using cutting-edge technology. The flagship product—an AI-driven, non-invasive cardiac test supported by the ACC/AHA Chest Pain Guidelines called the Heartflow FFRCT Analysis—provides a color-coded, 3D model of a patient’s coronary arteries indicating the impact blockages have on blood flow to the heart. Heartflow is the first AI-driven non-invasive integrated heart care solution across the CCTA pathway that helps clinicians identify stenoses in the coronary arteries (RoadMap™Analysis), assess coronary blood flow (FFRCT Analysis), and characterize and quantify coronary atherosclerosis (Plaque Analysis). Our pipeline of products is growing and so is our team; join us in helping to revolutionize precision heartcare.
Heartflow is a publicly traded company (HTFL) that has received international recognition for exceptional strides in healthcare innovation, is supported by medical societies around the world, cleared for use in the US, UK, Europe, Japan and Canada, and has been used for more than 750,000 patients worldwide.
We are looking for a Senior Data Scientist with a strong foundation in both statistical data analysis and deep learning to drive our data-centric AI initiatives. In this role, you will be at the forefront of understanding how diverse real-world imaging conditions and technical variables affect image appearance and, in turn, the accuracy and reproducibility of deep learning algorithms.
You will own the analytical pipeline to quantify these effects and explore data-driven methods to address them. You will work cross-functionally to translate these insights into downstream algorithm development and product decision-making, including development of annotation protocols to curate high quality annotations for training and validation. If you are drawn to deriving insight from large and messy real-world imaging data, and to communicating and crystalizing solutions from these insights, this is the role for you.
Key Responsibilities
- Tooling & Pipelines: Build robust, reproducible analytical pipelines and visualizations over population-scale data, to characterise dataset distributions and model vulnerabilities across diverse patient populations.
- Correction & Harmonization: Explore, develop, and validate methods for harmonizing complex data-related factors and image variations.
- Data-Centric Deep Learning: Translate findings about data variance and model behaviour into actionable data curation requirements, training-time robustness strategies, and architectural recommendations for downstream algorithm development.
- Data Curation and Annotation: Develop annotation protocols for algorithm training and validation for internal product development and FDA submissions.
- Cross-Functional Collaboration and Communication: Partner with Research Scientists, Machine Learning Engineers, Systems Engineers, Process Engineers, Product and Regulatory teams. Derive and present clear analyses from messy data to drive decision-making, and provide artifacts other functions consume.
Required Qualifications
- Education: Masters or PhD Degree in Data Science, Computer Science, Medical Image Analysis, Statistics, Biomedical Engineering, or a related quantitative field.
- Experience: 5+ years (or 3+ with a PhD) of industry experience in Data Science, Machine Learning, or Image Analysis.
- Measurement Science: Working command of reproducibility and agreement statistics — variance components, Gage R&R, intraclass correlation, repeatability and reproducibility coefficients, Bland-Altman — and the judgement to separate correctable bias from irreducible variance.
- Medical Imaging Expertise: Deep understanding of medical image data structures and the physical/clinical realities of imaging. Familiarity with image processing tools and building algorithms for medical imaging data.
- Data Analysis & Statistics: Expert proficiency in Python and statistical data analysis ecosystems (e.g., pandas, scipy, statsmodels, seaborn/matplotlib). Proven ability with large, complex, and messy datasets.
- Deep Learning Experience: Hands-on experience developing or fine-tuning deep learning algorithms (preferably using PyTorch) for computer vision tasks (segmentation, classification, detection) applied to medical images.
- AI-Augmented Workflow: Demonstrated proficiency using modern agentic tools and LLMs (e.g., GitHub Copilot, Gemini, Claude) as a daily force multiplier to accelerate software development, rapidly prototype data solutions, and build reproducible pipelines.
- Communication: Exceptional ability to distill complex, multi-dimensional data analyses into clear, strategic insights for cross-functional stakeholders.
Preferred Qualifications
- Published research specifically related to domain generalization, image harmonization, or out-of-distribution (OOD) detection in medical imaging.
- Experience working with large multi-vendor CT imaging datasets, with an understanding of acquisition and reconstruction protocols.
- Proven track-record in diagnosing data- and annotation-related algorithm performance gaps, and designing data-driven solutions.
- Experience with large-scale data querying and cloud storage (e.g., AWS, SQL).
- Experience developing SaMD products and contributing to regulatory filings.
- Experience with biostats to support FDA submissions.
A reasonable estimate of the base salary compensation range is $170,000 to $240,000, bonus, and equity. #LI-IB1 #LI-Hybrid
Skills
As published by greenhouse · 19 questions · 1 written answer
Basics
First Name, Last Name, Email, Phone, Resume/CV, Cover Letter, Location
Short answers (6)
- Preferred First Name optional
- LinkedIn Profile optional
- Website optional
- What are your salary expectations for this role?
- If answered Yes, please provide the name of the relative and the name of the organization in which they are employed. optional
- If answered yes, please provide the name of the relative. optional
Pick from a list (12)
- Are you legally authorized to work in the United States?
- Will you now or in the future require sponsorship for employment visa status (e.g. F-1 STEM OPT, H-1B, TN, E-3, O-1)?
- How did you hear about this job? optional
- Are any of your immediate family members practicing health care professionals that may use or purchase Heartflow products?
- Do you have any immediate family that work at Heartflow?
- Have you ever been or currently debarred by the U.S. FDA or excluded by the OIG?
- Do you have a Masters or PhD Degree in Engineering, Computer Science, Statistics, Data Science or related degree?
- Do you have 5+ years (or 3+ with a PhD) of industry experience in Data Science, Machine Learning, or Image Analysis?
- Which state do you currently reside in?
- Are you currently based in San Francisco Bay Area?
- If not local to the San Francisco Bay Area, are you willing to relocate to the San Francisco Bay Area?
- As part of this position, you will be required to be in the San Francisco office 3 days a week. Are you ok with this requirement?
Written answers (1)
- If yes, please provide details