Data Scientist
Summary
Build statistical and ML models on large pharma datasets in Databricks/AWS to inform business decisions, partnering with cross-functional teams.
- Develop statistical models, machine learning models, and advanced analyses using large-scale datasets in Databricks
- Access, prepare, and engineer features from data processed through Apache Spark and AWS data services
- Partner with business groups to understand pharma-specific problems and translate them into data science approaches
- Build and validate ETL/data pipelines as needed to support modeling and experimentation workflows
- Communicate modeling results, insights, and recommendations clearly to both technical and non-technical business stakeholders
- Apply pharma domain knowledge to ensure models and analyses are relevant and interpretable in a business context
- Collaborate with data engineers and analysts to productionize models and integrate outputs into reporting/decision workflows
- Monitor model performance over time and iterate as needed
Requirements
- 3–5 years of data science / applied statistics / machine learning experience specifically within the pharma industry
- Hands-on experience with Databricks for data science/ML workflows
- Working knowledge of AWS data services
- Strong experience with Big Data processing using Apache Spark, PySpark.
- Experience building ETL pipelines to support data science workflows
- Strong Python and SQL skills; experience with ML libraries
- Demonstrated ability to understand pharma business needs and speak to pharma business groups
- Strong communication skills with demonstrated ability to present technical findings to business stakeholders
- Bachelor's or master's degree in data science, Statistics, Computer Science, or a related quantitative field
- Experience with Delta Lake, Snowflake, or similar modern data platforms
- Familiarity with MLOps practices and model deployment/monitoring
- Prior experience supporting pharma commercial, clinical, or R&D data science functions