data engineer for biomedical research
Summary
Designs, develops, and maintains data pipelines, APIs, and backend services to support biomedical research workflows using Python, FastAPI, and cloud platforms like AWS/Azure/GCP.
Описание:
AstraZeneca develops pharmaceutical products and conducts early drug discovery research. Its computational infrastructure supports biomedical research and the company’s drug discovery pipeline.
Задачи:
- Design, develop, and maintain applications, including APIs, backend services, and user interfaces supporting scientific workflows
- Write clean, efficient, well-tested, and well-documented code following modern software engineering best practices
- Build and integrate databases with efficient data access patterns and API layers
- Maintain and enhance existing applications while developing new features and capabilities
- Build and maintain data processing pipelines that ingest, transform, and integrate scientific data across the organization
- Implement ETL workflows, data validation, and quality checks for reliable data delivery
- Work with various data formats, sources, and storage systems supporting research data needs
- Optimize data pipelines for performance, reliability, and scalability
- Collaborate with scientists to understand computational and data requirements
- Work with senior engineers and the technical lead on design approaches and implementation strategies
- Participate in code reviews, contributing to code quality and knowledge sharing
- Document technical decisions, system architecture, and data workflows
- Troubleshoot and resolve issues across applications and data pipelines independently
- Participate in CI/CD pipeline development and deployment processes
- Collaborate with IT teams on infrastructure, security protocols, and production deployments
- Support incident response and monitoring of production systems
Требования:
- Bachelor’s degree with 5+ years or Master’s degree with 3+ years of professional software development experience
- Demonstrated delivery of production applications or data systems
- Strong proficiency in Python for application development and data processing
- Experience with FastAPI, Flask, Django, pandas, NumPy, and scikit-learn
- Hands-on experience building and maintaining data pipelines, ETL workflows, and data processing systems at scale
- Experience with SQL and NoSQL databases, including schema design, query optimization, and data access layers
- Experience with RESTful APIs, backend services, and application integration with data systems
- Familiarity with AWS, Azure, or GCP, Docker, and Git
- Strong problem-solving and debugging skills across application and data infrastructure
- Good communication skills and ability to collaborate with technical and scientific stakeholders
- Nice to have: experience in scientific computing, bioinformatics, or pharmaceutical/biotech environments with biomedical data
- Workflow orchestration tools such as Airflow, Prefect, Nextflow, or Snakemake
- Data platforms such as Databricks or Snowflake
- Data lakes, data warehouses, and large-scale data storage architectures
- TypeScript/JavaScript and React, Vue, or Angular
- Microservices architecture, API design patterns, and DevOps practices
- Production AI/ML model integration or deployment
- Go, Rust, or C++
- Security best practices and compliance requirements in regulated environments
Условия:
3 Days in office and 2 remote per week.