Point your AI agent at freehire and let it find you a job.

Get the CLI →

EPAM Systems

NewBe an early applicant

Lead AI Engineer with Spark, AWS Services

Posted Updated 1 view
Discussion

We're looking for a Lead AI Engineer skilled in Spark and AWS Services to become part of the RBQM Production Pod within the program. In this position, you'll construct and sustain data pipelines that drive AI/GenAI applications supporting Risk-Based Quality Management for clinical trials. The primary focus of this role includes RAG document ingestion, vector indexing, and developing data APIs for AI applications.

Responsibilities

  • Architect and construct RAG document ingestion pipelines (chunking, embedding, vector indexing) to support clinical trial quality data
  • Establish and oversee vector databases (AWS OpenSearch) to support RAG-driven AI workflows
  • Create batch and streaming ETL/ELT pipelines from the ground up for unstructured clinical data (PDF, DOCX, clinical reports)
  • Construct and expose data APIs that AI applications can consume
  • Enhance chunking strategies, embedding generation, and retrieval performance within RAG architectures
  • Oversee data quality, lineage, and governance across AI/ML data pipelines
  • Set up and sustain AWS data infrastructure (S3, Lambda, Glue, Athena, Step Functions, DynamoDB)
  • Partner with Data Scientists and Backend Developers as part of a unified pod team

Requirements

  • Minimum 7 years of practical, large-scale data engineering experience
  • Strong background in RAG document ingestion pipelines (chunking, embedding, vector indexing)
  • Skilled in using AWS OpenSearch as a vector database for RAG workflows
  • High-level command of Python, along with SQL and Spark SQL
  • Experience transforming unstructured data (PDF, DOCX) for use in RAG/LLM applications
  • Working knowledge of AWS Services: S3, Lambda, Glue, Athena, Bedrock, Step Functions, API Gateway, CloudWatch, DynamoDB
  • Understanding of Docker-based containerization
  • Ability to develop custom pipelines from the ground up, going beyond simple configuration of pre-built services
  • Proficiency in English at a B2+ level

Nice to have

  • Experience within the pharmaceutical or life sciences sector
  • Exposure to Snowflake and Pinecone (as an alternative vector database)
  • Understanding of SageMaker processing jobs
  • Proficiency with CI/CD tools (Jenkins, Git/Bitbucket) and infrastructure-as-code tools (CDK or Terraform)
  • Familiarity with clinical data standards (CDISC, ADaM, SDTM)

Benefits

CONTINUOUS UPSKILLING, LEARNING & DEVELOPMENT

  • Diversity of tasks and projects
  • Assessment center for objective review of competency level
  • Personal development plan
  • Mentoring programs and leadership development
  • Certification and professional development support
  • Access to learning platforms including more than 2,500 internal courses
  • English courses taught by certified teachers

CORPORATE BENEFITS

  • Extra leave days
  • Referral bonuses

COMPENSATION PACKAGE

  • Competitive compensation paid in USD
  • Regular salary and performance reviews

MEDICAL & HEALTHCARE

  • Private health insurance
  • Well-being events

WORKING ENVIRONMENT

  • Recreation areas and kitchens
  • Tea, coffee and snacks
  • Sports equipment and game consoles
  • IT Equipment
  • Microsoft’s Software Assurance Home Use Program (HUP)

Skills

What Lead AI Engineering jobs ask for — and how much of it you have →
Apply

See also

AI Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available