Point your AI agent at freehire and let it find you a job.

Get the CLI →

EPAM Systems

NewBe an early applicant

Lead Data Science Engineer

Posted Updated 2 views
Discussion

We are looking for a Lead Data Science Engineer to build reusable data-sharing adapters that connect a governed lakehouse to external cloud data platforms while ensuring secure, auditable access. You will lead design and delivery across ingestion patterns, access controls, and quality gates, collaborating with engineers to ship reliable integrations.

Responsibilities

  • Design and implement a UniForm lakehouse write layer with dual-format metadata to support multiple consumers
  • Build and validate ingestion pipeline patterns for operational datasets into an analytical warehouse and lakehouse
  • Implement change-data-capture patterns with Kafka for real-time and near-real-time lakehouse updates
  • Develop dependency-aware bookkeeping and data lineage tracking patterns for pipeline observability
  • Engineer modular, version-controlled adapter components that can be reused across new data source integrations
  • Configure and validate external table access patterns for Iceberg and Delta formats in Snowflake with zero-copy reads
  • Implement tenant-scoped access controls compatible with external catalog governance requirements
  • Implement and certify a Delta Sharing adapter for zero-copy sharing to Databricks consumers
  • Configure sharing endpoints, registration, and agreement management for Databricks integrations
  • Validate end-to-end data freshness and sharing latency against defined SLA targets
  • Register connector types and implement RBAC and tenant-scoped authorization for governed data-out paths
  • Integrate metering hooks compatible with billing requirements for governed data-sharing flows
  • Apply AI-assisted development tools in daily engineering work and demonstrate effective usage during technical review

Requirements

  • 5+ years of data science or ML engineering experience with Python, pandas, and scikit-learn
  • Experience building data pipelines with SQL and Git-based workflows
  • Experience delivering production-grade RAG applications development using embeddings and chunking strategies
  • Proven leadership experience guiding technical design and code quality across shared components
  • Strong project delivery skills with reusable integration patterns and clear documentation
  • Solid software engineering skills in modular design, testing, debugging, and API/system design tradeoffs
  • Hands-on ML operations knowledge including experiment tracking, monitoring, drift detection, and performance metrics
  • Strong LLM fundamentals across tokenization, context windows, sampling parameters, and evaluation methods
  • Proficiency with AI-assisted development tools such as Claude Code, GitHub Copilot, Cursor, or similar
  • Upper-Intermediate English proficiency (B2)

Nice to have

  • Google Cloud Platform experience with BigQuery and cloud storage fundamentals
  • Large Language Models (LLM) integration experience with API usage, streaming, and cost controls
  • Vector database experience for retrieval and similarity search workflows
  • Skills integrating with Snowflake and Databricks data-sharing mechanisms

Benefits

  • International projects with top brands
  • Work with global teams of highly skilled, diverse peers
  • Healthcare benefits
  • Employee financial programs
  • Paid time off and sick leave
  • Upskilling, reskilling and certification courses
  • Unlimited access to the LinkedIn Learning library and 22,000+ courses
  • Global career opportunities
  • Volunteer and community involvement opportunities
  • EPAM Employee Groups
  • Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn

Skills

See also

Data Science jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available