Data Scientist
Summary
A Data Scientist at Oxylabs in Vilnius owns an applied ML backlog on the company's web data platform, building entity resolution, LLM-based structured extraction, and topic modeling to turn unstructured scraped data into product-ready outputs. Core stack: Python, SQL, Spark, plus Dagster, dbt, Superset, and Trino.
The team waiting for you:
Our Engineering team manages a powerful data platform that has unlocked a growing backlog of applied ML work. As a Data Scientist, you will take dedicated ownership of this backlog, driving initiatives like entity resolution, LLM-based structured extraction, and topic modeling. We are rapidly scaling our source acquisition and ingestion workflows, making this the perfect time to join and shape our capabilities. By turning unstructured data into structured, product-ready outputs, you will solve the core challenge of ensuring data accuracy and reliability at scale.
In this role, you will:
- Develop and maintain entity resolution techniques to match and link records across sources with inconsistent identifiers.
- Prototype, build, and evaluate LLM and ML-based extraction and inference models that turn unstructured or free-text data into structured outputs.
- Automate scraping logic and source queue generation using data-driven models to streamline and scale source acquisition workflows.
- Build, iterate on, and monitor applied ML models to proactively identify data or concept drift.
- Enrich and enhance datasets with product-ready calculated fields and derived metrics.
- Assist in building automated or ad-hoc QA validation processes to validate model output accuracy, reliability, and consistency.
Your skills & experiences
- 4+ years of previous experience as a Data Scientist, Data Analyst, or Data Engineer, with a strong background in data modeling, ML, and NLP.
- Excellent programming skills in Python, proficiency in SQL, and hands-on experience with Spark.
- Deep, structural thinking with the ability to communicate clearly and align with various stakeholders.
- A collaborative, self-driven mindset with a keen attention to detail and a propensity to dig into deeper layers to inspire improvements.
Nice to Have:
- Experience with Dagster, dbt, Superset, or Trino.
- Previous experience working closely in a team with data engineers.
- Strong business acumen and an understanding of how insights convert to value and unlock new revenue.
- Excellent written and spoken English.
Tech stack:
- SQL,
- Python,
- Spark,
- Dagster,
- dbt,
- Superset,
- Trino.
Salary & Benefits:
- Gross salary: 3500 - 6800 EUR/month. Keep in mind that we are open to discuss a different salary based on your skills and competencies.
- Growth & Learning: 40+ internal learning options, external conferences, mentorship, and year-round knowledge-sharing.
- Health & Well-being: Private health insurance, psychotherapy, on-site well-being consultants, 24/7 gym access, and a wellness app.
- Celebration & Community: Team events, an overseas workation, quarterly team-building budgets, and plenty of ways to mark milestones together.
Skills
As published by lever · 4 questions · 3 written answers
Basics
Resume/CV, Full name, Email, Phone, Current location, Current company, Link to your profile (Linkedin/Github etc.) URL
Pick from a list (1)
- Would you be able to work at least 3 days per week at the Vilnius office?
Written answers (3)
- What specific algorithms, techniques, or Python libraries have you successfully used for entity resolution when merging data with inconsistent identifiers? optional
- Name 1-2 specific methods or tools you use to automatically validate and QA structured data extracted by LLMs to prevent hallucinations in production. optional
- What specific metrics, tools, or approaches do you rely on to actively monitor for data or concept drift in deployed ML models? optional